Krishna's Blog — AI, Technology & Software Engineering

Decision Models Explained: Jev vs Laya vs OpenAI Decisions API (and the Rest of the Field)

What is a decision model?

Most AI models you know write text. A decision model doesn't. It answers a question your code can branch on, with a number that says how confident it is.

Think of it this way. Your support agent gets a ticket. It needs to know: which team owns this, and how urgent is it? A chat model would write a paragraph explaining its reasoning. Your code then has to parse that paragraph, handle the case where it rambles, and retry when it ignores your format.

A decision model skips all that. You send the ticket plus two typed questions: "which team?" (a choice among Billing, Technical, Sales) and "how urgent?" (a score from 1 to 5). It returns two answers with probabilities. One request, no parsing, no retries. Your code checks the confidence and routes the ticket.

The industry calls these "System One" models — fast, automatic judgment, like the instinctive thinking a human does without deliberation. They're not meant to replace big generative models. The generative model still drafts the reply. The decision model makes the small calls in between: which tool to invoke, whether to escalate, whether an action is safe.

The pattern is simple: state goes in, typed decisions come out. A typical pipeline looks like state + bounded question → typed decision → policy check → action, instead of the old prompt → generated text → parser → schema validation → retry → action. Every step you remove is a class of bugs that disappears.

The tradeoff is real too. A decision model answers one question well, but it can't reason through a complex problem, and it can't explain itself in words. Use it for judgment, not for thinking.

The big three

Jev: the one that started it

Jev is TypeSafe AI's System One model. The startup came out of stealth on September 15, 2026 with $40M in seed funding, and its co-founder Diogo Almeida is a former OpenAI researcher who worked on InstructGPT.

You send Jev a state (text or JSON) plus typed questions. Three primitives: Choice (pick one option), Score (a position on a scale), and Noul (a yes/no probability). It evaluates them in parallel and returns one structured answer per question, with probabilities. The name comes from Jevons Paradox: make judgment cheap, and software will use it everywhere.

Pros: First mover with a real ecosystem — MCP servers, Claude Code plugins, and agent tools already route micro-decisions through it. RLCD-calibrated confidence. Pricing runs $0.042 per million input tokens on OpenRouter ($0.0578 via AI/ML API — it varies by provider), output free. Strongest zero-shot accuracy in independent benchmarks (0.932 in Respan's community run).

Cons: Cloud only — your data leaves your infra. Closed weights, no fine-tuning announced as of late September. Launched in early access behind a waitlist, so availability was gated at first. Vendor lock-in.

Laya: the open-weight challenger

Laya, from Convai Innovations, released September 18, 2026 under Apache 2.0. It speaks the same typed-question language (Choice/Score/Noul) over a Jev-compatible API — some tools even fall back from Jev to Laya.

The difference is it's open and local. The English checkpoint is a 421M ModernBERT backbone; a multilingual checkpoint is smaller still, running on CPU in under 1 GB of RAM. Reported latency is about 33ms per decision on a Tesla T4 GPU — roughly 7x faster than Jev's 236–276ms measured on the same hardware, credited to non-autoregressive scoring.

Pros: Apache 2.0, runs on your own hardware. Fine-tunable on your own data. The model card reports ECE dropping to 0.081 after fitting calibration temperatures (from 0.466 as shipped — so calibrate before thresholding). Zero inference cost after setup. On the fine-tuned typed-decisions benchmark, Laya's own checkpoint scores 0.766 versus Jev's 0.727.

Cons: Small models, and the honest catch: zero-shot accuracy is 0.362 — near random — so you must fine-tune for your task. The authors say it plainly: "a fast base to specialise, not a zero-shot decision engine." Context capped at 512–1024 tokens, so decisions must be small. Its own README admits Jev does better on questions with 50+ options, and that option order in the prompt can flip answers. The ecosystem is growing but thin.

OpenAI Decisions API: the incumbent moves in

OpenAI announced the Decisions API at DevDay 2026, about two weeks after Jev and Laya. It focuses the GPT-6 Luna model on questions you define, each with a finite list of allowed answers. Pass context as text or images, get back a structured choice.

Stated use cases: content moderation (allowed / needs review / blocked), support routing, and agent control flow (pick the next tool or step from a fixed menu). OpenAI demonstrated it driving a computer-use agent on the keynote stage at sub-second speed.

Pros: Same vendor, same billing, same API keys if you're already on OpenAI. Handles images natively (Jev is text-only). Launch coverage cites ~150ms response times versus ~1.6s for a standard Luna call.

Cons: Limited preview only — as of October 2, OpenAI's own pricing page showed no Decisions API row, the docs returned 404, and a live probe of the endpoint returned "not enabled for this user," so neither price nor limits are confirmed yet. Also the least field-tested of the three. Community pushback is fair: constraining the answer space guarantees format, not correctness. You still need to evaluate it on your own tasks.

The rest of the legit field

Cloudflare Clef and Clef-flash. 27B and 9B models on Qwen bases, Apache 2.0 weights, handling text, JSON, images, and video with 65k context. Priced at $0.24 and $0.09 per million input tokens. Cloudflare calls them fully Jev-API compatible (drop-in replacement), and on its own benchmarks Clef-flash answered in a median 38.8ms against 524.1ms for Jev — though Jev still won two rows (When2Call, BRIGHT), and no independent head-to-head exists yet. The heaviest open-weight option for teams that want both size and self-hosting.

Perplexity pplx-decider-v1-27b. A 27B Qwen-based decision model with weights on Hugging Face under an Apache 2.0 tag, a context limit near 262k, and hosted pricing of $0.04 per million tokens — slightly undercutting Jev. The context king of the field. (Its 85.71% benchmark figure appears only in a company X post — unverified.)

AWS Strands Decider 2B. From AWS Strands Labs, a 2B Qwen3.5-based model with Apache 2.0 weights — and it ships with training data and scripts, so you can see exactly how it was taught. Self-hosted only, median 115ms on an RTX 3090, 167 of 231 on the public JevBench set. Best for teams that want a transparent, reproducible starting point.

Fastino GLiDE and GLiNER2.5-Decide. Fastino Labs has two entries: GLiDE, a closed API model that does adaptive thinking on uncertain cases, and GLiNER2.5-Decide, an open 340M DeBERTa-v3 encoder that runs on CPU (38ms on a V100, 167ms on a 48-core CPU) — including air-gapped setups. The enterprise air-gap answer.

Together AI Tev1 4B. A Jev-inspired Qwen3.5-4B fine-tune with the training recipe published — the team reported it cost $17 to train. Open weights, and Together AI hosts it too (per-model pricing not published). Caveats from its own README: it's experimental, answers with a single letter rather than Jev's typed format (so Jev client code needs an adapter), caps questions at 2–24 options, and its reported scores come from reused development benchmarks. Scored 0.901 accuracy in Respan's independent run — the closest open challenger to Jev.

Kev. From Jared Palmer (creator of Turborepo). Open, Apache 2.0, built on Qwen2.5-0.5B with a LoRA adapter and a learned readout head, answering yes/no, choice (2–255 options), and score questions in one forward pass — roughly 160ms for six questions on an Apple M5. The most transparent evaluation of the open bunch, with training code and a pre-registered test. Scored 0.833 accuracy in Respan's run. Caveat: calibration holds only in-distribution, so unfamiliar inputs deserve your own test set.

AnyJev. From Nokia Applied Research, Apache 2.0. It doesn't train a new model — it turns any LLM into a Jev-style decision model, with closed-form calibration from as few as 100–300 labeled examples and no retraining. The pragmatic choice if you already have a model you trust.

Head to head

Jev Laya OpenAI Decisions API
Vendor TypeSafe AI Convai Innovations OpenAI
Released Sep 15, 2026 Sep 18, 2026 DevDay 2026 (limited preview)
License Proprietary cloud API Apache 2.0, open weights Proprietary cloud API
Size / architecture Not disclosed 421M ModernBERT (EN), smaller multilingual GPT-6 Luna, constrained decoding
Input Text only Text / JSON state Text and images
Latency 70–500ms (per TypeSafe); 140ms measured server-side ~33ms on T4 GPU, CPU option ~150ms (per launch coverage)
Price $0.042/M input on OpenRouter (varies by provider), output free Free (self-hosted) Not published as of Oct 2
Confidence RLCD-calibrated 0.081 ECE after fitting (0.466 as shipped) Not publicly detailed
Context 32K tokens per request (per spec table; one source cites 64K) 512–1024 tokens Not publicly detailed
Ecosystem Richest (MCP, plugins, agents) Growing, Jev-compatible OpenAI platform
Best for Agent judgment at scale Private/on-device decisions, fine-tuning Teams already on OpenAI

And the wider legit field at a glance: Clef (27B, $0.24/M, Jev-API compatible per Cloudflare's own benchmarks) for max capability; Perplexity decider (27B, 262k context) for long inputs; Strands Decider 2B (transparent recipe, 115ms on RTX 3090) for reproducibility; GLiNER2.5-Decide (340M, CPU/air-gap, community Jev-API wrapper) for locked-down environments; Tev1 4B (0.901 accuracy in Respan's independent run) as the strongest open challenger to Jev; Kev (0.5B, 0.833 accuracy) for laptop-friendly local runs; AnyJev for wrapping your own existing LLM.

Which one should you pick?

Start with one question: where does the decision live?

Already on OpenAI? The Decisions API is the path of least resistance — same vendor, same billing. But it's preview-only and the least battle-tested. Wait for broad release, or pilot it on low-stakes routing first.

Decisions touch private data or the edge? Laya, GLiNER2.5-Decide, or Strands Decider. Local inference means nothing leaves your network. The honest price is evaluation effort: fine-tune and calibrate on your own labeled data before trusting the numbers.

Need it now, at scale, in an agent? Jev has the ecosystem and the best zero-shot numbers. The price is cloud dependency and $0.042/M input tokens.

Want maximum control for minimum cost? Tev1 or Kev with the $17-style training recipe: open weights, your data, your hardware.

The bigger story matters more than any single winner. Agents spent 2025 writing essays to make choices. In late 2026 they started answering in structured decisions — and the category went from one product to a dozen vendors in about two weeks. The pattern to remember is the same regardless of which you pick: decision first, generation second.

Disclaimers

  • This post is not financial advice and not an endorsement of any vendor. It is a framework for comparing decision models, written from the information available on October 6, 2026.
  • Accuracy and calibration figures cited are from community and vendor benchmarks, not a single independent comparison across all models — no such benchmark exists yet. Evaluate on your own tasks before trusting any number.
  • Prices, availability, and preview status change without notice — check the vendor page before you build.
  • The vendor names mentioned in this post are not sponsors of this blog, and no affiliate or referral links are included. Product names are trademarks of their owners and are used here for identification only.

References

  1. What Is Jev AI? A Decision Layer for AI Agents — Hugging Face
  2. Jev: The AI That Decides Instead of Talks — Sivaramarao Mondi, Medium
  3. The Ultimate Jev Resource List 2026 — scriptbyai
  4. Laya AI Model: How It Works, Run It Locally, and Evaluate It — Hugging Face
  5. Laya: The Open-Source AI Model That Decides Before It Generates — YouTube
  6. Laya open-source repo — Convai Innovations, GitHub
  7. OpenAI DevDay 2026 for Developers — thewebtier
  8. OpenAI Introduces Decisions API — digitalmarketreports
  9. OpenAI's Decisions API vs Jev: Inside the Decision-Model Architecture — Firecrawl
  10. OpenAI launches Decisions API to take on Jev — Pasquale Pillitteri
  11. What Is Jev? TypeSafe's Decision Model, Tested Against LLMs — aimlapi
  12. Jev Alternatives: Clef, Perplexity and Strands Decider — cellcog.ai
  13. Decision AI Models Explained — marktechpost
  14. 10 Best Jev Alternatives Compared in 2026 — respan.ai
  15. Laya model card — convaiinnovations/laya, Hugging Face
  16. Convai ships Laya — aiweekly.co
  17. Laya release notes — ai-tldr.dev
  18. OpenAI Decisions API pricing: the real cost per decision — eesel.ai
  19. OpenAI Decisions API explained — eesel.ai
  20. OpenAI's Decisions API: GPT-6 Luna Picks One Answer — orcarouter.ai

Comments & Reactions