Two infrastructure vendors shipped open-weight decision models in the same week. Here is where a calibrated structured-choice model actually belongs in a production agent workflow — and the three places it will make things worse.
Last Updated: October 6, 2026
Decision models are a real architectural component, not a smaller LLM. They belong in exactly four places in a production agent: intent routing, tool selection, confidence gating before an irreversible action, and exception triage. They will hurt you in three: anything requiring extraction, anything with an open-ended option set, and anywhere you trust the confidence score without calibrating it on your own labels. The cost saving is real but small — roughly $500 per million decisions against a Flash-class baseline. The latency and the calibrated confidence score are the actual reasons to adopt. Below a few hundred thousand decisions a month, the second model costs more to govern than it saves.
On 1 October 2026, two infrastructure vendors independently shipped the same product category. Cloudflare released Clef and Clef-flash, open-weight decision models hosted on Workers AI with an RL fine-tuning path attached. The same day, AWS's Strands Labs released Strands Decider 2B, a 2B open-weight model that runs on a laptop and returns confidence-scored choices. Both are downstream of TypeSafe AI's Jev, launched two weeks earlier on 15 September.
The launch recaps all say the same thing: decision models are fast and cheap. True and uninteresting. The question an operator actually has is narrower — which specific call in my workflow should stop being an LLM call? That question has a defensible answer, and it is not "all of them".
What a decision model gives up, and why that is the point
The architecture is the honest starting point. Strands took a pre-trained LLM torso and removed the language-modelling head entirely, replacing it with a pointer head of just over a million parameters that scores each offered option against the hidden state at the answer position. Cloudflare's approach is adjacent: a prefill-only pass over a frozen Qwen backbone, then parallel scoring of the valid schema choices, with no autoregressive generation step at all.
Three consequences follow mechanically, and they are the whole value proposition:
It cannot emit an invalid option. The output space is the schema. There is no parsing step, no retry-on-malformed-JSON, no "the model returned a tool name that does not exist". TypeSafe makes the strong version of this claim and is right about the mechanism — schema matching is structural rather than empirical.
It returns a probability per option, trained as a probability. Cloudflare post-trained Clef with a Brier loss specifically to refine probability calibration, on top of label-smoothed cross-entropy. Strands measures calibration as a first-class target alongside accuracy, reporting Brier score on JevBench's public set across the training trajectory. An LLM asked "how confident are you, 0 to 1?" is doing something categorically different and worse.
Multiple questions about one state cost roughly one pass. This is the under-discussed property. Cloudflare's threat intelligence team classified a rendered domain in 2.2 seconds end-to-end with Clef against 4.7 seconds for gpt-oss-120b — and got more categories back, because the multi-category answer came out of a single scoring pass rather than a generation loop.
What you give up is everything else. Strands is blunt: the single parallel pass makes it significantly worse at complex problems than reasoning models, and unsuited to coding, chatbots, or summarisation. It is a typed function call with an opinion, not an agent.
The four places a calibrated decider beats an LLM call
Every one of these shares a shape: the option set is finite, enumerable at request time, and you need the answer in the hot path.
1. Intent routing at the front door. A user message arrives and needs to reach one of six to twenty handlers. Clef posts 94.20 macro-F1 on BANKING77 and 97.43 on CLINC150+OOS — the second number matters more, because CLINC150+OOS includes out-of-scope detection, which is the "none of these, escalate" branch you will need in production. Note that Clef-flash drops to 66.77 on that same benchmark. The cheap model is not uniformly the smaller trade-off.
2. Tool selection from a fixed registry. If your agent has thirty MCP tools, you are currently paying a frontier model to read thirty descriptions and emit one name. Clef reports 98.47 case-exact on BFCL and 69.19 nDCG@10 on ToolRet. Caveat the vendor blog does not emphasise: on When2Call — the eval closest to "should a tool be called at all right now" — Jev scores 80.97 against Clef's 72.37. Picking the right tool and knowing whether to pick one are different skills, and the category leader loses the second one.
3. Confidence gating before an irreversible action. This is the strongest placement and the one I would implement first. The Strands launch example is a pre-tool-call interception: before the agent calls get_weather, the decider answers two yes/no questions — are the argument values grounded in anything the user actually said, and is it premature to call this tool before clarifying. The handler returns a typed action: Proceed, Deny, Confirm, or Guide. The Strands team's own framing is the line worth stealing: a decision this cheap can sit in a path where an LLM call never could.
4. Exception triage on the output side. Agent run completes, trace exists, something needs to decide whether a human looks at it. This is a bounded scoring problem over a long input, which is what a 64k-context decider is for. Clef's 65,536-token context window comfortably holds a full trace.
Where the decider sits in one request
Message, session context and tool registry are assembled by ordinary code. No model involved.
One call, several typed questions: which handler, which tool, is this in scope. Probabilities returned per option.
Above threshold, proceed. Below, escalate to the LLM or to a person. The threshold is config, not a prompt.
Extraction, generation, argument construction, multi-step planning. This is the part a decider cannot replace.
Grounded arguments? Premature? Policy-compliant? Proceed / Deny / Confirm / Guide before anything irreversible fires.
The three places it will quietly make things worse
"Quietly" is the operative word. None of these fail loudly. They fail as a slow drift in your escalation rate that nobody attributes to the router for six weeks.
Trusting the confidence score because the vendor trained for calibration
Calibration is a property of a model on a distribution. Cloudflare calibrated Clef against internal synthetic datasets that permute field orders, prompts and schema structures; Strands measures Brier on JevBench's public set. Neither distribution is your support inbox. A model that is well calibrated in aggregate can be systematically overconfident on the one intent class that generates your refunds. Teams pick a 0.8 threshold off the launch blog, ship it, and discover months later that 0.8 on their traffic means 0.62 accuracy on the tail classes.
Anything that requires extraction. A decider picks from options and assigns scores. It does not pull an invoice number out of a PDF, normalise a date, or construct tool arguments. If your "routing" step is secretly also populating fields, you have an extraction problem wearing a routing costume, and swapping in a decider will produce a correct choice with no payload.
Anything where the option set is open-ended or very large. The architecture scores enumerated candidates, so cardinality is a hard constraint rather than a soft one — Jev caps at 255 options and falls back to a two-stage score-then-choose system above that, with a corresponding slowdown. If your tool registry is dynamic, tenant-specific and in the thousands, you need retrieval in front of the decider, and now you are maintaining a retrieval index too.
Anywhere the decision is actually a negotiation. "Should we approve this exception?" is sometimes a bounded policy check and sometimes a multi-constraint judgement that needs to read three documents and weigh them. The first is a decider. The second is an LLM, and compressing it into a choice question produces a confident answer to a question you did not ask.
What the cost math actually says
Here is the per-decision arithmetic for a realistic routing call: 800 input tokens of state plus schema — a conversation snippet and a tool catalogue. Published list prices, my modelling. The LLM baseline is gpt-oss-120b on the same Workers AI platform at $0.35 per million input tokens and $0.75 per million output, which is the comparison Cloudflare itself drew.
| Option | Cost per 1M decisions | Published latency | What you give up |
|---|---|---|---|
| Clef-flash (9B, Workers AI) | ~$72 | 38.8 ms median / 122.4 ms p95 | Out-of-scope detection drops sharply vs. Clef (66.77 vs 97.43 macro-F1) |
| Clef (27B, Workers AI) | ~$192 | 209.3 ms median / 238.6 ms p95 | 5x the flash latency; still no text output, no extraction |
| Strands Decider 2B (self-hosted) | GPU-hours, not per token | ~115 ms median (RTX 3090); ~153 ms (M3 MacBook, small tasks) | You own the serving, the scaling and the uptime |
| Jev (TypeSafe, hosted) | ~$34 at $0.042/MTok, output free | 524.1 ms median in Cloudflare's run; 70–500 ms claimed by TypeSafe | Closed weights; 255-option cardinality ceiling |
| gpt-oss-120b, structured output + short trace | ~$370–$580 | Not published on this benchmark | Nothing — this is the baseline you are replacing |
Two things fall out of this table that the launch posts do not say.
First, the latency numbers in each column were produced by the vendor who benefits from them, on their own harness. Cloudflare's 524.1 ms median for Jev sits awkwardly against TypeSafe's own claimed 70–500 ms range, and TypeSafe notes their evals are run from laptops on the US West Coast. Treat all of it as directional. The p95 column is the one that will decide whether this survives in your hot path, and it is the least reported.
Second, the money is not the story. Replacing a Flash-class call with Clef-flash saves roughly $508 per million decisions on list prices. At 300,000 decisions a month — which is a genuinely busy workflow — that is about $152. You will spend more than that in the first week writing the eval harness. Anyone selling you a decision model on cost reduction at mid-volume has not done the subtraction.
Run a calibration test before you re-plumb anything
This takes about a week and costs almost nothing, and it is the difference between an architectural improvement and an incident. Run it against your existing production traffic, in shadow, before a single request changes path.
Five-day calibration protocol
Pull 300–500 real decisions your current LLM router made. Have a human label the correct answer. Stratify deliberately: include the rare classes and the out-of-scope cases, not a random sample that is 80% one intent.
Rewrite your routing prompt as typed questions. Resist collapsing three questions into one; cheap parallel questions are the economic advantage. Run both Clef-flash and Clef, or Decider 2B at two quantisations.
Bin by predicted confidence, plot observed accuracy per bin, compute Brier. Do it globally and then per class. Expected calibration error above ~0.05 on any class you care about means no auto-proceed on that class.
Not median, and not from the vendor's region. Include network. A 38.8 ms model behind a 240 ms round trip is a 280 ms model.
Pick the confidence floor that yields your required precision, then define what happens below it. If the answer is "escalate to the LLM", you have not removed the LLM — you have added a fast path, which is the correct and honest framing.
One detail that saves a round of rework: both Clef and Jev share an API shape, and Cloudflare built Clef to be fully Jev-API compatible. Write your schema against that shape and you can swap models in the test harness without rewriting the integration. The Strands CLI gives you the same thing locally — a strands-decider ask command that takes a state and a choice question and returns per-option probabilities.
pip install strands-decider
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \
--state "Help! My payouts have been failing for 3 days!" \
--choice "Which team should handle this?=billing,sales,retail"
# choice_0 -> billing (confidence 0.768)The second-model tax nobody is pricing in
A decision model is not a library. It is a second model in your stack, and it inherits every obligation the first one has: a version to pin, weights to host or a vendor endpoint to monitor, an eval suite that has to run in CI, a calibration drift check on a schedule, and an owner.
The benchmark tables make the drift risk concrete. Across Cloudflare's own published numbers, Clef-flash beats the larger Clef on home appliances (97.73 vs 82.95) and API-Bank (93.11 vs 91.93), while collapsing relative to it on CLINC150+OOS (66.77 vs 97.43). Performance is not monotonic in model size, and it is not monotonic across task types. That means you cannot reason about an upgrade from the parameter count; you have to re-run your own evals on every version bump. Strands is on v19 of its architecture, which tells you how fast this is still moving.
So the honest buying rule: adopt a decision model when you need sub-100ms decisions in a hot path or you need a calibrated confidence number to gate an action. Both of those are capabilities an LLM call cannot give you at any price. Do not adopt one because the token cost looks lower, unless you are past roughly a million decisions a month and have an eval pipeline already running.
The gating case is the one that is about to get urgent. BCG's 2026 Applied AI Index, drawn from a survey of more than 1,300 CxOs and senior leaders, found that 42% of companies expect to grant agents autonomy by 2030 while only 5% have the full set of critical controls in place today — and firms with all six controls generate three times as much agentic AI value as those with one. A calibrated gate in front of every irreversible tool call is one of the few controls that is now cheap enough to put everywhere. That is a stronger argument for this category than anything on the cost side.
If you want a second opinion on where the gate belongs in a workflow you already run, that is the kind of thing we scope in an architecture review — and our model coverage notes track how these benchmarks move. When you are ready to test one against your own traffic, start here.
Frequently Asked Questions
What is a decision model, and how is it different from an LLM?
A decision model takes a state plus a schema of typed questions and returns a probability for every allowed option, with no text generation. Strands built theirs by removing the language-modelling head from a 2B torso and replacing it with a pointer head. Cloudflare's Clef runs a prefill-only pass then scores schema choices in parallel. The result cannot hallucinate an invalid option, but also cannot extract, summarise or write code.
Can Clef or Strands Decider 2B replace my LLM router entirely?
Only if your router does nothing but choose from a fixed list. Most production routers also normalise parameters or construct tool arguments, and a decision model cannot do either. The realistic pattern is a fast path: the decider handles high-confidence routing, and anything below your threshold escalates to the LLM. You are adding a tier, not removing one.
How much does a decision model actually save on cost?
On an 800-token routing decision at list prices, Clef-flash works out around $72 per million decisions against roughly $370–$580 for a gpt-oss-120b call with structured output. That is about $508 saved per million. At 300,000 decisions a month the saving is roughly $152 — less than the engineering time to build the eval harness. Adopt for latency and calibration, not for the token bill.
Is the confidence score trustworthy out of the box?
It is trustworthy on the distribution it was calibrated against, which is not yours. Both vendors train and report calibration explicitly — Cloudflare with a Brier loss during post-training, Strands by tracking Brier on JevBench. Neither guarantees anything about your intent classes. Build a reliability diagram on 300+ of your own labelled examples, per class, before you set a threshold. Global calibration hides per-class overconfidence.
Clef or Strands Decider 2B — which should I test first?
Different jobs. Clef is a hosted 27B multimodal model with a 64k context and a vision encoder, priced at $0.24 per million input tokens on Workers AI — the right choice if your state includes images or long traces and you want someone else running it. Decider 2B is a 2B model you run on your own CPU or GPU, with training data and scripts published, which makes it the better starting point for experimentation and for data that cannot leave your network.
When should I decide not to add a decision model at all?
When your decisions need extraction, when the option set is dynamic and in the thousands, when the choice is really a multi-constraint judgement, or when your volume is low enough that a second model's governance overhead exceeds the saving. Also skip it if you do not yet have an eval pipeline — benchmark performance across these models is not monotonic in size, so every version bump needs re-testing on your own data.
Research digest
AI Research Briefing
Honest insights on AI agents, Small Language Models, and local RAG. No hype. Only when we have something worth sending.
- No hype, just measurable outcomes
- Read by 2,400+ engineers
- Unsubscribe anytime





