In-depth articles on technology shaping what comes next.

Constrained Decoding: Turn LLMs Into Fast Decision Models

Constrained decoding turns an LLM into a fast classifier. Learn how logit masking, calibration, and temperature scaling make decision models work well.

A light beam splitting into five colored rays through gates, illustrating constrained token output.
Masking the vocabulary forces the model's entire output distribution through a handful of allowed gates.

Here's a cheap trick that's quietly useful: if you only care about the first token an LLM would emit, you can ask it a multiple-choice question and get the answer in a single forward pass. Constrained decoding — masking every token in the vocabulary except the handful you're willing to accept — turns a generative model into something that looks an awful lot like a classifier. No JSON parsing, no retry loops, no eleven autoregressive steps to produce eleven characters. One pass, one softmax, one answer with a probability attached to each option.

This idea has been floating around under names like "decision models" or "system one" inference, and it recently caught fire on Hacker News with a walkthrough showing how to build one from a 1.7B-parameter Qwen model in about forty lines of Python. The machine-learning greybeards in the comments were, predictably, screaming into their keyboards: that's a classifier, we've had those since the perceptron. They're right, and also slightly missing the point. What's new isn't the concept — it's that you get a zero-shot classifier out of a general-purpose language model without training anything. The interesting engineering question is when this constrained approach beats just letting the model talk, and when it quietly lies to you with overconfident probabilities.

Two Ways to Get an Answer Out of an LLM

The default mode of talking to a language model is generation. You ask a question, the model emits tokens one at a time, and somewhere downstream you parse whatever came out. If you need structured output, you bolt on a schema: JSON mode, grammar-constrained sampling, outlines-style finite-state machines. These work, but the model still walks token by token through the whole response. A trivial multiple-choice answer might take eleven decoding steps, and each step costs a full forward pass over billions of parameters.

The alternative is to never let it walk at all. After the prompt is processed, you look at the logits over the vocabulary at the final position, throw away everything except the token IDs corresponding to your options — say "A", "B", "C", "D", "E" — and softmax over just those. The argmax is your prediction; the softmax values are scores. Total cost: one forward pass, same as the prefill you were already paying for. Here's the core of it, adapted from the Qwen-based approach that's been making the rounds:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Qwen/Qwen3-1.7B"
options = ["A", "B", "C", "D", "E"]
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name, torch_dtype="auto", device_map="auto"
)
# First token the model would emit for each option
option_token_ids = [
tokenizer.encode(opt, add_special_tokens=False)[0] for opt in options
]
prompt = (
"What color is the sky?\n"
"A. Red\nB. Blue\nC. Green\nD. Purple\nE. I don't know\nAnswer:"
)
messages = [{"role": "user", "content": prompt}]
text = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True,
enable_thinking=False
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
with torch.no_grad():
logits = model(**inputs).logits[0, -1]
# Constrained decoding: softmax over only the option tokens
probs = torch.softmax(logits[option_token_ids].float(), dim=-1)
print(options[probs.argmax().item()])  # -> B

That's the whole engine. The vocabulary has over 150,000 tokens and we've reduced the decision to five numbers. No sampling temperature games at generation time, no parser that can fail, no way for the model to hallucinate an option that isn't on the list. The output space is closed by construction.

It's a Classifier, and That's Fine

Let's give the skeptics their due. What we've built is a discriminative classifier over a fixed label set — a descendent of ideas that go back to Rosenblatt's perceptron in the 1950s, through logistic regression, through every softmax-headed neural net of the deep learning era. If you squint, constrained decoding is just a linear readout over the final hidden state, which is what a classification head has always been. The MLE who's been begging their team to just train a classifier for years has every right to feel a little screamy watching this get rebranded as a "decision model."

But history also shows why the new version matters. Old classifiers were narrow: you collected labeled data, you trained, you got a model that knew one task and nothing else. The reason LLM-based systems keep eating tasks that "should" use a proper classifier is the zero-shot property — the base model has already absorbed enough of the world that a prompt is the training data. When I tested a setup like this against a CommonsenseQA holdout, the 1.7B model landed at roughly 59% accuracy with zero fine-tuning, climbing to about 62% after a quick tune on the training split. That's not state of the art, but it cost an afternoon and no labeled data of my own. A purpose-built classifier might do better; it would also have required a pipeline, a dataset, and a retraining story every time the labels change.

There's a useful framing here that borrows from the old generative-versus-discriminative debate. Generative classifiers model the whole distribution and are wasteful but flexible; discriminative ones model the decision boundary and are efficient but rigid. Constrained decoding is an odd hybrid: a generative model pressed into discriminative service at inference time. You get the efficiency of the discriminative readout with the breadth of the generative pretraining. That combination genuinely wasn't available before, even if the pieces are ancient.

Where Constrained Decoding Wins

  • Latency. One forward pass instead of N autoregressive steps. On small models this is the difference between 'fast enough for a request path' and 'needs a queue.' Practitioners running pure decision models in the browser report responses under 200ms — try that with generative JSON.
  • Structural correctness. The model literally cannot produce output outside the allowed set. No malformed JSON, no 'The answer is probably B because...' rambling, no guardrail layer needed to catch format violations.
  • Throughput. Because every request is a single pass of identical shape, batching is trivial and predictable. Generating variable-length answers wrecks your batching efficiency.
  • A score per option. You get the whole distribution, not just the winner. That opens the door to abstention logic: if the top probability is below a threshold, route to a human or a bigger model.

The abstention point deserves emphasis. A generative answer is a single artifact you either trust or don't. A probability distribution over options lets you build routing: high-confidence cases go through automatically, low-confidence ones escalate. That's the pattern behind a lot of production triage systems — support ticket routing, content moderation pre-screening, intent detection — and it's where this technique earns its keep. If your task naturally decomposes into "pick one of K labels, and tell me how sure you are," constrained decoding is almost certainly the right tool.

Where Generative Output Wins

Now the other side. The moment your task doesn't fit a fixed label set, constrained decoding falls apart. If the answer is a free-form entity, a number, a piece of code, or anything compositional, you need generation — possibly with structured-output constraints, but generation nonetheless. There's also a subtler loss: reasoning. When a model generates a chain of thought before answering, it often does materially better on hard questions. The single-pass decision head gives you no scratchpad. You're asking for system one thinking, fast and intuitive, and you're getting exactly that — including its characteristic failures.

There's also the phrasing sensitivity problem. A constrained classifier over "A/B/C/D/E" is really measuring the model's preference for the text of each option in that position. Reword option C slightly, reorder the list, or change "Answer:" to "The best answer is" and you can shift the scores. Generative answers with reasoning tend to be more robust to surface perturbations because the model has to commit to content, not just to a token. If you're evaluating either approach, perturb the presentation and see what breaks — it's a cheap robustness test and an uncomfortable one.

Two paths across a landscape, one direct and one winding, symbolizing fast versus deliberative inference.
Constrained decoding is the direct path; generation with reasoning is the long trail that sometimes reaches higher ground.

The Calibration Problem: Your Confidence Scores Are Lying

Here's the trap that catches everyone who builds one of these. You get probabilities out of a softmax, so they must be probabilities, right? They are not. They're the model's confidence that a given token comes next, which is a statement about language, not about correctness. Ask the model "Where would you most likely find a bat?" with options including "Cave" and "Baseball game," and it will assign something like 0.998 to "Cave" — an ambiguous question with no defensibly certain answer, answered with near-total certainty.

Bin the predictions by confidence on a real eval and the picture gets worse. In one CommonsenseQA run, the 0.9–1.0 confidence bucket was correct only about 70% of the time, and the 0.8–0.9 bucket barely hit 40%. A well-calibrated model should be right about 90% of the time when it says 0.9. This one is systematically overconfident — which, if you think about it, mirrors the old observation that modern deep nets are overconfident in general. Guo et al. showed back in 2017 that a plain ResNet's softmax outputs are badly miscalibrated compared to the shallow nets of the 1990s. Everything old is new again; we've just reinvented the problem at 1.7 billion parameters.

The fix, happily, is also old and almost embarrassingly simple: temperature scaling. Divide the logits by a single learned scalar T before the softmax:

def scaled_probs(logits, option_ids, temperature):
selected = logits[option_ids].float() / temperature
return torch.softmax(selected, dim=-1)
# Fit T on a validation set by minimizing negative log-likelihood
# of the correct option. For the CommonsenseQA run above, T ~= 3.8
temperature = 3.7973
probs = scaled_probs(logits, option_token_ids, temperature)

T greater than 1 flattens the distribution; T less than 1 sharpens it. Fitting one number against a held-out set turned that miserable calibration table into something honest: the 0.9–1.0 bucket now sits around 95% accuracy, the 0.5–0.6 bucket around 55%. The accuracy doesn't change at all — argmax is invariant to monotonic scaling — but the scores now mean what you thought they meant. If you intend to route on those probabilities, calibrate first or don't bother collecting them.

Is Tuning Temperature on a Benchmark Cheating?

A fair question raised in the discussion: isn't fitting T so the model looks calibrated on a benchmark a bit p-hacky? It would be, if you fit it on the test set. Done properly — fit T on a validation split, report calibration on a held-out test split — it's just one-parameter regression, and it's the standard of care in the literature for exactly this reason. The discipline that matters is the same one as everywhere else in ML: keep your splits honest, and be suspicious of any calibration claim measured on data that touched the fitting. If your deployment domain drifts from your eval domain, your T drifts too, so recalibration belongs in your maintenance loop alongside everything else.

This is really an instance of a broader disease: the gap between what a system is specified to mean and what it actually computes, a gap I've argued before is where bugs live. "Confidence" is specified as a probability of correctness; the implementation gives you next-token plausibility. Temperature scaling is a patch over that gap, not a resolution of it. Keep the distinction in mind every time you're tempted to wire softmax outputs into an alerting threshold.

A Practical Decision Procedure

When I'm choosing between the two modes now, I run through a short checklist:

  1. Is the output a fixed set of labels? If yes, constrained decoding is on the table. If no, generate.
  2. Do I need reasoning to get acceptable accuracy? Prototype both. If the single-pass head lags chain-of-thought by more than you can tolerate, generation wins despite the latency.
  3. Do I need per-option scores for routing or abstention? If yes, constrained decoding plus temperature scaling is nearly free and generation gives you nothing comparable.
  4. How big is the label set? Five options is trivial; five thousand is retrieval territory. With large label spaces, embed your labels, shortlist with vector search, and only then use the model to pick among finalists — the classifier-scaling trick that's been standard since the recommendation-system era.
  5. Will the labels change often? The zero-shot property is the whole point. If labels churn weekly, a retrained classifier is a maintenance burden; a prompt edit is not.

The model's softmax is a claim about what token comes next, not about what's true. Calibrate it, or don't trust it.

One more consideration: model size interacts with the choice. A small model's single-pass head is cheap enough to run on every request, even client-side; people are already shipping decision models that run in a browser with sub-200ms responses. A cascaded design — small constrained model for the easy 80%, large generative model for the hard tail — often beats either extreme on both cost and accuracy. It's the same instinct behind speculative decoding, applied at the system level rather than the token level.

The Recommendation

My position: if your task is genuinely a fixed-choice decision — routing, triage, intent, multiple-choice evaluation — build the constrained decision head and don't look back. The latency win alone justifies it, the structural guarantees eliminate a whole class of parsing failures, and the per-option scores give you routing logic that generation can't match. But treat the raw softmax as an uncalibrated instrument. Fine-tune on your actual task if you can scrape together even a modest dataset, fit a temperature on a clean validation split, and verify calibration on data the fitting never saw.

Reserve generative output — with structured-output constraints where you need them — for tasks that are genuinely compositional or that benefit from visible reasoning. And don't let anyone sell you a "decision model" as a novel category: it's a classifier wearing an LLM's overcoat, descended from sixty years of discriminative modeling, and it's most useful precisely when you respect that lineage enough to do the calibration work the old-timers always insisted on. The tooling is new. The discipline isn't.