AI & Architecture

Two Grammars: Why Constrained Decoding Doesn't Replace a Domain Contract

Cirrus Tempo Engineering  •  March 16, 2026

The word "grammar" is doing two unrelated jobs in AI engineering right now, and the overlap is quietly causing a category error. Say "we built a grammar to constrain our AI" and one engineer hears XGrammar, Guidance, llguidance — token-level structured output. Another hears a domain rulebook that the prompt, the validator, and the UI all cite as ground truth. Both are real, both are valuable, and they are not the same thing wearing different clothes. They operate at different layers of the stack, they guarantee different properties, and — this is the part that causes the confusion — one of them cannot do the other's job no matter how far you push it.

This piece exists to put a name on the difference, because the rest of this series leans on a specific meaning of "grammar" that is easy to mishear as the popular one.


When most engineers hear "AI grammar" in 2026, they think of constrained decoding — the family of techniques behind XGrammar, Guidance, Outlines, and Microsoft's llguidance. The mechanism is elegant: take a context-free grammar, a JSON schema, or a regular expression, compile it into an automaton, and run that automaton in lockstep with the model's token-by-token generation. At each step, the automaton masks the probability distribution so that only tokens which keep the output on a legal path remain reachable. The model is not asked to produce valid JSON. It is made incapable of producing anything else.

This is a genuine engineering achievement, and it solves a real problem: it guarantees well-formedness. The output will parse. The braces will balance. The enum will be one of the listed values. For a huge class of integration headaches — the model wraps its JSON in prose, forgets a closing bracket, invents a field type that isn't in the schema — this class of tool makes the failure mode disappear entirely, at the syntax level, before a single byte reaches your system.

Notice the boundary of that guarantee, though: it is a claim about shape, enforced during generation, by an engine that has no idea what your domain means.


The other meaning: grammar as a domain's externalized truth

The grammars this series describes — one for integration composition, one for the Python mapper's vocabulary, one for platform-health diagnostics — are a different animal entirely. They are not compiled into token automatons. They are YAML documents, hand-authored, that encode a product's accumulated domain rules: which system roles may participate in which integration intents, which identifiers a Python sandbox will actually expose, which finding codes a diagnostic rule engine is permitted to emit. The composition grammar is one file, 430 lines, thirteen sections. A thin C# facade projects it into a runtime validator, into the AI's system prompt, into the UI, and — across the Python/C# seam — into a parity test that fails the build if the two sides disagree.

These grammars guarantee something a token automaton cannot reach: legality. Not "does this parse," but "is this true, here, in this domain, given everything else that's already been decided."

The gap between the two is not a matter of degree. It's a matter of kind, and a concrete example makes it vivid.


The example that separates them cleanly

Imagine constraining an integration-planning model with a perfect JSON-schema grammar. Token masking can guarantee, with total certainty, that the model emits something shaped exactly like:

{ "intent": "ReplicateChangeOrder", "steps": [ { "role": "PrimarySpine", "operation": "upsert" } ] }

Every key present, every value the right type, the enum members drawn from a fixed list. The output is, by construction, perfectly well-formed.

It can also be perfectly illegal. Suppose PrimarySpine is a real, valid EntityRole — just not one that the ReplicateChangeOrder intent permits, because that intent requires a WriteBackResponder somewhere in the chain and this plan doesn't have one. No context-free grammar over tokens can express that constraint, because it isn't a fact about this token given the previous few tokens. It's a fact about this role, given that intent, given the rest of the domain's accumulated rules about how those two concepts relate — a cross-referential, stateful, business-level claim that lives in the grammar's intent-to-role table and nowhere a token automaton can see.

This is exactly the gap GrammarConformanceValidator exists to close, and it is the gap that makes the domain grammar a contract rather than a cage. The cage can guarantee the plan is shaped like a plan. Only the contract can tell you whether it's a plan you'd actually let near a customer's data.


Four axes of the difference

Laid side by side, the distinction holds along several dimensions at once — which is a good sign that it's a real distinction and not a rhetorical one:

Where it operates. Constrained decoding lives inside the generation loop, at the syntax layer, masking tokens in real time. The domain grammar lives outside it — feeding the prompt before generation, and checked by a deterministic validator after. It never touches a token probability; it touches a finished artifact.

What it guarantees. Constrained decoding guarantees the output parses. The domain grammar guarantees the output is legal — that its claims about roles, intents, accessor names, or finding codes are true statements about a specific, versioned domain model. Parsing is a precondition for legality, not a substitute for it.

What it's made of. A constrained-decoding grammar is a generic structural shape — a CFG, a JSON schema, a regex — entirely reusable across any domain you point it at. Swap XGrammar for Guidance tomorrow and nothing about your business changes. The domain grammar is your business, externalized into an inspectable file. You cannot swap it for someone else's; doing so would mean replacing your product's accumulated rules with a stranger's.

Who reads it. A constrained-decoding grammar is read by exactly one consumer — the inference engine, at the moment of generation. The domain grammar is deliberately multi-consumer: the prompt builder cites it, the validator cites it, the UI cites it, the cross-language parity test cites it. Its value comes precisely from the fact that many things which are not the model also have to agree with it.


They are not rivals — and pushing one upstream doesn't retire the other

None of this makes constrained decoding less valuable; it makes it a different layer of defense, and a genuinely complementary one. A healthy architecture would happily use both: XGrammar or Guidance to make malformed JSON structurally unreachable, and the conformance validator to catch the well-formed-but-illegal plans that slip through that first net. The first layer shrinks the volume of garbage the second layer has to reject. It does not — cannot — eliminate the second layer's job, because the second layer's job was never about shape.

This is worth saying plainly, because it's the natural next question for anyone who has internalized the SSOT argument: if the grammar is already the single source of truth for valid plans, why not compile it straight into the decoder and make illegal output physically impossible? The honest answer is that "physically impossible at the token level" and "domain-illegal" are properties of two different objects. The intent-to-role table is a relational fact about the finished shape of a plan — which roles co-occur with which intents across an entire steps[] array — not a fact derivable from the next token given the last few. Some of that could, with heroic effort, be flattened into a sufficiently baroque CFG. Most of it can't be, without effectively re-deriving the validator inside the grammar compiler — at which point you've built the same contract twice, once in a form a human can read and once in a form only an automaton can run.

Push the domain grammar upstream and you get a smaller rejection rate at the validator. You do not get to delete the validator. The contract's job was never to make bad tokens unreachable. It was to make bad claims — about roles, about accessors, about finding codes — auditable, rejectable, and impossible to quietly disagree with across the four or five surfaces that all have to tell the truth about them at once.


The vocabulary this buys you

Once "grammar" splits cleanly into shape and truth, a lot of confused conversations resolve themselves. "Should we adopt XGrammar?" stops being a referendum on whether your architecture needs a grammar — of course it might help, at the layer it operates on — and becomes a much narrower, more answerable question: does our current failure mode look like malformed output, or like well-formed-but-illegal output? If it's the former, a constrained-decoding library is probably the right next investment. If it's the latter — if your AI keeps producing JSON that parses cleanly and still gets rejected at the validator, or keeps citing finding codes that don't exist — no amount of token masking will touch it, because the thing that's wrong was never a fact about tokens.

That's the contribution of naming the two grammars separately: not a verdict on which one to use, but the vocabulary to notice, quickly, which kind of problem you actually have.