Grammar Is Discovered, Not Designed
Cirrus Tempo Engineering • April 15, 2026
The instinct, when you decide your AI system needs a grammar, is to sit down and write it.
This is understandable. A grammar feels like rigor. A YAML file with well-named sections feels like engineering. The act of defining rules before you start generating gives you the comfortable sense that you have thought carefully about the problem. You have, in a way — but you have thought carefully about the problem as you currently understand it, which is not the same thing as the problem as it actually is. The domain you are trying to formalize has not finished teaching you yet. Writing the grammar now means formalizing ignorance.
This essay is about what to do instead: observe first, correct as you go, and treat the grammar as the last artifact you produce from a domain exploration — not the first.
The grammar is always a model of your understanding
Every rule in a domain grammar is a claim: this combination of things is legal; that one is not. The claim is only as good as the evidence behind it, and when you write a grammar at the start of a project, the evidence is thin. You are encoding assumptions — about how roles compose, about which failure modes are distinct, about which boundary cases are worth naming — that the domain has not yet confirmed or corrected.
The trouble is that a premature grammar does not feel premature. It feels like architecture. It sits in your repository, its rules are cited by your AI prompts and your validators, and everything downstream is written in terms of it. The grammar has acquired the authority of infrastructure before the domain had a chance to correct it.
When the correction eventually comes — and it always does — you have a choice between two expensive options: refactor the grammar and everything that cites it, or leave it wrong and route around it. Both happen in practice. Neither is cheap. The second is worse, because routing around a grammar you know to be wrong means you are now maintaining a lie at the center of your system.
The alternative is not to skip the grammar. It is to earn it.
How a vocabulary migrates — and what it carries with it
The first grammar-like artifact for this platform was not a task grammar. It was a disambiguation framework built around an ambiguity funnel — an attempt to formalize how a system navigates from a vague business need down through levels of crystallizing specificity to an executable integration. That is a problem-solving concern: how do you recognize and narrow what someone is trying to do? Intent — Push, Synchronize, Webhook — was the vocabulary it started with, because intent is the first thing that becomes legible when a business need starts to crystallize.
That framework was largely unsuccessful as a disambiguation tool. The problem turned out to require different machinery than a static grammar could provide. But the attempt was not wasted: it produced an intent vocabulary that was precise enough to be useful, and that vocabulary — the named intents, their rough definitions, the early intuitions about how they differed — migrated into the earliest forms of the task grammar.
In its original problem-solving context, intent was an interpretive signal — something you were clarifying and narrowing toward. When intent migrated into the task grammar, it carried that framing with it. The reasoning felt sound: if you know the intent, you know what you are building — the roles that participate, how many of each are allowed, which execution lanes are active, what operations each role can perform. Intent as a complete lookup key. One concept, one section, all the rules.
The problem is that in a task-execution context, intent is not a signal you are interpreting — it is a constraint you are enforcing. And enforcing requires structural precision that interpreting does not.
Cardinality turned out to be intent-independent. PrimarySource is always exactly one, regardless of whether the integration pushes, synchronizes, or receives a webhook. That rule has nothing to do with intent — it is a structural constraint on the integration itself. Encoding it under intent meant either duplicating it across every intent definition or silently getting it wrong when a new intent was added.
Lane activation could not be expressed as a simple property of intent either. Push always activates the Insert lane only. Synchronize activates Insert unconditionally, and Update and Delete conditionally — controlled by operator sync options set at configuration time. That conditionality had no place in a grammar organized around intent lookups.
Role operations were the deepest problem. PrimaryTarget in the Insert lane performs an Insert; in the Update lane it performs an Update. The operation is determined not by the role and not by the intent, but by the lane — a third axis the grammar had no way to represent.
The four sections that exist in the grammar today — intent_roles, role_cardinality, intent_lanes, role_operations — are not a decomposition that was planned upfront. They are the four axes the domain insisted on, one by one, as execution scenarios ran into the limits of a vocabulary that had been built for interpretation, not enforcement.
What the domain looks like before the grammar exists
Before the grammar for this platform existed, the rules lived in four or more files simultaneously. Role/operation/cardinality constraints were duplicated across the planning service, the wizard, the materializer, and the validator. Drift between copies was not an accident — it was the expected state. File A said role X allows operation Y; file B said it doesn't. Whichever file the current code path happened to consult determined the behavior.
This looks like a mess. It was not. It was the domain in the process of teaching itself what it needed. The rules that kept appearing in every file, consistently, were the rules the domain had actually stabilized on. The rules that drifted were the ones that hadn't been fully resolved yet. The duplication was painful, but it was also information: when two copies of a rule agree across every file, that rule is probably right. When they disagree, the rule is still being worked out.
The SSOT grammar — ~430 lines, 13 sections — was not designed from first principles. It was distilled from four files' worth of accumulated rules, pruned of the ones that had never stabilized, and committed only after the domain had essentially written the same rule in multiple places independently. The YAML was the last step of a process that the duplication had already completed. Which is another way of saying: the grammar was written after it was already true.
How the domain teaches the grammar: a concrete arc
The clearest demonstration of this principle is in the execution grammar, and it happened in a single day.
The execution grammar gives the AI a closed set of anchor codes to cite when analyzing a failed integration run. The catalog covers HTTP status codes, network-layer failures, mapping errors, auth failures. The discipline is strict: the AI may only cite codes from the catalog. It may not invent strings like HTTP_401_TOKEN_EXPIRED — strings that look plausible but that no downstream consumer can reliably key off of.
On a Tuesday in late May, a deliberately misconfigured integration kept failing. The error message, repeated across every record in the trace, was literal and unambiguous: "The SSL connection could not be established, see inner exception." The operator clicked Diagnose.
The analyzer's first response cited HTTP404. The trace contained no 404 anywhere. The model had asserted "the trace explicitly includes an HTTP 404 error" — a fabrication. It was force-fitting the nearest available anchor to evidence that did not support it.
The second attempt cited NET002 — timeout. There was no timeout. The model quoted elapsedMs and configuredTimeoutMs as if those fields had appeared in the trace. They had not appeared at all.
Two attempts, two different wrong answers, for the same SSL failure that the trace had named in plain English.
The reason was a catalog gap. The catalog had NET001 (TCP refused before any bytes exchanged), NET002 (request sent, no response), HTTP5XX (response received, server error). A TLS handshake failure happens after TCP connects and before any HTTP request is sent. It did not fit any of those codes, and its remediation — certificate chain, protocol mismatch, cipher negotiation — is completely different from the remediations the nearby codes imply. With no correct anchor, the model had no correct answer to give.
A third attempt, after a prompt change taught the analyzer to say "I don't have an anchor for this" rather than force-fit, produced PLATFORM001 at low certainty with a verbatim quote of the SSL exception. This was the correct behavior given what the catalog contained. It was also the first honest response in the sequence.
Then NET003 was added:
- code: NET003
category: connectivity
severity: error
persistence: unknown
origin: target
emission: runtime
summary_template: "TLS/SSL handshake with {systemName} failed before any HTTP exchange"
evidence_keys: [systemName, uri, exceptionMessage, innerExceptionMessage]
The fourth attempt, run on the same trace, produced this:
[ERROR] NET003 — TLS/SSL handshake with UtilityApp failed before any HTTP exchange (cert, chain, protocol mismatch)
- Why this code: The error message explicitly states, "The SSL connection could not be established, see inner exception," and the inner exception further details, "Cannot determine the frame size or a corrupted frame was received."
Every evidence key cited was a verbatim quote from the trace. The suggested next step named the certificate chain, supported protocols, and endpoint configuration — the right things, in the right order.
The domain had contained this failure mode for as long as TLS-secured connections existed. The grammar did not contain it until the domain had forced the issue: two wrong answers, one honest unknown, one catalog edit. The grammar learned what it needed to know from the sequence of corrections, not from upfront design.
Corrections flow into the grammar, not just into the model
The NET003 arc is clean because it involves a missing anchor — the gap is obvious in retrospect. The harder case is when the grammar has the right anchor but the wrong evidence surface: when the anchor exists, the model picks it correctly, and still gives a wrong answer because the evidence it is given to cite is incomplete.
The MAP001 anchor covers Python mapper expressions that raise exceptions. The catalog was written with pythonExceptionType and pythonExceptionMessage as evidence keys. Those two fields are sufficient to identify that the mapper failed and what kind of exception it was. They are not sufficient to answer the next question the operator always asks: what specifically do I fix?
In one diagnostic session, the failing expression was:
outbound['id'] = xfr(inbound['UtilityAppEndpointListEntry']['UtilityId'])
The exception was NameError: name 'xfr' is not defined. The analyzer correctly cited MAP001 — anchor selection was perfect. The operator then asked: how do I trim the field in the mapping?
The analyzer responded with prose: "use Python's strip() method or another appropriate string manipulation function."
The operator was staring at a literal expression in the UI. The analyzer had that expression in the trace. The correct answer was one line:
outbound['id'] = inbound['UtilityAppEndpointListEntry']['UtilityId'].strip()
Instead the operator received a description of what to do and was left to translate it back into the syntax they were working in — the exact translation the analyzer was positioned to perform, because it already had the expression, the variable names, and the accessor paths.
Two things changed in the catalog after this session. First, undefinedIdentifier was added as an evidence key to MAP001, giving the model a canonical slot for "the literal name that was undefined." Second, a few days later, nullSubKey was added after a different incident — a TypeError of the shape float() argument must be a string or a real number, not 'NoneType' where the accessor whose value resolved to None was the actionable detail.
Neither of these evidence keys was predictable from first principles. Both emerged from watching the analyzer fail and asking: what specific piece of information, had it been canonical, would have aimed the operator at the right fix in the fewest words?
The grammar's shape is the accumulated answer to that question, asked incident by incident.
Two wrong positions before the right framing
It is not only individual anchors that emerge this way. Sometimes it is the grammar's entire structural approach to a problem.
The investigation grammar uses a hybrid emission model: some anchor codes are emitted by the runtime (when the evidence is unambiguous — an HTTP 401 is a 401), and some are chosen by the AI from the closed catalog (when the mapping from raw evidence to code requires judgment). The partition seems obvious once you see it. It was not obvious when the grammar was being designed.
Two extreme positions were tried and rejected first.
"Runtime emits everything" required a giant switch statement to classify cases the runtime genuinely cannot classify. TaskCanceledException on an HTTP call — was it a configured timeout firing, a remote hang, or upstream throttling silently closing the connection? The runtime sees one exception type for all three. To emit a stable code, the runtime would have had to invent classification heuristics dressed up as facts.
"AI chooses everything" invited fabrication. When the model was free to invent codes in early prompt experiments, it produced strings like HTTP_401_TOKEN_EXPIRED and OAUTH_CLIENT_REVOKED — plausible-looking, internally consistent, and invisible to every downstream consumer that was keying off a closed catalog.
The hybrid resolved the tension: runtime emits codes it can prove from cheap signals, AI picks from a closed set for everything the runtime cannot classify without judgment. Neither position was in the grammar at the start. Both wrong positions had to be tried, and their failure modes had to be observed, before the right framing became legible.
When the grammar is not the problem
There is a failure mode that is easy to confuse with a grammar gap but is not one: the evidence never reaching the grammar at all.
On the same day as the NET003 incident, the analyzer was given a trace where the platform log contained a full SSL stack trace — exception type, inner exception, line numbers — but none of that text appeared in the trace the AI was given. The flow executor had been calling the HTTP method with no surrounding try/catch at the stage level, so the exception had bubbled past the per-step handler to an outermost catch that appended one line of ex.Message to a top-level message list. The full SSL text was swallowed.
The grammar was correct. The escape valve fired. The operator saw "uncertain, here is the raw exception text."
The fix was not a catalog edit. It was a production code change: a try/catch wrapping the HTTP call, capturing the exception into a structured ExecutionStage with category, persistence, and origin from the runtime classifier. Once the evidence was reaching the grammar, the grammar could do its job.
The lesson is narrow but important: before concluding that your grammar is incomplete, verify that the evidence the domain is producing is actually flowing into the prompts and validators that consume the grammar. A grammar gap and an evidence plumbing gap produce the same symptom — wrong or uncertain answers — but have entirely different fixes.
When the domain is ready to be formalized
Resisting the grammar does not mean never writing one. It means writing it at the right moment, which is later than the instinct suggests.
The signals that a domain is ready to be formalized are observable:
The same rule keeps appearing in different places, written independently, and consistently. When you find yourself writing the same constraint in the planner, the wizard, and the validator without coordinating between them, that constraint has stabilized.
You can explain why something is wrong without looking it up. If a proposed plan violates a rule and you immediately know which rule and why, that rule is real. If you have to reason it out from first principles each time, it is not yet a rule — it is a heuristic you are still developing.
The informal rules and the actual system behavior have diverged far enough to be painful, and you know which informal rules are right. The four-file drift was painful. The direction of the fix was known: the rules that had stabilized across files were the ones worth keeping. The pain of the drift and the clarity about which rules were real arrived at the same time. That is the moment to formalize.
The health grammar's own header puts it directly: "no speculative grammar." Every section of the grammar that exists today was demanded by something that happened, not anticipated by something that might.
The grammar as accumulated correction
The 430-line composition grammar did not emerge from a requirements session. It emerged from the deleted IntegrationPattern enum that could not hold metadata, from four files' worth of duplicated rules that drifted into inconsistency, from the planning patterns that decompose at the wizard boundary in ways that only became clear once the materializer existed to decompose them.
The execution grammar did not emerge from an architecture diagram. It emerged from HTTP404 cited against an SSL failure, from elapsedMs quoted from a trace that contained no timeout, from an operator asking how to trim a field and receiving prose when they needed syntax.
Every section that exists earned its existence by being needed. Every evidence key that is now canonical was absent when something specific went wrong and was added because that specific thing going wrong had a specific piece of information that would have made it right.
This is what it means for grammar to be an emergent property: not that it arises spontaneously without effort, but that its shape is determined by the domain's behavior over time rather than by your predictions about that behavior before you start. You do not discover the grammar by sitting with the domain in theory. You discover it by running, observing, correcting, and reading the corrections back into the grammar as the evidence accumulates.
Resist the notebook. Let the corrections come first.