AI & Architecture

Domain Grammars as a Single Source of Truth for AI-Authored Integrations

Cirrus Tempo Engineering  •  June 8, 2026

If the defining engineering challenge of product-level AI is establishing a typed contract between the model and the system, the immediate follow-up question is operational: Where does the contract live?

When product teams rush to ship AI features, they almost always pick one of three wrong answers:

  • They scatter the contract across conventions: Duplicating validation logic across prompt blocks, UI forms, and database schemas.
  • They hardcode truth into structural components: Burying domain rules deep inside code execution paths.
  • They declare that "the prompt is the contract": Surrendering systemic integrity to a text block in the hope that the LLM will follow formatting guidelines.

Every single one of these approaches fails under the weight of day-to-day software iteration. Prompt conventions leak, hardcoded rules drift, and "the prompt as the contract" inevitably breaks when a downstream system undergoes a minor API modification.

The architecture we settled on relies on a single, declarative YAML grammar, loaded by a highly performant C# facade, and consumed simultaneously by AI prompt generators, runtime validators, and client-side design engines.

The grammar is the single source of truth (SSOT). Everything else cites it.


The Hard Cost of Code Drift

Before adopting a unified grammar approach, our platform handled integration rules the traditional way. Rules governing system connections, entity roles, operation flags, and array boundaries were distributed across multiple layers of the application.

The system composition rules were duplicated across at least four distinct files:

  1. The system prompt block instructing the LLM how to build a flow.
  2. The UI components governing manual drag-and-drop actions in the wizard.
  3. The runtime validation service checking the payload before execution.
  4. The database constraints mapping the structural properties.

This layout introduced a recurring, high-friction bug class: architectural drift. A developer would update an integration requirement to allow a specific system role to execute an automated write-back operation. They would update the validation code, but forget to modify the AI prompt template. On the next user execution, the AI wizard would continue suggesting the old, restricted workflow configuration, leading to instant execution rejections at the validation boundary.

The system was constantly arguing with its own features. The metadata was true in the validator, suspicious in the UI, and an outright lie to the AI.


The Architecture: One Grammar Family, One Facade Pattern

To eliminate drift, we consolidated our entire integration configuration space into a unified, 430-line composition grammar. This file details exactly how components are allowed to plug together, structured across 13 distinct semantic sections.

The composition grammar isn't a freestanding document; it's a dialect that declares and inherits from a mode-wide base task grammar. The base makes the contract mode-aware — it requires every task-execution dialect to declare an artifact schema, a validator target, and a preview/apply ceremony, and it explicitly forbids a task dialect from carrying problem-solving sections (moves, correlations, narration_hints). The mirror structure exists on the diagnostic side. The grammar isn't just a single source of truth about rules; it's a single source of truth about which rules belong to which mode.

To make this configuration accessible to the runtime without parsing bottlenecks, we abstract the document behind an optimized C# facade, GrammarRules.

// A thread-safe, cached facade projecting our declarative grammar
public static class GrammarRules
{
    private static readonly Lazy<Dictionary<string, List<string>>> _intentRolesCache =
        new Lazy<Dictionary<string, List<string>>>(LoadIntentRoles);

    private static Dictionary<string, List<string>> LoadIntentRoles()
    {
        var yamlDoc = ResourceLoader.ReadEmbeddedYaml(CompositionGrammarResource);
        return YamlParser.ProjectSection(yamlDoc, "intent_roles");
    }

    public static IReadOnlyCollection<string> GetRequiredRoles(string integrationIntent)
    {
        if (_intentRolesCache.Value.TryGetValue(integrationIntent, out var roles))
            return roles;
        return FallbackSafetyNet.GetDefaultRolesFor(integrationIntent);
    }
}

This architecture completely decouples the system logic from the rule specification. Components never read the raw YAML file directly, and they never re-implement the rules in code. They query the GrammarRules facade.


A Worked Example: The Three-Way Rule Lifecycle

When a rule lives in a single file, adding a system capability requires changing only that file. Imagine adding a new IntegrationIntent called ReplicateChangeOrder. We update the composition grammar:

# composition grammar — intent_roles section
intent_roles:
  ReplicateChangeOrder:
    - PrimarySpine
    - EnrichmentLookup
    - WriteBackResponder

Once this block is written, the change flows immediately into three runtime environments:

1. The AI Prompt Composition

When a user launches the creation assistant, our AI engine calls GrammarRules.GetRequiredRoles("ReplicateChangeOrder"). The prompt builder maps those exact structural limits directly into the system message. The AI does not guess which roles are valid; it is handed the active vocabulary list derived from the code build itself.

2. The Deterministic Validator

When the AI outputs a completed structural blueprint, the platform channels it into the GrammarConformanceValidator. This component loops over the proposed execution steps, passing them through the GrammarRules facade. If the model attempts to invent a disallowed combination, the validator instantly catches the mismatch.

3. The Test Suite Enforcement

Every pull request runs an automated test suite that enforces strict structural discipline:

// Grammar facade parity test
[Fact]
public void Facade_Projections_Must_Match_Grammar_Declarations()
{
    var rawYamlIntents = RawYamlReader.ExtractKeys("intent_roles");
    var facadeIntents = GrammarRules.GetAllRegisteredIntents();

    // If a developer adds an intent to the YAML but forgets to expose it
    // in the C# facade, the build fails immediately.
    Assert.Equal(rawYamlIntents, facadeIntents);
}

Breaking Language Barriers: The Cross-Runtime Contract

The ultimate validation of this SSOT approach occurs when the contract must span completely different software ecosystems.

In our platform, field-level data manipulation is powered by a Python sandbox, while our orchestration engines and AI mapping features live in C#. To lock this boundary down, we run a cross-language grammar layer: a shared mapper vocabulary.

# mapper vocabulary — the cross-language contract
categories:
  numeric: [abs, round, min, max, int, float]
  text: [str, strip, lower, upper, replace, split, join]
  iteration: [len, sorted, enumerate, zip]
forbidden_globals:
  - __import__
  - eval
  - open

This single file serves as the definitive reference point for four distinct execution boundaries:

  1. Python Runtime Control: The sandbox reads the YAML to build its environment whitelist. If a function isn't in the contract, the sandbox refuses to execute it.
  2. C# Prompt Injection: The prompt builder reads the same YAML to compose the rules shown to the LLM during mapping setup.
  3. C# Pre-flight Validation: Before code is sent over the wire to the Python worker pods, the C# facade scans the accessor chain to catch typos and illegal names.
  4. Python CI/CD Parity Tests: On every commit, a Python test evaluates the active runtime environment against the YAML contract.

If a developer drops an unauthorized function into the Python code without declaring it in the global contract file, the build halts immediately:

Contract declares names the runtime does not expose: ['bogus_drift_canary']. Either add them to get_safe_environment() or remove them from the vocabulary contract.


The Pattern Generalizes: A Family of Grammars, Not a One-Off

The composition grammar was the seed, not the ceiling. Our diagnostic platform already emits a DiagnosticBundle whose findings[] array ships pre-classified: every anomaly arrives with a stable code (DB001 … CL007), a severity, a category, a trigger condition, and an evidence shape.

Read that description again — a table of stable codes, severities, categories, triggers, and summaries is structurally identical to a grammar section. For a long time it didn't live in one: the rule table sat inside the diagnostic bundle service, reproduced informally in the operator UI's labels and again in the AI prompt's explanation of each code. Three surfaces, one rule table, the same drift risk we had already paid down once.

The fix was the same move applied to a second mode: migrate the table into a health_finding_rules section of a dedicated health grammar, and let the runtime rule firing, the operator UI, and the AI's narrative prompt all generate from that section. That has since shipped. The platform now runs a family of grammars under one discipline — a composition grammar that bounds what the AI may build, and a health grammar that bounds what the AI may claim — each the typed contract for its own mode, each enforced by the same edit-YAML → expand-facade → verify-by-test → consume-in-production lifecycle. The health grammar projects its own analyzer block into the bundle-analysis prompt through the same facade discipline the composition grammar established.


The Honest Limit: SSOT Does Not Make a Composed Prompt Consistent

Everything above is a claim about drift between the grammar and the code that cites it, and that claim holds up. It is worth being equally precise about what it does not cover — the boundary turns out to be more interesting than the pattern.

A single source of truth is a property of a single document. Our prompts are not built from a single document; they are composed. The planning prompt builder emits the base task grammar's block and then appends the composition dialect's rules on top of it. The bundle-analysis prompt builder does the same with the health dialect. What actually reaches the model is a concatenation — and the concatenation is not itself a grammar document that anything validates.

The failure this permits has a specific shape. The base layer declares the properties an output must have; the dialect layer shows an example of how that output should be rendered. Each layer is internally coherent. Each passes its own conformance test. Read together — which is the only way the model ever reads them — they can be asking for two different things, and nothing in the pipeline notices. Our grammar-mode conformance tests validate each YAML document in isolation. No test asserted that base plus dialect, concatenated, request one artifact rather than two.

What makes this class of defect genuinely hard to catch is that it does not present as a failure. A capable model silently resolves the ambiguity, usually by picking the reading you intended, and the contradiction sits in production looking exactly like working software. It surfaces only when something shifts the model's disposition toward ambiguity — a version upgrade, a different provider, a change in decoding behavior. Then the prompt starts being read more literally, the latent contradiction resolves the other way, and what looks like a regression in the new model is a defect that was there the whole time.

The lesson does not retire the pattern; it sharpens it. SSOT is a property of each document, not of the composition. If you assemble prompts from layered grammars, the assembly needs a conformance test of its own — something that reads the composed block and asserts that its layers agree on the artifact they are asking for. The single-source guarantee stops precisely at the boundary where two sources begin being concatenated, and that boundary is invisible until a model finds it for you.


Knowing Where to Limit the Scope

A single source of truth is highly effective, but it is not free. Every developer on the codebase must accept the rigidity of the model: you are no longer allowed to "patch in a quick validation condition" directly in code. You must model the change through the grammar layer first.

Avoid deploying formal grammars when:

  • Micro-surfaces with static rules: If a component will only ever use a tiny, completely unchanging set of parameters, a formal YAML facade is overkill.
  • Purely algorithmic execution paths: When execution rules are based entirely on dynamic, state-dependent procedural code, forcing those behaviors into a static layout produces unreadable abstractions.
  • Use cases that haven't shipped yet: When you spot a future capability that will obviously need a grammar section eventually, the temptation is to lay the groundwork now. Don't. A section nothing reads isn't a contract; it's a guess wearing structure's clothing. Each new section should land in the same release as the use case that reads it — never before.

The Value of Curation

The true strength of an AI-assisted integration engine does not stem from its ability to handle unconstrained processing flows. It stems from its ability to enforce domain boundaries relentlessly.

A declarative grammar functions as an architectural anchor. It encapsulates years of production engineering experience into an explicit, machine-readable validation contract. By treating this document as our absolute source of truth, we ensure that our application logic, our platform testing pipelines, and our generative AI components remain permanently aligned. The grammar makes safety a permanent trait of the architecture itself.


Back to Insights