A non-deterministic model behind a deterministic gate
We assume the model is fallible and built for it. Nothing an AI produces reaches the runtime without passing a formal grammar, and what ships is Python you read, test, and commit to source control — not an opaque artifact you have to trust. The same bound applies when the AI is diagnosing a problem rather than authoring one: it works from a closed catalog of findings, and cannot invent a diagnosis.
Two claims that get confused
What we do not claim
That the AI's output is deterministic. It is not. Ask the same question twice and you may get two different transforms. Any vendor telling you otherwise is describing something you can disprove in two minutes.
What we do claim
That the gate is deterministic. The grammar engine accepts or rejects a given artifact the same way every time, and nothing reaches the runtime without passing it. The model proposes; the grammar disposes.
That distinction is the whole design. It is also the stronger claim, because it says we treated the model as an unreliable component and put an engineered control in front of it — which is what you would do with any unreliable component in a regulated system.
What the AI actually hands you
A field mapping in Samba is a Python expression. When AI drafts one, the artifact you review is the same artifact a human would have written — readable, diffable, and coverable by a unit test.
A generated transform
# Target field: member.coverage.effective_date
# Source: eligibility_response.benefit[0].period.start
#
# Drafted by AI, validated by the grammar engine, reviewed by a human,
# and committed to source control like any other change.
parse_date(src['benefit'][0]['period']['start'], fmt='%Y%m%d').isoformat()
There is no proprietary DSL between you and that expression, and no model artifact you cannot open. If your validation process requires a human to review every transform before it reaches production, this is a review your people can actually perform.
And when the grammar rejects one
The interesting case is the failure. The grammar encodes what is legal in the domain, not merely what is well-formed — so a syntactically perfect expression that violates a domain rule is rejected with a specific reason, before it can be saved.
{
"accepted": false,
"rule": "transform.no_unbounded_external_call",
"detail": "Expression calls requests.get(); transforms may not perform
network I/O. Fetch the value in a prior endpoint step and
reference its output instead.",
"expression": "requests.get(lookup_url).json()['code']",
"correlation_id": "7f3a1c94-2b8e-4d21-9c05-1a6e8f0b2d77"
}
No silent fallback, no substituted default, no best-effort guess. When something is missing or mismatched, Samba fails explicitly and logs the failure.
How we know it still works
The grammar engine tells you that one output is admissible. It does not, on its own, tell you that behaviour is repeatable under change — and in a validated environment that is the harder question, the one that actually blocks AI adoption: how do you know it still behaves the same after a change?
Our answer is a versioned scenario corpus with recorded fixtures, replayed deterministically offline as a build gate. Every release runs it. A change that alters behaviour fails the build rather than reaching a customer.
Versioned scenarios
Each scenario is a real integration shape with known-correct behaviour, versioned alongside the code it exercises. The corpus grows every time a defect teaches us something the existing scenarios did not cover.
Recorded fixtures
Model responses and upstream payloads are recorded once and replayed, so a test run does not depend on a live model or a live endpoint. The same inputs produce the same comparison every time.
Deterministic offline replay
Replay happens with no network and no inference call. That is what makes the gate deterministic even though the thing it is testing is not.
Run as a build gate
The corpus runs on every build, not on a schedule and not by hand. A behavioural regression fails the build, which is the only enforcement mechanism that actually holds over time.
We publish the method rather than the corpus itself — the scenarios encode domain rules and are closer to product than to content. Customers and assessors in an active evaluation can review it directly with us.
The same discipline, pointed at diagnosis
Authoring is only half of what the AI does. The other half runs when something breaks — reading a failed execution, or a snapshot of the whole deployment, and telling you what it sees. That is a far more dangerous place to put a language model, because a confident wrong answer during an incident is worse than no answer at all.
So it is bounded the same way the authoring path is bounded — structurally, not by asking it nicely. The AI cannot invent a diagnosis.
A closed finding catalog
Findings come from a versioned catalog with stable codes. The AI selects from that set; it cannot mint a new one. A finding you read today means the same thing six months from now, which is what makes findings correlatable across incidents at all.
The runtime proves what it can
Where a fact is cheap to establish — an HTTP status, an exception type — the runtime emits the code itself and the AI only narrates it. The model is given latitude exactly where the evidence is genuinely ambiguous, and nowhere else.
Every claim carries its certainty
Confirmed, likely, possible, or ruled out. A hypothesis is labelled a hypothesis. You are never handed a guess wearing the costume of a fact, which is the failure mode that makes AI diagnosis untrustworthy in the first place.
Findings cite their evidence
Each one points at the specific anchor in the bundle that supports it. Expand any finding and read the raw evidence yourself — the narration is a convenience, not something you are asked to take on faith.
What it looks at
For a single failed run: which pipeline stage failed and why, whether the failure is transient or permanent, and — where it is identifiable — the specific field mapping expression or HTTP request that caused it.
For the deployment as a whole: a diagnostic bundle collected in seconds covering pod and runtime state, database, Vault, scheduling, recent execution history, and a live view of the Kubernetes cluster itself. Bundles carry schema metadata, counts, timestamps, and status flags — never payload bodies, secret values, or customer data. They are safe to attach to a support ticket, which is precisely why they are shaped that way.
Where a finding has a known remedy, the recovery action is a signed, scoped operation with a defined blast radius, a rollback path, an admin confirmation step, and an audit-log entry — never a free-form command. The console offers only the actions the system's current state actually warrants, and tells you what each one will not do before you confirm it.
Where the inference happens
AI authoring runs against your Azure OpenAI endpoint, under your key, inside your tenant. Cirrus Tempo does not host inference on your behalf. The model, endpoint, and request shape are configured per deployment and can change without us shipping a release.
Being precise about the limits: one inference provider is implemented today, and it is Azure OpenAI. The provider sits behind a single-method interface that nothing above it depends on, so a second one is contained work — but it is code, not a configuration toggle, and we will not describe it as one. Model-portable by design; Azure OpenAI today.
The platform also runs entirely without AI. Every AI surface is additive and degrades cleanly, so an operator who declines external inference still has a fully functional integration platform. If a disconnected deployment is a hard requirement for you, say so early and we will tell you plainly where that sits on the roadmap rather than implying it ships today.