Put Deterministic Systems Before Generative Explanations

Choose the right operation
Start by classifying each task as computation, retrieval, or generation. The category determines which component should own the answer.
| Operation | Use it when | Typical examples |
|---|---|---|
| Compute | Published rules and inputs should determine one reproducible result | Totals, date conversions, eligibility rules, risk thresholds |
| Retrieve | An authoritative source already contains the current fact | Account status, inventory, policy text, approved product data |
| Generate | The task allows several useful ways to express the same grounded facts | Explanations, summaries, examples, tone changes |
This is a design rule, not a ban on AI. A model may help a user understand a tax calculation, but it should not make up the tax rate. It may explain why an appointment is unavailable, but the schedule service should decide whether the slot exists. It may summarize a policy, but the policy store must supply the current text and effective date.
Mixed tasks should be split. For example, a support request may require retrieving the customer’s plan, computing whether a usage limit has been reached, and generating a clear explanation. Keeping those operations separate makes each one testable.
Define a contract between the systems
The generative layer should receive a small, explicit data contract instead of unrestricted access to application state. The contract should contain the result, the inputs that matter, the rule or data version, warnings, and an identifier that lets the team trace the request.
{
“result”: {“eligible”: false, “reason_code”: “LIMIT_REACHED”},
“facts”: {“current_usage”: 103, “plan_limit”: 100},
“source_version”: “plans-2026-09”,
“warnings”: [],
“trace_id”: “req_7f31”
}
The prompt can tell the model to explain only those fields and to avoid inventing missing details. The application can also reject any generated response that changes a protected value. If the model says the plan limit is 120 while the contract says 100, the system should fail the response rather than hope a user catches the mismatch.
Treat retrieved text as untrusted
Retrieval does not automatically make an answer reliable. A document may be outdated, incomplete, or written by someone who was never authorized to set the rule. It may also contain text that looks like an instruction to the model. The retrieval service should therefore return provenance with the content: source owner, publication or effective date, region, version, and access decision. The application should decide which sources are authoritative before the model sees them.
Keep retrieved content in a clearly marked data field. Do not mix it into the same instruction channel that defines the model’s job. If a support note says to ignore policy and approve a refund, the system should quote or summarize the note as evidence, not obey it as a command. This boundary reduces prompt-injection risk and makes stale-source errors easier to diagnose.
Preserve a calculation receipt
A deterministic result is useful only when the team can reproduce it later. Store a compact receipt containing normalized inputs, calculation or source version, timestamp, outputs, and warnings. Avoid storing sensitive raw inputs when a reference or redacted form is enough.
The receipt serves three purposes. Support can reconstruct what the user saw. Engineers can compare results before and after a rule change. Auditors can distinguish a calculation defect from a wording defect. Without that boundary, teams often search model logs for an error that began in a time-zone converter, rules engine, or stale database record.
Validate before generating
Validation belongs on both sides of the deterministic boundary. Validate incoming data before computing, and validate the result before it reaches the model. A date parser should reject an impossible date. A currency calculation should declare its rounding mode. A retrieval step should reject expired records when freshness matters.
Use fixtures for difficult but known cases. Good fixture sets include boundary dates, daylight-saving transitions, rounding edges, missing fields, conflicting source records, and version changes. Each fixture should assert the structured result rather than the prose that happens to describe it.
The same discipline applies to the generative layer. The UK government’s AI Insights guidance notes that these systems can produce plausible but inaccurate results and recommends testing outputs against ground truth or expert judgment. For this architecture, the structured result is part of that ground truth.
Make failure visible
A failed calculation should not become an invitation for the model to improvise. Define failure states the interface can show directly: missing input, unsupported region, stale source, conflicting records, calculation timeout, or version mismatch. The explanation layer can translate a known error into plain language, but it should not replace the missing result.
This is especially important when partial data looks complete. Suppose a scheduling service returns a date but omits the time zone. The model may still write a polished confirmation. A stricter pipeline treats the absent time zone as a blocked state and asks the user to resolve it.
NIST’s Generative AI Profile identifies confabulation as a distinct risk and places testing, evaluation, verification, and validation across the AI lifecycle. Visible failure states put that principle into the product itself: the system admits when it lacks a valid result instead of hiding uncertainty behind fluent text.
Test the boundary instead of the prompt alone
Prompt tests are necessary, but they cover only one component. Add tests at the handoff between deterministic and generative systems.
- Contract tests verify that required fields, types, units, and version identifiers are present.
- Mutation tests change one protected value and confirm that the explanation changes only where expected.
- Adversarial tests place misleading instructions in retrieved text and confirm that the system treats the text as data, not authority.
- Regression tests replay saved receipts after changes to calculation code, retrieval sources, prompts, or models.
- Presentation tests confirm that the interface labels computed facts, retrieved facts, generated explanation, and warnings clearly.
Measure the components separately. Track calculation errors, retrieval freshness, contract failures, generation refusals, protected-value mismatches, and user corrections. A single satisfaction score cannot tell the team which layer failed.
Use a staged rollout
Teams can introduce this pattern without rebuilding the whole product. Start with one high-consequence answer that currently depends on a model. Identify the facts that must be reproducible, move them into a small service or validated retrieval step, and pass a structured result to the existing prompt.
Run the old and new paths in shadow mode. Compare structured results first, then compare explanations. When the deterministic layer is stable, expose the source version and warnings to support staff. Finally, show users only the provenance that helps them understand or correct the result.
A practical release gate is simple:
- The same valid inputs produce the same protected result under the same version.
- Every result can be traced to a receipt and source version.
- Missing or conflicting data produces an explicit failure state.
- The model cannot alter protected values without detection.
- The team can test the explanation independently from the calculation or retrieval step.
Keep the model in the role it handles well
Generative models are valuable because they can adapt an explanation to a user’s question, language, and level of knowledge. That flexibility is strongest when the model receives facts it does not need to invent.
The architecture also makes product decisions easier. Teams can update a prompt without changing business rules. They can replace a model without rewriting calculation code. They can correct a source record and regenerate the explanation from the same receipt. Most importantly, they can tell whether a problem came from the facts or the words.
The rule is straightforward: compute what must be exact, retrieve what must be current, and generate what can be expressed in more than one useful way. Put those boundaries into the system before polishing the prompt.
Related posts

How AI Is Transforming Digital Workflows in 2026
Artificial intelligence is rapidly becoming part of the infrastructure behind modern digital work. Its impact is visible in software development, research, publishing, marketing, analytics, automation.
Read more
How to compare two AI tools for one job
To compare two AI tools honestly, you must evaluate them against one specific operational job. When an option bills against an API meter, you must.
Read more
How to claim and optimize your AI Dust listing
Claiming takes two minutes. Optimizing it properly takes a bit longer — here's where to start.
Read more