What is deterministic AI?

A deterministic AI system produces the same specified result for the same input and relevant context within a defined execution boundary. State which parts are deterministic and which conditions must remain fixed.

Updated October 11, 2026

What does the definition cover?

This is the working definition used by this reference. For a business process, specify whether “result” means an extracted field, calculation, decision, sequence of actions or final response. Also specify the state and versions included in “same input.”

Environment matters in real computation. PyTorch warns that complete reproducibility is not guaranteed across releases, individual commits or platforms, and that CPU and GPU results may differ even with identical seeds. Read PyTorch’s reproducibility guidance.

Which parts of an AI system can you evaluate separately?

Use these boundaries to make an evaluation claim precise:

This table is our recommended evaluation structure. Do not infer that a whole workflow is deterministic merely because one calculation is.

Is temperature zero enough?

No. Thinking Machines Lab’s introduction distinguishes theoretically deterministic sampling at temperature zero from inference systems whose results can still vary in practice. PyTorch likewise documents limits to reproducibility across execution environments. Thinking Machines Lab · PyTorch.

Use the LLM determinism guide to define the conditions for your deployment.

How do correctness and auditability differ?

Under this site’s terminology, determinism concerns repeated execution, correctness concerns whether a result satisfies the task’s requirements, and auditability concerns evidence that lets someone reconstruct and examine what happened.

A hypothetical rule that always adds the wrong tax percentage is repeatable but wrong. A log that merely says “done” does not show which rule or input produced the outcome. These examples illustrate the distinctions; they are not claims about a named system.

Ask for separate evidence: repeated results, checked reference answers, and execution records. METR’s evaluation research reports models exploiting scoring systems instead of performing the intended task, reinforcing the need to inspect the work behind a success score. Read METR’s research.

How can AI interpretation fit with explicit execution?

In a proposed invoice process, AI extracts candidate fields; checks establish which fields the workflow accepts; explicit rules determine routing and calculations; a person resolves exceptions. To assess that design, ask where interpretation ends, what is approved for execution and what changes after human intervention.

The architecture needs both a defined boundary and evidence from the intended workload. Our companion reference explains hallucination prevention controls.

What should you request from a supplier?

Request a written scope, fixed input and version records, repeated uncached runs, expected results, exception examples and evidence of external actions. Use our testing guide and the short determinism definition.

Reading scope

Frequently asked questions

Does deterministic mean accurate?

No. In this reference, determinism means the same specified result under the same declared conditions. Accuracy requires a separate check against the task’s requirements.

Can an AI workflow have a deterministic execution layer?

Evaluate that claim at the stated boundary: inspect the approved rules, fixed inputs and context, repeated results and exception paths. Also test how variable interpretation can affect what enters the layer.

Sources

  1. Reproducibility, PyTorch (2026-05-14)
  2. Defeating Nondeterminism in LLM Inference, Thinking Machines Lab (2025-09-10)
  3. Recent Frontier Models Are Reward Hacking, METR (2025-06-05)