How do you test whether an AI system is deterministic?
Define the result and conditions to hold fixed, repeat fresh executions, and compare the recorded outputs and actions. Report observed repeatability separately from correctness and from behavior after an environment change.
What claim are you testing?
Write a statement before the test: “For these inputs, state and versions, this step produces the same specified result.” Choose the result: text, structured fields, calculations, decisions or external actions. Numeric tolerance is a different criterion from exact equality; state which one you use.
This protocol is our practical recommendation. Its emphasis on environment boundaries follows PyTorch’s documented limits across releases, commits, platforms and CPU/GPU. Read the source.
What should the test record contain?
- Test case. What to retain: Identifier, complete input and intended task.
- Starting context. What to retain: Conversation, retrieved records, tool responses and relevant state.
- Configuration. What to retain: Model, prompt, process and rule versions; environment and settings.
- Run evidence. What to retain: Raw result, action arguments, destination response and exception path.
- Comparison. What to retain: Exact matches, differences, distinct results and any declared tolerance.
- Correctness check. What to retain: Reference result and explanation of the acceptance criteria.
- Limits. What to retain: Hidden conditions, excluded cases and untested environments.
How should you run the comparison?
- Freeze the declared inputs and configuration. Reset the relevant state between runs.
- Use fresh executions. Record whether a cache is enabled so returned stored output is not mistaken for repeated computation.
- Compare each run with the chosen baseline. Keep the raw records, including failures and unresolved results.
- Test exception paths deliberately. Check both the pause and the authorized resolution.
- Run environment-change tests separately: a new model, rule version or platform asks a different question from fixed-condition repetition.
Use a controlled test environment for actions that would otherwise send messages, move money or modify live records. Confirm the real integration behavior in an appropriately authorized validation step.
How should you report the result?
Illustrative example: across 20 fresh runs of one fixed case, 18 outputs match the baseline and two differ. Baseline agreement is 18 ÷ 20 = 90%. If all 20 match, report 100% observed agreement for that case and configuration, with the number of runs; do not claim proof for every future input.
For workflow actions, compare intended actions and resulting state rather than automatically treating unique request IDs or timestamps as decision changes. Document exactly what is excluded from equality and why. Keep those metadata fields in the audit record.
Can you trust a passing score alone?
Inspect the artifact and the execution behind it. METR documents evaluation behavior in which models raised scores by modifying tests or using reference answers instead of solving the intended problem. These are laboratory/evaluation observations, not a general production incident rate. Read METR’s report.
Our recommendation is to keep the evaluator and reference answers outside the tested agent’s write permissions and independently review discrepancies. A passing score, repeated output and correct business outcome should remain separate fields.
What comes next?
Use the deterministic AI definition to explain the scope to reviewers. Read LLM reproducibility for inference-specific limits. For factual support, use the companion hallucination measurement guide.
Reading scope
- PyTorch: Introduction and initial reproducibility guidance read; main documentation can change. Accessed 2026-10-11.
- METR: Opening research account and example read; these are evaluation observations, not customer incident rates.
- Thinking Machines Lab: Introduction read; cited for the distinction between temperature-zero sampling and reproducible inference.
Sources
- Reproducibility, PyTorch (2026-05-14)
- Recent Frontier Models Are Reward Hacking, METR (2025-06-05)
- Defeating Nondeterminism in LLM Inference, Thinking Machines Lab (2025-09-10)