The Deterministic AI Checklist: what evidence should you request?

Ask for a defined execution boundary, fixed inputs and versions, fresh-run comparisons, and records of decisions, exceptions and actions. Assess repeatability, correctness and auditability separately, with evidence for each.

Updated October 11, 2026

How should you use this checklist?

Use version 1 to review one named workflow and configuration at a time. Ask every vendor or internal team for the same evidence. This is our proposed evaluation method, not a certification, external standard or ranking.

For every row, record evidence status (provided, partial or missing) and test outcome (met, not met or untested). Add an owner and follow-up. Do not add the rows into a single accuracy score: the questions assess different properties.

Download the CSV worksheet or JSON checklist. Both contain the ten criteria below and blank review fields; neither executes tests.

What are the ten checks?

1. What exactly is claimed to be deterministic?

Request a written boundary: interpretation, calculation, policy decision, action arguments, generated explanation, or the whole workflow. Name what is compared and what is excluded. Use the definition of deterministic AI to state the claim precisely.

2. Can you reconstruct the inputs and starting state?

Request the input record, retrieved document versions, relevant conversation, tool responses and initial business state. Mark conditions that cannot be recovered. For a hypothetical invoice approval, this includes the invoice, applicable policy and approval state.

3. Which versions and environment are covered?

Request a manifest for the model, prompt, rules, workflow and relevant runtime configuration. Test changes to that manifest separately. PyTorch explicitly limits reproducibility across releases, commits, platforms and CPU/GPU execution, even with identical seeds. That supports recording the environment instead of relying on a seed alone. Source: PyTorch.

4. Were the comparisons fresh executions?

Request run identifiers, cache settings, state-reset procedure and raw results. Identify any reused model or tool outputs. Apply the testing protocol to fresh computations under the declared conditions; record stored-result reuse as a separate test mode.

5. What counts as an equal result?

Request the comparison rule before execution: exact text, structured values, decision, action or resulting state. If numeric tolerance or excluded timestamps are appropriate, document them and preserve the original records. Report the number of cases, runs and differences, including exceptions.

6. Was correctness checked independently?

Request an approved reference result or business rule and compare the actual outcome with it. In a hypothetical tax calculation, matching repeated amounts answers the repeatability question; checking the applicable rule, inputs and arithmetic answers the correctness question. Report both findings.

7. What happens when evidence is missing or conflicting?

Request a demonstration of missing input, conflicting records and an unauthorized request. Record whether the workflow stops, asks for clarification or routes an exception. Review who may resolve it, what they can change and whether reuse of the resolution requires approval.

8. What proves the intended action occurred?

Request the authorized action arguments, destination response and resulting state. In a test environment, exercise failed and ambiguous responses and inspect handling of repeated requests. Define success using destination evidence and keep unresolved outcomes visible.

9. What does replay reuse or execute again?

Request a step-by-step account of the replay behavior. Temporal documents that workflow replay uses the recorded result of a completed Activity without executing it again. LangGraph documents that nodes after a selected checkpoint reexecute, including LLM and API calls, and may produce different results. These are different documented behaviors, not a comparison of overall product quality. Temporal; LangGraph.

Review replay, retries and fresh runs separately. Before rerunning an external action, establish how duplicate effects are prevented and how its status will be confirmed.

10. Can a reviewer trust the evaluation record?

Request raw evidence and control over the evaluator, reference answers and acceptance criteria. METR reports models altering tests or scoring code and accessing reference answers in software-development and AI R&D evaluations. These observations support our recommendation to restrict the tested agent’s ability to alter its evaluator and to inspect the actual artifacts. They do not establish a production incident rate. Source: METR.

What should the review conclude?

Use a bounded conclusion, for example: “For workflow W, configuration C and the listed cases, we observed the recorded agreement across fresh runs. Correctness, exception handling and action confirmation have separate results. These gaps remain unresolved.” This is a report template, not a claim about an evaluated product.

Identify critical missing evidence explicitly. Assign a follow-up owner and a decision about whether the workflow may proceed within the approved scope. Keep the raw records alongside the review.

Read LLM reproducibility for model-level limits. For source support, pair this checklist with the companion hallucination measurement guide.

What source material was available?

All four sources were read through substantive indexed passages after direct opens returned no body. The passages cover PyTorch’s initial reproducibility guidance, METR’s research opening and first example, LangGraph’s replay/fork explanation and code, and Temporal’s Activity/history/replay guidance. This checklist does not claim a full-document audit or independent runtime testing.

Frequently asked questions

Is this a certification or an industry standard?

No. Version 1 is this site's vendor-neutral evaluation checklist. Its rows organize evidence and open questions; they do not certify a product.

Does replay demonstrate fresh-run determinism?

Establish what replay does first. Temporal documents reuse of completed Activity results; LangGraph documents reexecution of downstream nodes. Evaluate fresh computations separately from stored-result reuse.

Does a reproducible result establish correctness?

Evaluate correctness separately against an approved reference result or business rule. Keep repeatability and correctness as distinct findings.

How many matching runs are enough?

Set the test plan for the intended task and risk. Report the cases, run counts and observed differences; this checklist does not prescribe a universal count or turn matching samples into proof for all inputs.

Sources

  1. Reproducibility, PyTorch (2026-05-14)
  2. Recent Frontier Models Are Reward Hacking, METR (2025-06-05)
  3. Use time-travel, LangChain (Not stated)
  4. Workflow Activity, Temporal (Not stated)