Is an LLM deterministic?

An identical prompt or a temperature-zero setting does not establish that an LLM deployment is deterministic. Evaluate repeatability within a specified model, configuration and execution environment.

Updated October 11, 2026

Why is the prompt only part of the input?

Thinking Machines Lab describes practical LLM inference that can vary even when sampling is theoretically deterministic at temperature zero. Its introduction separates the sampling setting from the computation used to produce the result. Read the introduction.

PyTorch’s reproducibility guidance makes another boundary explicit: identical seeds do not guarantee agreement across CPU and GPU, releases or platforms. These are limits of the stated guarantee, not a claim that every pair of runs must differ. Read PyTorch’s guidance.

What should you record?

Our recommended evaluation record includes the complete request and conversation, system instructions, model identifier, generation settings, supplied evidence, software versions and the execution environment you control. For a hosted service, identify conditions the provider does not expose; keep those as limitations of your test.

For an agent, also record tool results and initial workflow state. Decide before testing whether the target is identical text, structured values, decisions or actions. Do not silently replace an exact-match requirement with “similar meaning.”

What does temperature zero establish?

It describes a sampling configuration; it is not, by itself, evidence of end-to-end reproducibility. The Thinking Machines introduction explicitly distinguishes that theoretical sampling behavior from practical inference. Source.

Do not use a seed or temperature setting as a substitute for a scoped deployment claim. PyTorch recommends controls within a specific platform, device and release and notes that deterministic operations may be slower. That trade-off should be evaluated on the intended workload. Source.

How should you test a deployment?

Run fresh executions from the same declared starting conditions. Compare the chosen result exactly; store differences and configuration records. Keep tests after model or environment changes in a separate group because they examine a broader claim.

As a suggested exercise, repeat a small set of representative prompts several times and inspect the raw results. Choose the number of runs according to the decision you need to support; this guide sets no certification threshold.

What should the report conclude?

Report “all observed outputs matched under these conditions” when that is what the experiment shows. A finite test does not examine every possible input. If outputs differ, inspect whether the declared inputs or hidden conditions changed before attributing the mechanism.

See how to test determinism for a complete record format, and what deterministic AI means for the distinction between inference and business execution.

Reading scope

Sources

  1. Defeating Nondeterminism in LLM Inference, Thinking Machines Lab (2025-09-10)
  2. Reproducibility, PyTorch (2026-05-14)