The strongest argument for deterministic execution deserves to lead: if a system’s behavior can be reproduced under specified conditions, reviewers have a clearer target to inspect. That is a serious engineering objective. I want more systems to state and test that promise.
Then I want the next question asked: what makes the reproduced result acceptable?
Why take repeatability seriously?
Thinking Machines Lab explains how changes in batching can alter floating-point reduction order, and why batch-invariant operations require a stable order for an individual element’s reduction. The technical point is concrete: the serving conditions belong in a repeatability claim. Thinking Machines Lab.
This is a better argument than treating a prompt setting as a complete system contract. It identifies a source of variation and a computation property intended to control it. The accessible technical passages support that mechanism; our article is not an independent reproduction or a universal performance claim.
What is the strongest objection to asking for more?
A builder might fairly respond that reproducibility is the property being offered. Why criticize a component for not proving a different property?
That objection is right. An inference component should be judged against its stated contract. My argument concerns the assurance case for the whole workflow: the buyer still needs evidence for the decisions and actions built around that component. Demanding a complete system account does not diminish a useful component guarantee.
What makes a repeated decision right?
Take a hypothetical reimbursement process. It can apply the same rule to the same supplied amount every time. To approve the outcome, a reviewer must also establish that the amount came from the right record and that the applicable rule was selected. Repetition alone does not answer those separate questions; this follows from the example’s logic, not from a measured product failure.
AWS’s Automated Reasoning documentation offers a concrete complementary property: checking responses against user-defined policies and explaining results using policy rules and variable assignments. That is evidence about consistency with the supplied policy. Its documented limitations also keep input integrity in view. AWS documentation.
What contract should the industry offer?
I would ask for three statements, each paired with its own evidence:
- What repeats? Specify the result, execution boundary and fixed conditions.
- What is checked? Identify authoritative inputs, applicable rules and acceptance criteria.
- What happens when a check fails? Show the stop, escalation and authorized resolution.
This is my proposed assurance structure. It supports reproducible execution, explicit business rules and human control of exceptions without pretending one property establishes all the others.
What would settle the debate in a review?
A demonstration that names these boundaries and leaves inspectable records would be more persuasive than an argument about labels. Keep the repeatability test. Add the correctness reference. Retain the exception record.
The definition of deterministic AI establishes the scope, while the Checklist and audit evidence packet turn it into questions a reviewer can answer.
The source record
Read the original evidence and the scope of our review.
- Defeating Nondeterminism in LLM InferenceThinking Machines Lab · 2025-09-10 · Accessed 2026-10-11Indexed technical passages on batch invariance and fixed reduction order read; full article unavailable.
- What are Automated Reasoning checks in Amazon Bedrock Guardrails?Amazon Web Services · Not stated · Accessed 2026-10-11Substantive indexed capability sections and opening limitations read; full page not retrieved.
