Treat this as an inference design to inspect: a fast candidate computation followed by a verification step. The engineering question is which result the verifier fixes and under which conditions.
What is the development?
Microsoft Research lists LLM-42: Enabling Determinism in LLM Inference with Verified Speculation under SOSP 2026, September 2026. The preprint’s listed revision is January 30, 2026; this is conference research coverage, not a claim that the technique first appeared this month. Publication listing · Preprint.
How does it work?
The authors describe a scheduling-based approach. A fast, nondeterministic path generates candidate tokens. A verifier checks a token window under a fixed-shape reduction schedule, commits consistent tokens and rolls back violations. The stated aim is to retain existing kernels and dynamic batching while controlling the result. Research summary.
Why does it matter?
The publication attributes variability to floating-point arithmetic interacting with batching and GPU reduction order. Its proposed control therefore reaches below the prompt and sampling settings. Research summary.
Our take: when comparing approaches, ask which computations are fixed, what verification costs under your workload, and which hardware and model versions the guarantee covers.
What remains to be tested?
We read the publication summary and preprint abstract, not the full paper or an independently reproduced benchmark. The retrieved summary provides no quantitative performance results or cross-hardware guarantee. Reproducibility and correctness should be tested separately.
Start with how to test determinism and use the Deterministic AI Checklist to organize the evidence.
The source record
Read the original evidence and the scope of our review.
- LLM-42: Enabling Determinism in LLM Inference with Verified SpeculationMicrosoft Research · 2026-09 · Accessed 2026-10-11Substantive indexed publication summary read; conference listing gives month only.
- LLM-42: Enabling Determinism in LLM Inference with Verified SpeculationarXiv · 2026-01-30 · Accessed 2026-10-11Abstract and revision metadata read; full paper not read.