Start with the serving stack, not just the prompt. Record the supported feature settings, then test the same request alongside different neighboring requests. Keep the resulting claim within that configuration.
What can teams try?
vLLM documents a batch-invariance option enabled with VLLM_BATCH_INVARIANT=1. It describes the feature as beta and intended to make outputs independent of batch size and request order. The retrieved documentation lists supported hardware and usage examples. vLLM documentation.
SGLang documents --enable-deterministic-inference, disabled by default. Its explanation connects batch-dependent GPU reduction order and floating-point arithmetic to inconsistent inference. It uses batch-invariant operators to address that source of variability. SGLang documentation.
What should you check before using it?
Support is configuration-dependent. vLLM’s retrieved page calls the feature beta. SGLang lists supported attention backends and notes that FlashInfer lacks radix-cache support in the documented deterministic mode. vLLM · SGLang.
Our take: repeat your workload with different request groupings and concurrency, while recording model, software, hardware and sampling settings. Check the output you actually need to reproduce, then measure the performance cost.
Is this a new release announcement?
No. This is a review of current documentation on October 11, 2026. Neither retrieved page displayed a publication date. We read substantive implementation guidance, but the recovered text was incomplete, and we have not run either implementation. Recheck the current documentation before deployment.
The testing guide provides a broader workflow protocol; is an LLM deterministic? explains the scope of the question.
The source record
Read the original evidence and the scope of our review.
- Batch InvariancevLLM · Publication date not displayed · Accessed 2026-10-11Indexed beta status, supported hardware and usage examples read; text truncated after offline example.
- Deterministic InferenceSGLang · Publication date not displayed · Accessed 2026-10-11Indexed mechanism, enable flag and backend support read; text truncated during MoE example.