Start with a run manifest. Before invoking replay, list the operations that can execute and the results that will be restored. That one step prevents a test report from treating a saved response as evidence of a fresh computation.
The distinction depends on the system. In LangGraph’s time-travel documentation, nodes after a checkpoint execute again, including LLM calls, API requests and interrupts. Earlier nodes are not rerun. In Temporal’s Activity documentation, completed Activity results are recorded in history and reused on Workflow replay. These are different execution boundaries. LangGraph · Temporal.
What should the manifest contain?
Here is a suggested record format. It is illustrative YAML, not a configuration file for either framework:
mode: checkpoint_resume
checkpoint: recorded-checkpoint-id
restored_results: [document_lookup]
reexecuted_steps:
- draft_answer
- validate_claims
external_actions: test_double
comparison: final_answer_and_validation
Replace the step names with the actual workflow steps. Store the input, application version, rule version and relevant environment alongside it. If a value is unknown, make that gap explicit before interpreting the comparison.
How do you run a useful comparison?
Use three separate checks where they apply:
- Inspect the recorded run. Confirm what the original trace actually contains. This is evidence inspection, not fresh execution.
- Resume from a checkpoint. Record restored and re-executed steps, then compare the specified downstream results.
- Run again from the input. Record the new execution conditions and compare outputs and actions against the stated contract.
These are our recommended test categories, not claims that both frameworks expose identical commands or guarantees. LangGraph’s documented reruns can produce different results. Temporal’s history reuse means a completed Activity need not execute again during replay. LangGraph · Temporal.
What about external actions?
For an initial test, use a controlled destination or a test double and record that choice. A receipt retrieved from history demonstrates what was recorded earlier; it does not demonstrate that a payment or message was sent again.
The trade-off is straightforward. A test double makes repeated tests easier to inspect, but it does not establish the behavior of the live destination. A later controlled integration test must address that boundary separately. This is an engineering recommendation, not a performance finding from the documentation.
What counts as a passing result?
Define that before running the test. Exact text equality, the same approved decision and the same external-action outcome are different criteria. Choose the one the claim promises, retain discrepancies and report the execution mode with the result.
Use the determinism testing guide for the comparison protocol and the Deterministic AI Checklist for the evidence packet. Our review recovered substantive documentation passages, but did not execute either framework or test a deployed agent.
The source record
Read the original evidence and the scope of our review.
- Use time-travelLangChain · Not stated · Accessed 2026-10-11Indexed Overview, Replay and Fork sections with examples read; extract truncated.
- Workflow ActivityTemporal · Not stated · Accessed 2026-10-11Substantive indexed text on execution, history and replay read; full live-page body unavailable.
