Give long-running agents a working account of the task
New research on explicit belief states gives agents a way to track what remains unknown and recognise stalled progress. The opportunity is a more useful investigation workflow, with a measurable cost for maintaining that state.
Trace the interfaces between applications and records before changing a system.
An agent investigating a disputed delivery needs to keep track of more than its conversation. It needs a current account of the evidence, the unresolved discrepancy and the next observation that could settle it. A longer transcript can preserve every unsuccessful search while leaving that task unclear.
A preprint released on 1 October 2026 proposes Progression of States (PoS), an inference-time framework for maintaining explicit belief states. Here, a belief is the agent’s revisable account of the task world. It is not a verified business record. The design combines that account with checks on its consistency and recovery when activity stops advancing the task.
The opportunity for an enterprise architect is a more capable investigation workflow: one that can continue through conflicting observations and missing evidence without repeatedly reconstructing the problem from its entire history. This is a design hypothesis for local testing, not an established enterprise outcome.
Make unfinished work visible
The paper’s method separates the current world estimate, the goal, missing knowledge and unfinished actions. It selects an active gap to focus the next step. A checking component reviews proposed state updates against the available evidence before they become the next decision context. Recovery responds to stagnation, recurrence or movement that leaves the relevant problem unresolved.
That separation suggests a useful application boundary. In a hypothetical delivery investigation, the order system says delivered while the customer says the parcel is missing. Those are two attributed observations. The agent should preserve the disagreement, identify the missing receipt and seek evidence that distinguishes a misdelivery from another explanation. Recording “delivery resolved” because a tool returned successfully would confuse a completed query with a completed investigation.
A small working record could hold the disputed delivery identifier, each source observation and its time, the open questions, the current evidence request and the handoff condition. Keep tentative interpretations visibly tentative. The record should make the next useful action easier to select and the eventual handoff easier to review.
A new observation changes the next useful action
Explain why a delivery is disputed and prepare a supported handoff.
Swipe or scroll for the full diagram →
Supported handoff prepared
A reviewer receives the conflicting records, receipt and unresolved address check. No refund or customer message has been authorised.
- Discrepancy recorded
- Order system says delivered. Customer says missing. The conflict is open, with both source records retained.
- Evidence still missing
- The delivery receipt has not been obtained. The existing order record cannot establish who received the parcel.
- Repeated search adds nothing
- The same status query returns the same record. No new evidence resolves the receipt question.
- Different source obtained
- A receipt names a reception desk at another building. Record its source and time, then check the address relationship.
- Supported handoff prepared
- A reviewer receives the conflicting records, receipt and unresolved address check. No refund or customer message has been authorised.
Constructed operational trace, not a PoS run or hidden model reasoning. Step through the records that would justify progress.
A fictional delivery investigation moves from an unresolved discrepancy to an evidence-linked handoff. Repeating an unchanged search leaves the task unresolved.
Reviewed 2026-10-04
Give recovery a direction
The authors’ implementation walkthrough describes a runtime that constructs a candidate update, checks it, permits a repair and commits only an accepted state. It also distinguishes a failed progress assessment from an assessment of no progress. That matters operationally: missing instrumentation should not be silently treated as evidence that an action was useless.
Its recovery instructions depend on the detected pattern and the unresolved requirement. They may suppress an ineffective transition, interrupt a cycle or redirect attention to the active gap. Recovery itself does not execute an action. It changes the context in which the next action is selected.
For the hypothetical investigation, repeating the same order-status search cannot produce a missing receipt. A useful recovery would seek an authorised evidence source with different information, or hand the case to the team that can obtain it. Rewording the same query indefinitely adds activity without resolving the question. Conversely, waiting for a promised document may be appropriate. A local progress rule must recognise legitimate waiting rather than forcing another tool call.
This also offers a better human handoff. A reviewer can receive the unresolved question, the observations supporting it and the source still needed. They need not infer the task’s position from a long sequence of searches. That benefit remains to be measured, including the effort required to correct an inaccurate working record.
Count the work of maintaining state
The authors report higher overall results across four benchmarks and three model backbones. The tests cover task execution and diagnosis. These are author-reported experiments, not an independent replication or a production service assessment.
There is a material compute trade-off. In the reported RCA-100 comparison, using Qwen3.7-Plus across 103 cases, average total input and output tokens per episode rose from 355,700 for the raw-history baseline to 1,800,130 for PoS, approximately 5.06 times as many. Task-agent tokens fell, but belief maintenance and checking increased the total. Token counts do not establish monetary cost, latency or staffing benefit.
The reproduction guide makes further limits explicit. The release includes PoS, a raw-history baseline and two ablations, but omits four other comparator implementations. Historical experiment outputs are not bundled, and the manuscript’s live API models have no fixed snapshots. Running the released code later therefore cannot be assumed to reproduce every published row. This article does not report a model run.
What supports a trial of explicit task state?
Swipe or scroll for the full diagram →
Explicit task state is a candidate design for investigations that repeatedly lose progress.
- Observed
- The authors report better overall results across their four benchmarks and three model backbones, with additional inference work.
- Inference
- A read-only delivery investigation could benefit from tracking missing evidence and recognising repeated unsuccessful searches.
- Evidence boundary
- No result here establishes delivery-case quality, response time, operating cost or safe financial authority.
Research evidence, architecture inference and local unknowns have different status.
Author-reported benchmark improvements motivate a bounded local trial. Enterprise outcomes and economics remain untested.
- PoS reported results · Metrics, Table 1 and inference overhead
Reviewed 2026-10-04
Start with a bounded investigation
Choose an existing read-only case workflow where repeated searches or lost context are observable problems. Compare the incumbent agent with explicit task state on the same cases, model configuration, tool permissions and time allowance. Include straightforward cases: maintaining an elaborate state may add cost where a single lookup already works.
Use cases with contradictory status records, an unavailable evidence source, a changed observation and a legitimate waiting period. Have the service team judge whether the final explanation is supported and whether the handoff preserves the unresolved work. Record completed cases, incorrect resolutions, repeated queries, elapsed time, all model calls and reviewer effort. These measures test the proposed operational benefit rather than rewarding a neatly populated state object.
Retain the controls already needed at the action boundary. The archive’s retry contract prevents repeated business effects, while approval preconditions address changes after review. Neither function should be delegated to a model-written belief. In the delivery trial, the agent can investigate and prepare a handoff while refunds and external messages remain separately authorised.
The architecture owner should fund a bounded comparison where maintaining task state could recover useful work. Expansion should depend on supported resolutions and better handoffs at an acceptable total cost. If the extra checking merely produces a more elaborate account of the same unresolved cases, keep the simpler workflow.