Request a scoping call Contact
← Research

Give long-running agents a working account of the task

New research on explicit belief states gives agents a way to track what remains unknown and recognise stalled progress. The opportunity is a more useful investigation workflow, with a measurable cost for maintaining that state.

Architecture / Conceptual study
Make the connections explicit.
  1. Applications
  2. Interfaces
  3. Records

Trace the interfaces between applications and records before changing a system.

An agent investigating a disputed delivery needs to keep track of more than its conversation. It needs a current account of the evidence, the unresolved discrepancy and the next observation that could settle it. A longer transcript can preserve every unsuccessful search while leaving that task unclear.

A preprint released on 1 October 2026 proposes Progression of States (PoS), an inference-time framework for maintaining explicit belief states. Here, a belief is the agent’s revisable account of the task world. It is not a verified business record. The design combines that account with checks on its consistency and recovery when activity stops advancing the task.

The opportunity for an enterprise architect is a more capable investigation workflow: one that can continue through conflicting observations and missing evidence without repeatedly reconstructing the problem from its entire history. This is a design hypothesis for local testing, not an established enterprise outcome.

Make unfinished work visible

The paper’s method separates the current world estimate, the goal, missing knowledge and unfinished actions. It selects an active gap to focus the next step. A checking component reviews proposed state updates against the available evidence before they become the next decision context. Recovery responds to stagnation, recurrence or movement that leaves the relevant problem unresolved.

That separation suggests a useful application boundary. In a hypothetical delivery investigation, the order system says delivered while the customer says the parcel is missing. Those are two attributed observations. The agent should preserve the disagreement, identify the missing receipt and seek evidence that distinguishes a misdelivery from another explanation. Recording “delivery resolved” because a tool returned successfully would confuse a completed query with a completed investigation.

A small working record could hold the disputed delivery identifier, each source observation and its time, the open questions, the current evidence request and the handoff condition. Keep tentative interpretations visibly tentative. The record should make the next useful action easier to select and the eventual handoff easier to review.

hypothetical fixture

A new observation changes the next useful action

Explain why a delivery is disputed and prepare a supported handoff.

Step 5 of 5

Swipe or scroll for the full diagram →

Operational provenance sequenceRequest, permissions, source versions, answer validation and human handoff are visible operational records, not hidden reasoning.01 / Discrepancy recorded02 / Evidence still missing03 / Repeated search adds nothing04 / Different source obtained05 / Supported handoff prepared

Supported handoff prepared

A reviewer receives the conflicting records, receipt and unresolved address check. No refund or customer message has been authorised.

Discrepancy recorded
Order system says delivered. Customer says missing. The conflict is open, with both source records retained.
Evidence still missing
The delivery receipt has not been obtained. The existing order record cannot establish who received the parcel.
Repeated search adds nothing
The same status query returns the same record. No new evidence resolves the receipt question.
Different source obtained
A receipt names a reception desk at another building. Record its source and time, then check the address relationship.
Supported handoff prepared
A reviewer receives the conflicting records, receipt and unresolved address check. No refund or customer message has been authorised.

Constructed operational trace, not a PoS run or hidden model reasoning. Step through the records that would justify progress.

A fictional delivery investigation moves from an unresolved discrepancy to an evidence-linked handoff. Repeating an unchanged search leaves the task unresolved.

Reviewed 2026-10-04

Give recovery a direction

The authors’ implementation walkthrough describes a runtime that constructs a candidate update, checks it, permits a repair and commits only an accepted state. It also distinguishes a failed progress assessment from an assessment of no progress. That matters operationally: missing instrumentation should not be silently treated as evidence that an action was useless.

Its recovery instructions depend on the detected pattern and the unresolved requirement. They may suppress an ineffective transition, interrupt a cycle or redirect attention to the active gap. Recovery itself does not execute an action. It changes the context in which the next action is selected.

For the hypothetical investigation, repeating the same order-status search cannot produce a missing receipt. A useful recovery would seek an authorised evidence source with different information, or hand the case to the team that can obtain it. Rewording the same query indefinitely adds activity without resolving the question. Conversely, waiting for a promised document may be appropriate. A local progress rule must recognise legitimate waiting rather than forcing another tool call.

This also offers a better human handoff. A reviewer can receive the unresolved question, the observations supporting it and the source still needed. They need not infer the task’s position from a long sequence of searches. That benefit remains to be measured, including the effort required to correct an inaccurate working record.

Count the work of maintaining state

The authors report higher overall results across four benchmarks and three model backbones. The tests cover task execution and diagnosis. These are author-reported experiments, not an independent replication or a production service assessment.

There is a material compute trade-off. In the reported RCA-100 comparison, using Qwen3.7-Plus across 103 cases, average total input and output tokens per episode rose from 355,700 for the raw-history baseline to 1,800,130 for PoS, approximately 5.06 times as many. Task-agent tokens fell, but belief maintenance and checking increased the total. Token counts do not establish monetary cost, latency or staffing benefit.

The reproduction guide makes further limits explicit. The release includes PoS, a raw-history baseline and two ablations, but omits four other comparator implementations. Historical experiment outputs are not bundled, and the manuscript’s live API models have no fixed snapshots. Running the released code later therefore cannot be assumed to reproduce every published row. This article does not report a model run.

illustrative fixture

What supports a trial of explicit task state?

Swipe or scroll for the full diagram →

The evidence stops before the proposed applicationObserved finding connects to an inference. A dashed boundary separates that inference from the untested application.ObservedInferenceUntested application
The evidence stops before the proposed applicationObserved finding and inference appear above a dashed boundary. The untested application remains below it.Observed findingInferenceUntestedapplication

Explicit task state is a candidate design for investigations that repeatedly lose progress.

Observed
The authors report better overall results across their four benchmarks and three model backbones, with additional inference work.
Inference
A read-only delivery investigation could benefit from tracking missing evidence and recognising repeated unsuccessful searches.
Evidence boundary
No result here establishes delivery-case quality, response time, operating cost or safe financial authority.

Research evidence, architecture inference and local unknowns have different status.

Author-reported benchmark improvements motivate a bounded local trial. Enterprise outcomes and economics remain untested.

Reviewed 2026-10-04

Start with a bounded investigation

Choose an existing read-only case workflow where repeated searches or lost context are observable problems. Compare the incumbent agent with explicit task state on the same cases, model configuration, tool permissions and time allowance. Include straightforward cases: maintaining an elaborate state may add cost where a single lookup already works.

Use cases with contradictory status records, an unavailable evidence source, a changed observation and a legitimate waiting period. Have the service team judge whether the final explanation is supported and whether the handoff preserves the unresolved work. Record completed cases, incorrect resolutions, repeated queries, elapsed time, all model calls and reviewer effort. These measures test the proposed operational benefit rather than rewarding a neatly populated state object.

Retain the controls already needed at the action boundary. The archive’s retry contract prevents repeated business effects, while approval preconditions address changes after review. Neither function should be delegated to a model-written belief. In the delivery trial, the agent can investigate and prepare a handoff while refunds and external messages remain separately authorised.

The architecture owner should fund a bounded comparison where maintaining task state could recover useful work. Expansion should depend on supported resolutions and better handoffs at an acceptable total cost. If the extra checking merely produces a more elaborate account of the same unresolved cases, keep the simpler workflow.

Filed under · Architecture · Agents · Task state · Inference-time compute Inference Institute · 04 Oct 2026

Related engagement

The decision behind this article

A clear design your team or chosen delivery partner can build from.

Explore AI Architecture →

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.