THE SIGNAL IN ONE SENTENCE

A coding agent can drown in its own receipts. Every file view, search result, stack trace and test log lands in the conversation, then gets carried into later model calls. The longer the job runs, the more of yesterday's terminal output the model rereads before deciding what to do next. Three researchers at Peking University, Zhensheng Zou, Guoqing Wang and Dan Hao, have tested a different memory arrangement. Their September 25 preprint calls it Latent Observations, Hard Actions, or LOHA. Older tool observations are compressed into continuous representations called soft tokens. The agent's own reasoning and tool calls remain ordinary text. The newest K observations also remain ordinary text, because a coding tool often needs the exact path, identifier or source line it just saw. A companion training method, Anchored Context Distillation, teaches the adapted model to read the compressed material while nudging its full-text behavior toward the unmodified base model. That is the clever part. The honest part is the bill. With the recent-text window set to three observations, the researchers report 43 percent less context per call for Qwen3-4B and 57 percent less for SWE-Master-4B-RL on SWE-bench Verified. Those agents also resolved fewer issues than their uncompressed bases: 12.1 percent versus 14.5 percent for Qwen3, and 21.8 percent versus 27.5 percent for SWE-Master. A larger recent window generally recovered some performance while giving back compression. In a single-run sweep, K equals eight reached 14.4 percent and 23.0 percent. That makes latent memory a dial, not a trophy. Turn it toward shorter contexts and the server can fit more work. Turn it toward visible text and the agent gets more exact evidence. The paper also reports a case where compression helps capability rather than merely price. Under a 32,768-token cap, Hard-Last-3 Qwen3 resolved 42 of a 199-instance subset, while the same adapted model using full text resolved 22. The compressed history kept more useful past information inside the limit. Yet each capped condition was run once, and the subset is not the complete benchmark. This is encouraging engineering evidence, not a production guarantee. The serving result is similarly specific. On one GPU with sixteen concurrent Qwen3 trajectories, LOHA completed 63.4 instances per GPU-hour compared with 33.4 for the adapted full-text agent. Median call latency fell from 5.07 seconds to 1.71 seconds. The current server also re-encoded old observations on every call, adding encoder work. A replay with per-observation caching estimated that most of that extra work could disappear, but an estimate is not a deployed measurement. Masking, the simpler approach of dropping older observations, reached 64.7 instances per GPU-hour in the same comparison. Compression preserved more historical information than masking, but it did not own the speed chart. Exact detail is where the method shows its seams. At the default sixteen-times compression ratio, a separate exact-recitation test recovered only 15.1 percent of 358 literals later quoted by actions. That is why LOHA keeps recent observations in text and allows old material to be fetched again through tools. A soft representation may preserve the gist of an earlier test failure while losing the punctuation needed for a precise replacement. It may remember that a configuration file mattered while blurring the value the agent must reproduce. The product implication is not to hide the original log. Keep an immutable human-readable trace outside the model context, record which observations were compressed, and let the agent or operator reopen the exact source. Otherwise a memory-saving feature becomes an audit-deleting feature. The paper also separates reading from behavior. A model can become better at reconstructing compressed text and worse at completing the coding task. In one small ablation, a student trained from a larger teacher exactly recited more literals but resolved fewer tasks than the self-anchored student. Anchoring matters because learning a new input representation can disturb the policy that chooses tools, edits code and decides when to stop. This is an important warning for every compressed-memory product. A good reconstruction score does not prove a good agent. The operational test is whether the complete trajectory reaches a correct patch without looping, quoting the wrong string or silently abandoning necessary evidence. For teams building coding agents, the practical design is a memory ladder. Keep the system instructions, task description, assistant actions and a recent evidence window as readable text. Compress older observations only after their exact source is safely stored. Attach an identifier to each compressed block so a reviewer can reopen it. Trigger retrieval when the agent needs a literal path, command, line or error message. Test several K values on the organization's own repositories because the right window depends on task shape, tool design and context limit. Report resolved tasks beside tokens, latency, throughput and retrieval count. A cheaper failed patch is still a failed patch. Human inspection needs its own evaluation. Can an engineer reconstruct why the agent changed a line? Can a security reviewer see whether a malicious instruction entered through a tool response before it was compressed? Can an incident investigator compare the latent view with the original observation? Can the system prove which exact text was available when a tool call was made? Soft tokens are not naturally readable, so the surrounding product has to supply this evidence. Compression may reduce what the model rereads. It should not reduce what the organization can audit. China matters here because the work comes from a Peking University team tackling a problem that every long-running agent faces: memory is becoming infrastructure. The result is not another giant model announcement. It is a systems proposal about where exact text belongs, which history can become approximate, and how training can add a new memory representation without replacing the agent's whole behavior. That is useful beyond coding. Research assistants, support agents and operations tools all accumulate observations. The boundary should remain the same: compress the material used for orientation, preserve or retrieve the material needed for exact action, and keep the original evidence available to people. The plain signal is pleasantly inconvenient. Latent memory can buy shorter contexts and more concurrent work. It can also lower task success and make the agent's past harder to inspect. The winning product will not compress everything and announce victory. It will expose the dial, preserve the receipts, measure the errors and know when one missing character is the whole job.

01

WHAT ACTUALLY CHANGED

Researchers Zhensheng Zou, Guoqing Wang and Dan Hao at Peking University submitted the preprint on September 25.

The proposed LOHA layout compresses older tool observations into soft tokens.

The system prompt, task description and the agent's own turns remain ordinary text.

The K most recent tool observations remain ordinary text for exact reference.

At the default sixteen-times ratio, roughly sixteen text tokens become one continuous embedding.

Anchored Context Distillation teaches the adapted agent to read the latent view while anchoring full-text behavior to the base model.

The experiments use Qwen3-4B-Instruct-2507 and SWE-Master-4B-RL on SWE-bench Verified.

With K equal to three, context per call falls by 43 percent for Qwen3 and 57 percent for SWE-Master.

Qwen3 resolves 12.1 percent with Hard-Last-3 versus 14.5 percent for its uncompressed base.

SWE-Master resolves 21.8 percent with Hard-Last-3 versus 27.5 percent for its uncompressed base.

A single-run sweep at K equal to eight reaches 14.4 percent for Qwen3 and 23.0 percent for SWE-Master.

Under a 32K limit, Hard-Last-3 Qwen3 resolves 42 of 199 tasks versus 22 for the adapted full-text condition.

On one GPU with sixteen concurrent Qwen3 trajectories, LOHA completes 63.4 instances per GPU-hour versus 33.4 for the adapted full-text agent.

Median call latency in that serving test falls from 5.07 seconds to 1.71 seconds.

The current server re-encodes old observations on each call, while a replay estimates that per-observation caching could sharply reduce encoder work.

An exact-recitation test recovers 15.1 percent of 358 later-used literals at sixteen-times compression.

The authors propose learning when to retain higher-fidelity observations and using objectives tied more directly to tool decisions and trajectories.

02

WHY THIS MATTERS

Tool results can dominate a coding agent's growing context and make every later call more expensive.

Keeping recent evidence in text recognizes that code edits often require exact strings rather than a semantic summary.

Compressed history can keep more useful evidence inside a strict context window.

Shorter per-call context can improve concurrency when the key-value cache is the serving bottleneck.

The reported throughput gain compares with an adapted full-text agent and does not establish the same gain over every baseline.

Masking is slightly faster in the reported Qwen3 serving test, so latent compression must justify its extra complexity through retained information.

Lower context use does not automatically mean lower total compute because the encoder also consumes work.

Re-encoding old observations can erase part of the efficiency gain unless representations are cached safely.

The default compressed agents resolve fewer full-window benchmark tasks than their uncompressed bases.

A larger recent-text window can recover performance while surrendering some compression.

Exact recitation remains weak, which makes retrieval of original observations a core capability rather than an optional patch.

A model can read compressed content better while behaving worse as an agent.

Human reviewers cannot inspect soft tokens directly, so compression creates a new auditability burden.

Security teams need the original observation to investigate prompt injection or tool-output manipulation.

A one-run subset result should guide more testing, not settle a deployment decision.

The price-performance setting should depend on task risk, context pressure and the cost of one missing detail.

FIG. 262COMPRESS AN AGENT MEMORY WITHOUT BURNING THE RECEIPTS
1STORE THE RAW TOOL RESULT→
2KEEP THE NEWEST EVIDENCE AS TEXT→
3ENCODE OLDER OBSERVATIONS→
4LINK EACH LATENT BLOCK TO ITS SOURCE→
5LET THE AGENT REASON OVER BOTH VIEWS→
6REOPEN EXACT TEXT BEFORE A PRECISE ACTION→
7GRADE THE PATCH→
8COMPARE SAVINGS WITH LOST TASKS→
9PRESERVE THE AUDIT TRAIL
The short context is a working view, not the record. Exact source text stays available whenever the agent or a person needs to verify one detail.

03

WHERE IT COULD HELP

  • Keep an immutable raw trace of every tool observation outside the model context.
  • Assign stable identifiers linking each latent block to its original text.
  • Keep system instructions, task goals, agent actions and recent observations readable.
  • Test several recent-window sizes instead of adopting K equal to three as a universal default.
  • Trigger exact retrieval before edits, commands or tool calls that quote literals.
  • Cache compressed observations per source version so old material is not re-encoded on every call.
  • Invalidate cached memory when the underlying file, tool result or repository revision changes.
  • Measure resolved tasks beside input tokens, encoder work, latency and throughput.
  • Compare latent compression with masking, summarization, pruning and ordinary retrieval.
  • Evaluate long tasks separately from short tasks that rarely approach the context limit.
  • Create detail-critical tests containing paths, identifiers, punctuation, hashes and exact replacement strings.
  • Log every time the agent reopens raw evidence after consulting compressed memory.
  • Give operators a readable timeline that never depends on decoding soft tokens.
  • Preserve provenance for observations that may contain untrusted instructions.
  • Require human review when compression precedes a high-consequence code or infrastructure change.
  • Publish the chosen K value, compression ratio, model, context cap and cache policy with every result.

KEEP A HAND ON THE WHEEL

This is a September 25 preprint evaluated on SWE-bench Verified with two related four-billion-parameter agents. Full-window results show lower resolve rates for the default compressed condition than for the uncompressed bases. The K equal to eight recency sweep is a single run. The 32K comparison uses a 199-instance subset with one run per capped condition. The serving test uses one GPU and sixteen concurrent Qwen3 trajectories, and the 1.9-times throughput claim is relative to the adapted full-text agent rather than every baseline. The current server re-encodes old observations on each call, so some lower encoder-cost figures are replay estimates with caching. Exact recitation remains low even at gentler compression. Soft tokens are not human-readable, and the paper does not provide a production security, auditability or incident-response study. Watch for independent reproduction, released training and serving code, evaluations on larger agents and real repositories, cached production measurements, human audit tests, prompt-injection analysis, retrieval quality and task-specific policies for choosing what remains exact.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on September 28, 2026.

PUBLICATION RECEIPT: Revision 1. Published September 28, 2026.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US