Long-running agents make context an accumulating operational cost
Source → Observed → Interpretation → Model implication
Improved token efficiency for longer agent runs
View source →Cursor reports that agents now run longer and carry more accumulated context between steps, increasing the importance of how its harness assembles, reuses, and partitions context.
Cursor says changes across prompt construction, context reuse, and work division reduced token costs for users by 7% without reducing agent quality.
Cursor removed roughly 66% of its system prompt as stronger models required less explicit instruction across model families. It says production A/B testing was important because offline evals overrepresent hard tasks relative to real user traffic.
Moving MCP tool definitions into dynamic context had reduced total tokens by 46.9% across sessions that invoked an MCP tool.
Cursor A/B tested which built-in tools needed to remain visible from the start, tracking tokens, cost, latency, tool-call errors, and overall agent usage. Dynamically loading the remaining tools cut static-context tool- description tokens by 60%.
Adding explicit cache breakpoints after stable request layers and moving variable setup behind those boundaries reduced Cursor's cold-cache-miss rate by 20%.
Numbering only every tenth line in file-read output reduced cache-read tokens by 1.6% with no reported reduction in quality.
Cursor says subagents can reduce token spend because each typically starts with fresh context and returns a result without adding its full working history to the parent context.
Cursor also reports a coordination tax when isolated subagents duplicate work or continue tasks that are no longer necessary. It removed instructions that strongly encouraged subagents for codebase exploration and constrained alternate-model selection to cases directed by the user or harness.
Context becomes an environmental cost that accumulates as agent runs lengthen. Cursor's production experiments show Selection operating on harness variants: prompts, tool visibility, cache layout, file representation, and delegation policy survive when they preserve reported quality while lowering token cost or avoiding errors and latency. The subagent findings refine Cooperation because delegation is not monotonically beneficial. Partitioning work can isolate context and reduce inference cost, but isolation also weakens shared state and can create duplicated or obsolete work. Cooperation therefore has an economic fitness function: context savings must exceed the coordination tax.
REFINES. Supports Selection with production A/B evidence that harness configuration materially affects agent economics without a reported quality loss. Refines Cooperation by identifying context partitioning as a benefit of subagents and coordination overhead as its countervailing cost. This makes efficient cooperation conditional on the surrounding context and coordination design, rather than on increasing the number of delegated agents.
What this does not establish
- The 7% user token-cost reduction aggregates changes across several harness layers and does not isolate the contribution of each intervention.
- Cursor does not publish sample sizes, experiment durations, model mix, confidence intervals, statistical significance, or its definition and measurements of unchanged agent quality.
- The 46.9% result applies only to sessions that invoked an MCP tool and should not be interpreted as a reduction across all agent sessions.
- The 60% result concerns static-context tool-description tokens, the 20% result concerns the rate of cold cache misses, and the 1.6% result concerns cache-read tokens; none is an end-to-end user cost reduction by itself.
- Cursor does not quantify the token savings from subagents, the frequency or cost of duplicated work, or the net effect of its revised delegation policy.
- Removing explicit encouragement of subagents demonstrates policy adjustment, not that a particular level or topology of cooperation is generally optimal.
- The results are first-party and specific to Cursor's production traffic, models, agent harness, pricing, caching behavior, and tool distribution.
Across longer engineering tasks, when do the context savings from subagent isolation exceed the added token, latency, duplicated-work, and reconciliation costs of coordination?