Scale Signal

Long-running agents make context an accumulating operational cost

September 23, 2026 · Cursor
Selection → CooperationContextCoordinationEconomicsExecution
Scale signal: Cursor describes A/B tests on a large production user base and reports an aggregate 7% reduction in user token costs without reduced agent quality. The article does not publish experiment sizes, traffic volume, time windows, confidence intervals, quality definitions, or statistical significance.
Evidence record

Source → Observed → Interpretation → Model implication

SOURCE

Improved token efficiency for longer agent runs

View source →
OBSERVED

Cursor reports that agents now run longer and carry more accumulated context between steps, increasing the importance of how its harness assembles, reuses, and partitions context.

Cursor says changes across prompt construction, context reuse, and work division reduced token costs for users by 7% without reducing agent quality.

Cursor removed roughly 66% of its system prompt as stronger models required less explicit instruction across model families. It says production A/B testing was important because offline evals overrepresent hard tasks relative to real user traffic.

Moving MCP tool definitions into dynamic context had reduced total tokens by 46.9% across sessions that invoked an MCP tool.

Cursor A/B tested which built-in tools needed to remain visible from the start, tracking tokens, cost, latency, tool-call errors, and overall agent usage. Dynamically loading the remaining tools cut static-context tool- description tokens by 60%.

Adding explicit cache breakpoints after stable request layers and moving variable setup behind those boundaries reduced Cursor's cold-cache-miss rate by 20%.

Numbering only every tenth line in file-read output reduced cache-read tokens by 1.6% with no reported reduction in quality.

Cursor says subagents can reduce token spend because each typically starts with fresh context and returns a result without adding its full working history to the parent context.

Cursor also reports a coordination tax when isolated subagents duplicate work or continue tasks that are no longer necessary. It removed instructions that strongly encouraged subagents for codebase exploration and constrained alternate-model selection to cases directed by the user or harness.

INTERPRETATION

Context becomes an environmental cost that accumulates as agent runs lengthen. Cursor's production experiments show Selection operating on harness variants: prompts, tool visibility, cache layout, file representation, and delegation policy survive when they preserve reported quality while lowering token cost or avoiding errors and latency. The subagent findings refine Cooperation because delegation is not monotonically beneficial. Partitioning work can isolate context and reduce inference cost, but isolation also weakens shared state and can create duplicated or obsolete work. Cooperation therefore has an economic fitness function: context savings must exceed the coordination tax.

MODEL IMPLICATION

REFINES. Supports Selection with production A/B evidence that harness configuration materially affects agent economics without a reported quality loss. Refines Cooperation by identifying context partitioning as a benefit of subagents and coordination overhead as its countervailing cost. This makes efficient cooperation conditional on the surrounding context and coordination design, rather than on increasing the number of delegated agents.

Epistemic boundaries

What this does not establish

OPEN QUESTION

Across longer engineering tasks, when do the context savings from subagent isolation exceed the added token, latency, duplicated-work, and reconciliation costs of coordination?