Kurtel is in early access — we are onboarding a small cohort of teams.Become a design partner
Research note

The right context makes or breaks a coding task — and matters more as tasks grow

A coding agent is only as good as what it knows when it starts. Research published in 2026 measures this precisely: every missing piece of context costs its share of the work, agents often fail to find it on their own, and the bigger the change, the more pieces there are to miss. The answer is not more context — it is the right context.

01Every missing fact costs exactly its share of the work

To change code in a real repository, an agent has to know a set of facts that depend on each other: which tests cover it, what it imports, how it is configured, which migration rules apply. In a 2026 study, researchers removed those facts from the agent one at a time. The result is strikingly linear: each withheld fact costs exactly the work it supports, and nothing more[1].

Each missing fact costs exactly the work it supports
0816243201234facts withheld from the agenttests passed (out of 32)observedlinear prediction

Hover a point to see the exact value.

Mohammadi et al., 2026. Fault injection over 30 trials: each withheld fact supports four tests; passed tests follow the linear prediction almost exactly (dashed).

Worse, agents rarely notice what they are missing. “An agent asked to do something does something, so a missing fact yields a confident wrong edit”[1]. Some agents reported the missing file in every trial; others never did[1]. And it does not need to sit right next to the code being changed: a fact given to the agent works just as well when it lives elsewhere in the repository — what matters is that the agent has it[1].

02Agents often fail to find it on their own

Before writing a patch, an agent has to find the files the task needs. A 2026 benchmark built from real coding workflows measured exactly that step: in roughly one case out of three (between 27% and 35% of tasks, depending on the setting), the agents' own exploration missed every file they needed[2]. Seeding them with the right context up front gave better results with less exploration — and ideal context still left substantial headroom[2].

1 in 3

tasks where the agent's own search missed every file it needed (27% to 35%)[2]

12.8×

more tokens for the same result, when the agent keeps going back to fetch its context[1]

Finding context the hard way is also expensive. In the working-set study, every agent setup passed all the tests, and each had about the same amount of context in front of it at any moment. Yet total token use differed by 12.8× — from under 300,000 to over 3.7 million[1]. The difference was how often each agent had to go back and rebuild its context: 5 tool calls for the cheapest, 79 for the most expensive — and every extra round re-sends the whole conversation[1]. Same result, same knowledge; the cost is in fetching it again and again.

Same result, twelve times the tokens
5 tool calls
0.29M
79 tool calls
3.75M

Mohammadi et al., 2026. All setups pass every test and hold about the same context at any moment; total input differs 12.8× because some rebuild their context far more often (5 vs 79 tool calls).

03The bigger the change, the more there is to miss

If each fact costs its share, a task that depends on more facts has more to lose. That is what happens as changes grow. The same model, GPT-5.2, solves 72.8% of single-issue tasks on SWE-Bench Verified — and 22.9% of release-level changes spanning 21 files on average[3].

Same model, bigger change: success collapses
Single-issue tasks
72.8%
Changes across ~21 files
22.9%

Le et al., SWE-EVO (v6, 2026): GPT-5.2 on SWE-Bench Verified (single issues) vs SWE-EVO (release-level changes spanning 21 files on average).

Large refactorings tell the same story: across tasks touching 11.4 files on average, the best model resolves only 41.2%[4]. Real work on a real codebase looks far more like these tasks than like isolated bug fixes — and that is where knowing the right facts decides the outcome.

72.8% → 22.9%

GPT-5.2, single issues vs changes across ~21 files[3]

41.2%

best resolve rate on refactorings touching 11.4 files[4]

04More context is not the answer

The obvious fix — give the agent more — backfires. On the same task, as the context grows with more material of the same kind, accuracy collapses: Claude Opus 4.5 falls from 96% to 15% between 8K and 256K tokens[5].

Same task, more context: accuracy falls as the context grows
0%25%50%75%100%8K16K32K64K96K128K256K
Claude Opus 4.5GPT-5.2Gemini 3 Flash

LOCA-bench (Zeng, Huang, He, 2026), Table 1. The task stays identical; only the amount of environment context grows. Context sizes are evenly spaced for readability.

And writing it all into a file does not work either: repository context files such as AGENTS.md or CLAUDE.md did not generally improve task success, while raising inference cost by over 20%[6].

What this means

Success depends on the agent having every fact the task relies on — and nothing it does not. Missing facts turn into confident mistakes; extra context dilutes attention and raises the bill. As tasks grow, both effects grow with them. The work that matters is selection: the few facts this task depends on, available from the start.

This is what Kurtel does for every task: it follows the code map to what the change touches, adds the team knowledge that applies to it, and sends nothing else. See how the context engine works →

Read next: why agent skills fall short on real codebases →

Sources

  1. [1]Mohammadi, Klein, Chadha, Arora, Bindschaedler, The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks, arXiv 2608.16630, Aug. 2026
  2. [2]Qin, Xie, Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents, arXiv 2607.24882, Jul. 2026
  3. [3]Le et al., SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios, arXiv 2512.18470 (v6, May 2026)
  4. [4]Shi et al., SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring, arXiv 2608.09802, Aug. 2026
  5. [5]Zeng, Huang, He, LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth, arXiv 2602.07962, Feb. 2026
  6. [6]Gloaguen, Mündler, Müller, Raychev, Vechev (ETH Zurich), Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?, arXiv 2602.11988, 2026

Figures are quoted from each paper. SWE-EVO first appeared in December 2025; figures come from its May 2026 revision.