Kurtel is in early access — we are onboarding a small cohort of teams.Become a design partner
Research note

Why coding agents repeat the same mistakes

Every team that works with coding agents knows the feeling: you correct the agent, and a few sessions later it makes the same mistake again. Research published in 2026 measures it across thousands of real sessions — and shows what kind of memory fixes it, and what kind does not.

01The same mistakes come back

A 2026 study analysed 20,574 real coding-agent sessions across 1,639 repositories, looking for the moments where a developer had to push back on their agent[1]. The most common problem was not bad code. It was the agent ignoring the developer's own rules — in 38% of cases[1].

What goes wrong between developers and their agents
Ignores the developer's rules
38.3%
Misreads what was asked
27.0%
Reports success too early
22.6%
Writes faulty code
17.8%
Misreads the project
11.6%
Goes beyond what was asked
10.2%
Malformed commands
2.9%

Tang et al., 2026. Share of misalignment episodes showing each symptom; an episode can show several, so shares add up to more than 100%.

And these problems do not stay in one session. When a session went wrong, the next session in the same repository went wrong 51.9% of the time, against 33.6% otherwise — about one and a half times as often[1]. Over time, even as problems overall became less frequent, ignored rules grew as a share of what goes wrong[1].

A bad session makes the next one worse
Previous session went fine
33.6%
Previous session went wrong
51.9%

Tang et al., 2026 — 20,574 sessions, 1,639 repositories. Probability that the next session in the same repository contains a misalignment.

02The developer is the one who fixes it — every time

When these problems got resolved, 91.5% of the time it was because the developer corrected the agent explicitly — and in another 5.5%, the developer simply took over and fixed it themselves. The agent caught its own mistake in only 3% of cases[1].

Who fixes it: almost always the developer
The developer corrects the agent
91.5%
The developer takes over and fixes it
5.5%
The agent corrects itself
3.0%

Tang et al., 2026. Among episodes with a visible resolution (1,504, i.e. 9.33% of all episodes). The three add up to 100%.

Put the two findings together: developers are already teaching their agents, one correction at a time — and that teaching is lost at the end of the session. The next session, or the next developer, starts over.

03Memory helps — when it is the right memory

Giving agents the lessons of past work does pay off. A 2026 benchmark of 1,100 coding tasks found that accurately summarised and retrieved past experience significantly improves resolution, while cutting runtime and token cost — particularly on harder tasks[2]. Another, run on real repositories, injected verified experience into five different coding agents across 111 tasks: four of them resolved more tasks — up to 74.3% from 69.8% for the best gain — and every one of them needed fewer steps[3].

Tasks resolved, without memory vs with verified past experience
Without memory With verified experience
kimi-k2.7-code
69.8% → 74.3% (+4.5)
glm-5.2
76.8% → 80.4% (+3.6)
glm-5
55% → 57.4% (+2.4)
qwen3.8-max
79.1% → 80.2% (+1.1)
deepseek-v4-pro
67.1% → 67.1% (±0.0)

Fan et al., VibeMemBench, 2026 — 111 real repository tasks, five coding agents. Scale starts at 40% to make the gaps visible. Every agent also needed 2.2 to 10.9 fewer steps per run.

04Storing everything does not work

The same studies are just as clear about the other side. Unfiltered or wrongly selected past context brings limited or even negative benefits[2]. When four existing memory systems — Mem0, SimpleMem, MemoryOS and A-MEM — had to build and retrieve that experience themselves, eleven of twelve pairings with coding agents did no better than no memory at all[3]. The leading cause: raw conversation transcripts stored without filtering, polluting the agent's instructions[3].

11 of 12

pairings of existing memory systems and coding agents no better than no memory[3]

1.5×

as likely to go wrong when the previous session did (51.9% vs 33.6%)[1]

A review of 65 papers on self-improving coding agents reaches the same warning: memories can hurt future performance when they are retrieved uncritically[5].

05The agent cannot be its own teacher

Can the agent simply learn from its own experience? A 2026 benchmark of continual learning for agents found a clear line: repeated rounds of learning bring genuine improvement when the feedback comes from outside — while self-feedback alone leads to recursive drift[4].

That is the missing piece. The outside feedback that matters is already there, in every session: the developer's correction when the agent gets it wrong — and the passing result when it gets it right.

What this means

Agents repeat mistakes because what the team learns — the corrections that fix a task, and the approach that made one succeed — disappears with the session. Keeping everything is not the answer: raw history pollutes more than it helps. What works is memory built from outside feedback, verified, and handed to the agent only where it applies.

This is how Kurtel learns: from the corrections your developers already make, and from the tasks that succeed — so the next agent avoids the mistake and goes straight to what worked. Each lesson is attached to the code it concerns and delivered to every agent on the team only when the task touches that code. See how it learns →

Sources

  1. [1]Tang et al., How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions, arXiv 2605.29442, 2026
  2. [2]Zhu et al., SWE Context Bench: A Benchmark for Context Learning in Coding, arXiv 2602.08316, 2026
  3. [3]Fan et al., VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks, arXiv 2609.23570, Sept. 2026
  4. [4]Zhong et al., SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks, arXiv 2604.20087, Apr. 2026
  5. [5]Zhou et al., Self-Evolving Coding Agents (survey), arXiv 2608.03392, Aug. 2026

Figures are quoted from each paper. Resolution shares in [1] refer to the 1,504 episodes with a visible resolution.