Kurtel is in early access — we are onboarding a small cohort of teams.Become a design partner
Research note

Why agent skills fall short on real codebases

Skills have become the default way to give coding agents knowledge they don't have. They do help — under the right conditions. But the research published in 2026 points to three structural limits, and they matter most exactly where teams need help: large, long-lived codebases.

What skills get right

A skill is a package of instructions an agent can load when a task calls for it. To keep the context light, only each skill's name and description are loaded when a session starts; the full body loads when the model decides it is relevant[6]. And when an expert writes them, skills work: across 86 tasks, 11 domains and seven agent configurations, curated skills raised the average pass rate by 16.2 points[1]. Focused skills of two or three modules even beat comprehensive documentation[1].

The limits are not in the idea. They are in who writes the skill, who decides when it applies, and what it costs.

01A skill is only as good as its author

The gains above come from skills written by experts. When the agent writes its own, they vanish: self-generated skills bring no benefit on average — models cannot reliably author the procedural knowledge they benefit from consuming[1]. One-shot generated skills can be well formed yet behaviourally weak, while expert-written ones are costly and may not match how agents actually execute tasks[3]. Whether downloaded or self-generated, skills are often unreliable, incomplete or outdated[4].

Average pass-rate gain from skills, by who wrote them
Written by experts
+16.2
Written by the agent
≈ 0

SkillsBench v1: 86 tasks, 11 domains, 7 agent configurations, 7,308 runs. Self-generated skills: “no benefit on average”.

So a team faces a choice: maintain skills by hand, at a real and constant cost — or let the agent maintain them, and lose most of the benefit.

02The agent decides what applies — and it can get it wrong

With skills, the agent picks which one to load by reading a short description — and nothing checks that choice. A skill that looks right is no guarantee it helps.

When researchers took failed tasks and traced each one back to the skill that caused it, they found 307 cases[2]. In 125 of them, the skill led the agent to do the task wrong or to miss part of it. In the other 182, the task got done — but slower and with more tokens than without the skill[2].

Even skills written by experts are not immune: on 16 of 84 tasks, the agent did worse with them than without[1].

307 failures traced back to the skill that caused them
125 functional failures — the task was done wrong
182 efficiency regressions — done, but slower and with more tokens than without the skill

Dong et al., 2026. Among efficiency regressions, excessive verification alone accounts for 67 cases.

03Every skill has a cost beyond its size

A loaded skill does not only add its own tokens. It changes what the agent does next — for instance, making it double-check work that did not need checking. In the study above, that alone explains 67 of the 182 tasks that ran slower with a skill than without[2].

And every extra token has a price. On the same task, Claude Opus 4.5 falls from 96% to 15% accuracy as its context grows from 8K to 256K tokens[5].

The drop is not caused by noise: nothing off-topic is added. In the benchmark, an agent must find every exam in a set of courses, and the context grows with more courses, announcements and emails of the same kind[5]. Some of it is needed. Much of it only looks relevant: courses that are exempt, courses with no exam, assignments and quizzes mixed into the same announcements[5]. That is exactly what a skill or a memory brings when it is on the right topic but not quite about this task — and it is the kind of context that hurts most, because nothing marks it as irrelevant.

Same task, more context: accuracy falls as the context grows
0%25%50%75%100%8K16K32K64K96K128K256K
Claude Opus 4.5GPT-5.2Gemini 3 Flash

LOCA-bench (Zeng, Huang, He, 2026), Table 1. The task stays identical; only the amount of environment context grows. Context sizes are evenly spaced for readability.

On code, the gap is real — and static files don't close it

Across the eleven domains of SkillsBench, curated skills helped least in software engineering: +4.5 points, against +51.9 in healthcare[1]. It would be easy to read this as “coding agents don't need project knowledge”. Real sessions say the opposite.

Gain from expert-written skills, by domain
Healthcare
+51.9
All domains (average)
+16.2
Software engineering
+4.5

SkillsBench v1. Software engineering shows the smallest gain of the eleven domains; the average across all domains is +16.2 points.

An analysis of 20,574 real coding-agent sessions across 1,639 repositories found agents repeatedly misreading the project, misinterpreting intent and breaking its rules — and 91.5% of the problems that got resolved were resolved by the developer correcting the agent explicitly. Over time, constraint violations even grew as a share of these problems[7].

Yet writing that knowledge into a file does not fix it. A 2026 study from ETH Zurich tested repository context files — AGENTS.md, CLAUDE.md — across several coding agents and models: whether written by an LLM or by the project's own developers, they did not generally improve task success, while raising inference cost by over 20%[8]. The instructions were followed; broad repository overviews simply did not help[8].

91.5%

of resolved agent misalignments needed the developer to correct the agent[7]

+20%

inference cost from repository context files, with no general gain in success[8]

The knowledge agents lack is real — developers supply it by hand, session after session. What does not work is handing it over as a static file loaded whether or not it applies.

What this means

The hard part was never writing knowledge down. It is three things skills leave to chance: keeping that knowledge correct without an expert maintaining it, deciding with evidence — not the agent's own judgment — which piece applies to this task, and sending nothing more than that piece.

This is the problem Kurtel is built around: memory learned from your team's real corrections, attached to the code it concerns, and selected for each task by an independent scorer. See how the context engine works →

Sources

  1. [1]Li et al., SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks, arXiv 2602.12670, Feb. 2026
  2. [2]Dong et al., Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents, arXiv 2608.11888, Aug. 2026
  3. [3]Liu et al., SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision, arXiv 2606.01139, 2026
  4. [4]Wang et al., SkillGrad: Optimizing Agent Skills Like Gradient Descent, arXiv 2605.27760, May 2026
  5. [5]Zeng, Huang, He, LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth, arXiv 2602.07962, Feb. 2026
  6. [6]Anthropic, Agent Skills — overview (progressive disclosure)
  7. [7]Tang et al., How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions, arXiv 2605.29442, 2026
  8. [8]Gloaguen, Mündler, Müller, Raychev, Vechev (ETH Zurich), Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?, arXiv 2602.11988, 2026

Figures are quoted from each paper's abstract. SkillsBench figures refer to its first version (86 tasks, 11 domains); a later version reports +16.6 points over 87 tasks.