The 42% vs 78% Result Is Really About the Harness
A HAL result makes the point quickly: on CORE-Bench, Claude Opus 4.5 scored 42% with CORE-Agent and 78% with Claude Code before later grading corrections. Same model, very different result.
The takeaway is not that AutoHarness caused this gap—it didn’t. It is that the system around a model—its prompts, tools, skills, and workflow rules—can matter as much as the model itself.
But that layer is still mostly maintained by hand. AutoHarness asks a simple question: what if the skill layer could learn from the work Claude Code is already doing?
What AutoHarness Actually Does
Four Ideas That Make It Interesting
1. It learns from real work
AutoHarness does not need a separate training dataset. When a session contains a useful correction, workaround, setup step, or repository-specific technique, the Reflector can turn that experience into a reusable skill.
For example, if a deployment fails because a production manifest must be regenerated first, the useful lesson is not “Claude made a mistake today.” It is a reusable rule for that repository’s deployment workflow.
2. It consolidates instead of piling up
A self-learning system becomes less useful if every session creates another near-duplicate skill. AutoHarness compares new lessons with what already exists. When possible, it patches or broadens an existing skill rather than creating a new one.
This gives the skill layer a lifecycle: learn, merge, update, mature, and eventually archive.
3. Skills survive through actual use
AutoHarness does not require an offline benchmark for every learned skill. Instead, it tracks adherence: whether a skill is actually recalled or consumed in later work. New skills begin in probation. Skills that never get used can be archived; mature skills compete for limited active capacity.
That is a useful online signal, but it should not be overstated. Usage proves that a skill entered the workflow; it does not prove that the skill caused a better task outcome.
4. It only changes what it created
This is one of the project’s most important boundaries. AutoHarness only manages skills marked as AutoHarness-authored. Your hand-written skills and installed skills stay outside its automated mutation boundary.
In other words, the system can evolve its own learned layer without becoming a general-purpose editor for all of your Claude Code configuration.
How the Learning Loop Works
The architecture separates learning from authority. The model can decide what may be worth learning, but it does not get unrestricted permission to write to the skill store.
| Component | Role |
|---|---|
| CAP · Capture | Collects the relevant session episode and redacts sensitive material before reflection. |
| REF · Reflect | Finds durable lessons, compares them with existing skills, and proposes a create, patch, update, or consolidation. |
| Promoter | The sole deterministic writer. It validates structure, ownership, paths, and provenance before anything lands. |
| MNG · Lifecycle | Tracks use, manages probation and capacity, and archives skills that no longer justify active recall. |
| LED · Ledger | Keeps an append-only record of why a skill was created or changed and what evidence supported it. |
The key design rule is simple: model proposes; deterministic code validates and writes. That boundary makes the system easier to audit and reduces the risk of a free-form model silently rewriting the live skill store.
Why This Design Matters in Practice
The value of AutoHarness is not simply that it can create skills. The useful part is that learning, validation, storage, and retirement are treated as separate jobs. That makes the skill layer easier to inspect and less dependent on someone manually maintaining every instruction.
Because learned behavior stays in ordinary files, teams can review what was added, trace why it changed through the ledger, and restore an archived skill when needed. The system also stays off the host recall path: Claude Code continues to recall native skills normally while AutoHarness maintains the layer beside it.
In practice, that turns AutoHarness from a skill generator into a maintenance system for procedural knowledge: learn what matters, keep what proves useful in real work, and retire what no longer earns attention.
Where AutoHarness Fits
There are several ways to maintain a self-learning skill layer. One option is to let it grow without bounds. Another is to gate every change with an offline benchmark. A third is to use a resident daemon and time-based cleanup.
AutoHarness takes a different route: it bounds the active skill layer, uses real-world adherence as its survival signal, requires no dedicated benchmark or oracle, and recomputes lifecycle state without a resident daemon.
That choice fits the kind of work Claude Code often does: open-ended, repository-specific tasks where a clean held-out benchmark may not exist. The trade-off is deliberate—adherence tells the system what remains relevant in practice, while the ledger preserves evidence for deeper evaluation later.
The result is a lightweight learning layer designed to live inside everyday development rather than beside it as a separate evaluation system.
The Bigger Idea
Today, Claude Code can help solve a difficult repository-specific problem and still lose much of that operational lesson when the session ends. AutoHarness asks whether those lessons can compound.
Not by training a new model. Not by building a benchmark for every private workflow. But by turning real corrections and successful procedures into inspectable skills, merging overlapping knowledge, protecting human-authored instructions, tracking provenance, and retiring what no longer gets used.
The model can stay the same. The experience around it does not have to.
Getting Started
AutoHarness requires python3 and installs through Claude Code’s plugin system. In Claude Code, run:
/plugin marketplace add tigerless-labs/autoharness/plugin install autoharness@autoharness/reload-pluginsAfter installation, keep working normally. Project-level learned skills appear under ., while general skills can live in ~/..


