How do we measure which tool coding agents choose?
NativeFold runs real AI coding agents against real repositories and records what they actually do.
- Agent-Native Market Share reports one thing publicly: which option each agent selects and wires into a project.
- The Fold Score is our private instrument: where a single vendor loses runs along the way.
Both rest on the same evidence rules and the same bar a study has to clear before any number is allowed to leave the building.
The market measured here is agent tool-selection events under a published prompt set. Share is the fraction of scored runs each option occupies.
Browser share, operating system share and search share are all usage-measured rather than revenue-measured. This is the same kind of number, for a market that did not exist three years ago.
What this page tells you, and what it deliberately does not
Published here
- What each number means and what it licenses you to say.
- That verdicts come from observed behaviour, not from what an agent claims.
- The stages where an agent can drop a tool, so you can see selection is one moment in a longer path.
- How we report uncertainty, and the claims we refuse to make with the data we have.
- The conditions a study must satisfy before a result is publishable.
Not published here
- Harness internals and environment configuration.
- Classifier rules and internal vocabulary.
- Per-stage failure detail and vendor-specific findings.
- Anything describing the counterfactual intervention mechanism.
- Raw transcripts and run-level traces.
The line is not arbitrary. A reader should be able to judge whether our numbers deserve belief, and challenge them on their merits, without this page functioning as a build guide. Published prompts, corpus hashes and stated intervals are what make the results auditable. The instrument that produced them is the company.
The measurement spine
Everything in this section applies identically to Agent-Native Market Share and to the Fold Score. The two instruments differ in what they look at, not in how evidence is admitted or how uncertainty is handled.
1.1 Revealed behaviour, not stated preference
We do not ask a model which tool it would use. We give an agent an unbranded implementation task in a real repository and watch what it installs, imports, configures and calls. When an agent invents a call to a method that does not exist, that invented call is the most honest available statement of what the agent expected to find.
1.2 The cell is the unit
A cell is one scenario, run by one agent, in one edition. Every number lives inside its cell. We do not average across scenarios, we do not pool agents, and we do not publish a category-wide blended figure. Share depends on where you stand, and a blend erases exactly the signal that matters.
1.3 Controlled, reproducible environments
Runs execute in controlled, reproducible environments. The environment is measured rather than assumed, because a meaningful share of how an agent behaves depends on it.
1.4 Evidence, not narration
A verdict is never taken from what the agent says it did. It rests on what actually happened in the run. Verdicts are PASS, FAIL or AMBIGUOUS. Ambiguous runs are set aside and counted; totals are recomputed on the scored set so exclusions cannot hide.
1.5 One setup per corpus
A corpus is collected under a single held-constant setup, and the model each run actually resolved to is recorded. A setup change mid-corpus is segmented and reported, never silently pooled: a time series over unstated model versions is not a time series.
1.6 Frozen corpora, pre-registered claims
Scenarios, prompts and the analysis plan are fixed and hashed before collection starts. Corpora are read-only once banked. Exploratory work is stamped as exploratory and can never be relabelled into a confirmatory result.
The seven gates
Adoption is not one event. An agent has to notice it needs a tool, choose one, get it installed and authenticated, make it work, come back to it, and leave the choice behind for the next agent. Each step is a fork where a vendor can lose the run, and each implies a different fix. We measure all seven; we publish what happens at only one of them.
Does the agent register that the task needs an external capability at all?
Does the agent look anything up before it decides, or choose from memory?
Which option does the agent commit to and wire into the project?
Does the tool end up genuinely installed, and can the agent place the credentials it needs?
Does a real request actually leave the machine and succeed?
Given a second, related task, does the agent return to the same tool or switch away?
Does the choice persist into something the next agent will read as context?
Selection alone produces Agent-Native Market Share. The public benchmark reports one gate: which tool an agent chooses and wires. Everything from install onward is where the private instrument lives, because that is where a vendor learns not just whether it was picked, but whether the choice survived contact with a real project.
How we handle the fact that agents are random
The same agent, the same prompt and the same repository do not produce the same run. Run-to-run variation is real, it persists even at the lowest randomness setting, and it is the single largest threat to any benchmark of this kind. Every rule below exists to stop noise being reported as a finding.
Intervals, always
Every share is reported with a two-sided 95 percent Wilson interval. A bare point estimate is never published. One run is a rumour.
No ranking inside overlap
When two intervals overlap, we report them as not separated at this sample size. We do not rank them, order them visually, or describe one as ahead. A ranking our data cannot support is a fabrication with a chart around it.
The governance floor
Options below a small share threshold are reported as detected but precision-limited. They appear, they are named, and no claim is built on their exact value.
Change requires a guard
An edition-over-edition movement is only described as a change when the intervals do not overlap. Otherwise the movement is shown and explicitly labelled as within noise. Directional language is gated on the same test.
Paired comparisons only
Two options measured on the same scenarios are compared as paired data. Treating them as independent samples throws away the pairing and manufactures separation that is not there.
Reliable by design: every result is published at a certified sample size
Reliability here means one thing: if we ran the same measurement again, we would get the same answer. A share is only worth publishing if it is reproducible, not a lucky draw.
A single run is a rumour. Coding agents are stochastic: the same agent, on the same prompt in the same repository, does not always make the same choice, and that variation persists even at the lowest randomness setting. Run a scenario once and one tool wins, which tells you almost nothing, because the next run might go the other way. So we run every scenario many times, and we do not pick that number by budget. We set it with a proven internal procedure that confirms the answer has stabilised and would hold on a fresh set of runs before any number leaves the building.
What we publish is the stable share, the interval around it, and where relevant the total number of runs behind a study. The per-scenario count that procedure certifies stays internal.
Classification
Each run is scored from observed behaviour into exactly one outcome: a named tool, a hand-rolled implementation (the agent built the capability itself), or no integration (the agent produced nothing that works). Outcomes that cannot be scored cleanly are set aside and counted, never guessed.
Non-vendor outcomes are first-class results, not residue. An agent choosing nobody is a market fact, and often the most commercially interesting one on the chart. That unclaimed territory has its own name below.
Wording is accounted for
A prompt is one wording, and wording can move outcomes. We account for that in how studies are designed and in how strongly a finding is allowed to be stated. A result that would not survive a reasonable rewording is not presented as if it would.
Sensitivity panel
Leading option, three wordings, same cell.
Spread across three wordings: 8 points. Each result sits inside the others' intervals, so the finding survives rewording and can be stated at full strength.
One spine, two instruments
Everything above is shared. From here the instruments answer different questions for different readers. The public one answers who. The private one answers why. The public map is complete on what it covers and silent on the rest, by design.
Agent-Native Market Share
The share of scored runs in which each option is selected and wired, per scenario, per agent, per edition.
- Selection only. Which option the agent commits to and wires, nothing downstream of that choice.
- Agents measured as separate populations, equal sample per agent, never merged into a single figure.
- Nothing is pooled. Every time series lives inside its own scenario card. There is no category-wide number anywhere on the page.
- Unclaimed share is share. Hand-rolled and no-integration always occupy their buckets, always carry their sub-labels, and are rendered in hatched amber so the eye reads them as unclaimed rather than as a vendor.
- Phantom Demand. Runs where the task required the capability and no vendor got picked. The agent hand-rolled a solution, wired a placeholder, or shipped without it. The need was there; the selection never happened. It is charted as a first-class figure, never blended into a vendor bar.
- The verbatim prompt sits on every chart, with its interval, and a short signed analyst note written by the founder.
- It never states that one option is better than another. It states what agents chose.
The Fold Score
Our private, per-engagement instrument. It shows a vendor where along the adoption path agents drop their product, and which changes move adoption.
The method and the per-stage diagnostics are shared under engagement, not published here.
- A per-vendor score for a vendor that has not consented is never published. A detailed scored claim about a named company without its agreement is the most attackable object we could host, and we do not host it.
- Engagement data never touches the public index. The index is self-funded or it is nothing.
What a study has to satisfy before any number ships
These are conditions, not aspirations. A study that fails one does not get published with a caveat; it does not get published.
Scenarios, prompts, buckets and the analysis plan are fixed, hashed and signed before collection, including the rule that would kill the study.
A partial cell is unusable, not a smaller cell. Collection that cannot finish a cell does not start it.
Ambiguous and failed runs are counted and published as counts. Totals are recomputed on the scored set. Hidden exclusions are the first thing a hostile reader looks for, and they should find nothing.
Corpus and prompt hashes published. Corpora are read-only after banking. Every deviation from the plan is logged with its reason at the time it happens.
No ranking inside interval overlap, no directional language without the change guard, no assertion below the noise floor.
No element of the public index is requested, funded, previewed or timed by any vendor. The index carries no advertising and no paid placement.
Every published card and every analyst note is founder-signed before it ships. Automation renders; it never interprets.
Example outputs
Layouts only. Every number below is invented to show the shape of the artefact. Real cards are published in editions, each with its verbatim prompt, its interval and its signed note.
ANMS · scenario card
Signup confirmation email · Next.js · EU data residency
Claude Code · Edition 00X
Signed note. Vendor A leads in this slice. Vendor B and hand-rolled are not separated at this sample size and are reported as tied. Nearly a quarter of runs end with no supplier at all.
ANMS · saturated cell
Password reset email · Django · no framework hint
Claude Code · Edition 00X
Signed note. No vendor holds this stack. Every scored run built the capability by hand. This is unclaimed territory in its purest form.
Fold Score · gate ladder for one vendor, one slice
Share of runs surviving each gate. The drop between two bars is where the vendor loses the customer.
Each rate carries an interval in the delivered report. The largest single drop names the gate to fix first. The per-stage detail underneath each drop is part of the engagement and is never published.
References
The design decisions above are not house preferences. Each traces to published work on the statistics of agent evaluation, or to the standard statistical literature the field itself relies on.
Cite this methodology
Want your category measured?
Get early accessNativeFold measures which developer tools AI coding agents actually select and wire. The benchmark is free and self-funded. No vendor pays for placement, preview or timing.
Methodology version 2. Amendments to this page are versioned and logged. Questions about the method, including challenges to a published result, are answered directly.