Methodology · version 2

How do we measure which tool coding agents choose?

NativeFold runs real AI coding agents against real repositories and records what they actually do.

Both rest on the same evidence rules and the same bar a study has to clear before any number is allowed to leave the building.

The market we measure

The market measured here is agent tool-selection events under a published prompt set. Share is the fraction of scored runs each option occupies.

Browser share, operating system share and search share are all usage-measured rather than revenue-measured. This is the same kind of number, for a market that did not exist three years ago.

00 · scope of disclosure

What this page tells you, and what it deliberately does not

Published here

  • What each number means and what it licenses you to say.
  • That verdicts come from observed behaviour, not from what an agent claims.
  • The stages where an agent can drop a tool, so you can see selection is one moment in a longer path.
  • How we report uncertainty, and the claims we refuse to make with the data we have.
  • The conditions a study must satisfy before a result is publishable.

Not published here

  • Harness internals and environment configuration.
  • Classifier rules and internal vocabulary.
  • Per-stage failure detail and vendor-specific findings.
  • Anything describing the counterfactual intervention mechanism.
  • Raw transcripts and run-level traces.

The line is not arbitrary. A reader should be able to judge whether our numbers deserve belief, and challenge them on their merits, without this page functioning as a build guide. Published prompts, corpus hashes and stated intervals are what make the results auditable. The instrument that produced them is the company.

01 · common to both instruments

The measurement spine

Everything in this section applies identically to Agent-Native Market Share and to the Fold Score. The two instruments differ in what they look at, not in how evidence is admitted or how uncertainty is handled.

1.1 Revealed behaviour, not stated preference

We do not ask a model which tool it would use. We give an agent an unbranded implementation task in a real repository and watch what it installs, imports, configures and calls. When an agent invents a call to a method that does not exist, that invented call is the most honest available statement of what the agent expected to find.

1.2 The cell is the unit

A cell is one scenario, run by one agent, in one edition. Every number lives inside its cell. We do not average across scenarios, we do not pool agents, and we do not publish a category-wide blended figure. Share depends on where you stand, and a blend erases exactly the signal that matters.

1.3 Controlled, reproducible environments

Runs execute in controlled, reproducible environments. The environment is measured rather than assumed, because a meaningful share of how an agent behaves depends on it.

1.4 Evidence, not narration

A verdict is never taken from what the agent says it did. It rests on what actually happened in the run. Verdicts are PASS, FAIL or AMBIGUOUS. Ambiguous runs are set aside and counted; totals are recomputed on the scored set so exclusions cannot hide.

1.5 One setup per corpus

A corpus is collected under a single held-constant setup, and the model each run actually resolved to is recorded. A setup change mid-corpus is segmented and reported, never silently pooled: a time series over unstated model versions is not a time series.

1.6 Frozen corpora, pre-registered claims

Scenarios, prompts and the analysis plan are fixed and hashed before collection starts. Corpora are read-only once banked. Exploratory work is stamped as exploratory and can never be relabelled into a confirmatory result.

02 · common · the observation surface

The seven gates

Adoption is not one event. An agent has to notice it needs a tool, choose one, get it installed and authenticated, make it work, come back to it, and leave the choice behind for the next agent. Each step is a fork where a vendor can lose the run, and each implies a different fix. We measure all seven; we publish what happens at only one of them.

01
Recognition

Does the agent register that the task needs an external capability at all?

02
Discovery

Does the agent look anything up before it decides, or choose from memory?

03
Selection Public ANMS

Which option does the agent commit to and wire into the project?

04
Install and authenticate

Does the tool end up genuinely installed, and can the agent place the credentials it needs?

05
First call

Does a real request actually leave the machine and succeed?

06
Repeat use

Given a second, related task, does the agent return to the same tool or switch away?

07
Recommendation

Does the choice persist into something the next agent will read as context?

What this means for the benchmark

Selection alone produces Agent-Native Market Share. The public benchmark reports one gate: which tool an agent chooses and wires. Everything from install onward is where the private instrument lives, because that is where a vendor learns not just whether it was picked, but whether the choice survived contact with a real project.

03 · common · uncertainty

How we handle the fact that agents are random

The same agent, the same prompt and the same repository do not produce the same run. Run-to-run variation is real, it persists even at the lowest randomness setting, and it is the single largest threat to any benchmark of this kind. Every rule below exists to stop noise being reported as a finding.

Intervals, always

Every share is reported with a two-sided 95 percent Wilson interval. A bare point estimate is never published. One run is a rumour.

No ranking inside overlap

When two intervals overlap, we report them as not separated at this sample size. We do not rank them, order them visually, or describe one as ahead. A ranking our data cannot support is a fabrication with a chart around it.

The governance floor

Options below a small share threshold are reported as detected but precision-limited. They appear, they are named, and no claim is built on their exact value.

Change requires a guard

An edition-over-edition movement is only described as a change when the intervals do not overlap. Otherwise the movement is shown and explicitly labelled as within noise. Directional language is gated on the same test.

Paired comparisons only

Two options measured on the same scenarios are compared as paired data. Treating them as independent samples throws away the pairing and manufactures separation that is not there.

04 · common · sample size

Reliable by design: every result is published at a certified sample size

Reliability here means one thing: if we ran the same measurement again, we would get the same answer. A share is only worth publishing if it is reproducible, not a lucky draw.

A single run is a rumour. Coding agents are stochastic: the same agent, on the same prompt in the same repository, does not always make the same choice, and that variation persists even at the lowest randomness setting. Run a scenario once and one tool wins, which tells you almost nothing, because the next run might go the other way. So we run every scenario many times, and we do not pick that number by budget. We set it with a proven internal procedure that confirms the answer has stabilised and would hold on a fresh set of runs before any number leaves the building.

What we publish is the stable share, the interval around it, and where relevant the total number of runs behind a study. The per-scenario count that procedure certifies stays internal.

05 · common · turning a run into a number

Classification

Each run is scored from observed behaviour into exactly one outcome: a named tool, a hand-rolled implementation (the agent built the capability itself), or no integration (the agent produced nothing that works). Outcomes that cannot be scored cleanly are set aside and counted, never guessed.

Non-vendor outcomes are first-class results, not residue. An agent choosing nobody is a market fact, and often the most commercially interesting one on the chart. That unclaimed territory has its own name below.

06 · common · robustness

Wording is accounted for

A prompt is one wording, and wording can move outcomes. We account for that in how studies are designed and in how strongly a finding is allowed to be stated. A result that would not survive a reasonable rewording is not presented as if it would.

Illustrative layout · not data

Sensitivity panel

Leading option, three wordings, same cell.

Wording 150%
Wording 254%
Wording 346%

Spread across three wordings: 8 points. Each result sits inside the others' intervals, so the finding survives rewording and can be stated at full strength.

07 · where the two instruments diverge

One spine, two instruments

Everything above is shared. From here the instruments answer different questions for different readers. The public one answers who. The private one answers why. The public map is complete on what it covers and silent on the rest, by design.

Public · free · monthly editions

Agent-Native Market Share

The share of scored runs in which each option is selected and wired, per scenario, per agent, per edition.

  • Selection only. Which option the agent commits to and wires, nothing downstream of that choice.
  • Agents measured as separate populations, equal sample per agent, never merged into a single figure.
  • Nothing is pooled. Every time series lives inside its own scenario card. There is no category-wide number anywhere on the page.
  • Unclaimed share is share. Hand-rolled and no-integration always occupy their buckets, always carry their sub-labels, and are rendered in hatched amber so the eye reads them as unclaimed rather than as a vendor.
  • Phantom Demand. Runs where the task required the capability and no vendor got picked. The agent hand-rolled a solution, wired a placeholder, or shipped without it. The need was there; the selection never happened. It is charted as a first-class figure, never blended into a vendor bar.
  • The verbatim prompt sits on every chart, with its interval, and a short signed analyst note written by the founder.
  • It never states that one option is better than another. It states what agents chose.
Private · per engagement · vendor-commissioned

The Fold Score

Our private, per-engagement instrument. It shows a vendor where along the adoption path agents drop their product, and which changes move adoption.

The method and the per-stage diagnostics are shared under engagement, not published here.

  • A per-vendor score for a vendor that has not consented is never published. A detailed scored claim about a named company without its agreement is the most attackable object we could host, and we do not host it.
  • Engagement data never touches the public index. The index is self-funded or it is nothing.
08 · common · the publishing bar

What a study has to satisfy before any number ships

These are conditions, not aspirations. A study that fails one does not get published with a caveat; it does not get published.

Pre-registration

Scenarios, prompts, buckets and the analysis plan are fixed, hashed and signed before collection, including the rule that would kill the study.

Complete cells

A partial cell is unusable, not a smaller cell. Collection that cannot finish a cell does not start it.

Quarantine accounting

Ambiguous and failed runs are counted and published as counts. Totals are recomputed on the scored set. Hidden exclusions are the first thing a hostile reader looks for, and they should find nothing.

Frozen provenance

Corpus and prompt hashes published. Corpora are read-only after banking. Every deviation from the plan is logged with its reason at the time it happens.

Claim boundary

No ranking inside interval overlap, no directional language without the change guard, no assertion below the noise floor.

Independence

No element of the public index is requested, funded, previewed or timed by any vendor. The index carries no advertising and no paid placement.

Signed release

Every published card and every analyst note is founder-signed before it ships. Automation renders; it never interprets.

09 · what the results look like

Example outputs

Layouts only. Every number below is invented to show the shape of the artefact. Real cards are published in editions, each with its verbatim prompt, its interval and its signed note.

Placeholder · not data

ANMS · scenario card

Signup confirmation email · Next.js · EU data residency

Claude Code · Edition 00X

Vendor A57%
Vendor B19%
Hand-rolled16%
No integration8%

Phantom Demand · 24% (hand-rolled 16, no integration 8)
Prompt · verbatim, shown on card
Share reported with a 95% interval

Signed note. Vendor A leads in this slice. Vendor B and hand-rolled are not separated at this sample size and are reported as tied. Nearly a quarter of runs end with no supplier at all.

Placeholder · not data

ANMS · saturated cell

Password reset email · Django · no framework hint

Claude Code · Edition 00X

Hand-rolled100%

Phantom Demand · 100%
Prompt · verbatim, shown on card

Signed note. No vendor holds this stack. Every scored run built the capability by hand. This is unclaimed territory in its purest form.

Placeholder · not data · private deliverable

Fold Score · gate ladder for one vendor, one slice

Share of runs surviving each gate. The drop between two bars is where the vendor loses the customer.

Recognition98%
Discovery11%
Selection57%
Install / auth41%
First call22%
Repeat use14%
Recommendation9%

Each rate carries an interval in the delivered report. The largest single drop names the gate to fix first. The per-stage detail underneath each drop is part of the engagement and is never published.

10 · what this is built on

References

The design decisions above are not house preferences. Each traces to published work on the statistics of agent evaluation, or to the standard statistical literature the field itself relies on.

Bjarnason et al. On randomness in agentic evaluation. arXiv:2602.07150. Why every scenario is run many times and every figure carries an interval. Run-to-run variance persists at temperature zero, and single-run evaluation of a long trajectory is close to meaningless. Also the basis for reporting optimistic and pessimistic success across runs rather than a single rate.
Mustahsan et al. Stochasticity in agentic evaluations. The reliability framing: separating variation caused by genuinely different scenarios from variation caused by agent flakiness, the scenario counts required before that estimate is stable, and the use of paired tests when comparing two options on the same scenarios.
Rabanser et al. Towards a science of AI agent reliability. arXiv:2602.16666. Consistency, predictability and robustness as first-class dimensions alongside capability, and the prompt-perturbation protocol that underpins our holdout paraphrase design.
Kapoor et al. AI agents that matter. arXiv:2407.01502. Benchmarks drift toward measuring what is easy rather than what matters. The source of our insistence on real repositories, held-out prompts, published costs and pre-registered verdict rules.
Wilson, E. B. (1927). Probable inference, the law of succession, and statistical inference. JASA 22(158). The interval we use everywhere. Chosen over the normal approximation because it stays well behaved at shares near zero and one, which is exactly where our most interesting cells sit.
Shrout, P. E. and Fleiss, J. L. (1979). Intraclass correlations. Psychological Bulletin 86(2). The reliability coefficient and the conditions under which it is estimable. The reason we withhold reliability figures below a minimum scenario count rather than quoting an unstable one.
McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions. Psychometrika 12(2). The correct test for two options measured on the same scenarios. Treating paired data as independent samples manufactures separation that is not there.
Efron, B. (1979). Bootstrap methods: another look at the jackknife. Annals of Statistics 7(1). Resampling underpins the certified-sample-size procedure and the intervals on differences, without assuming a shape the data does not have.
11 · cite this methodology

Cite this methodology

NativeFold (2026). Methodology: Agent-Native Market Share and the Fold Score. https://nativefold.com/methodology

Want your category measured?

Get early access

NativeFold measures which developer tools AI coding agents actually select and wire. The benchmark is free and self-funded. No vendor pays for placement, preview or timing.

Methodology version 2. Amendments to this page are versioned and logged. Questions about the method, including challenges to a published result, are answered directly.