The monthly benchmark of the developer tools AI coding agents choose.
The market measured here is agent tool-selection events under a published prompt set. Share is the fraction of scored runs each option occupies.
This service runs on AWS Lambda. When a new user signs up, send them a welcome email so a real message would go out from the function. Wire it end to end.
prompt committed at: 2026-07-18T01:30:53Z
On an AWS Lambda backend, Claude Code reached for AWS SES on every run and wired no third-party provider. This is a control: inside the platform, the platform's own email service owns the stack completely. The interval sits just short of total because a unanimous result still admits a little uncertainty.
On an AWS Lambda backend, Codex selects AWS SES almost universally, the same outcome Claude Code reaches on this scenario. The only behavioral difference between the two agents is which AWS SDK version they reach for, a sub-finding about tooling rather than a split in which provider is chosen. An earlier version of the measurement instrument under-counted this cell; the error was found and corrected before publication, and AWS SES resolves as the near-unanimous selection either way.
First edition for this prompt. A time series starts at the next edition.
Mentioned means the agent said the tool's name. Installed means it actually wrote code that uses it. A tool can be named and still never used.
Alternatives were named only in the closing message after the integration was written, as an offer to swap.
These counts cover the tools we track. A tool outside that list could have been named without being counted here, so treat each rate as a floor.
When a user signs up, send them a welcome email. For data-residency compliance we need the email provider to store data in the EU. Wire it end to end.
prompt committed at: 2026-07-19T07:42:17Z
On a Next.js backend with an EU data-residency requirement, Mailgun leads Claude Code's selections, with several other providers present. Like the Codex result on this scenario it is a genuinely contested cell, and the specific vendors present and their relative shares differ between the two agents.
On a Next.js backend with an EU data-residency requirement, Mailgun leads Codex's selections, with three other providers present in meaningful share. This is the page's most competitive cell: several vendors are genuinely in contention rather than one default dominating.
First edition for this prompt. A time series starts at the next edition.
Mentioned means the agent said the tool's name. Installed means it actually wrote code that uses it. A tool can be named and still never used.
Alternatives were named only in the closing message after the integration was written, as an offer to swap.
These counts cover the tools we track. A tool outside that list could have been named without being counted here, so treat each rate as a floor.
When a new user signs up, send them a welcome email confirming their account is ready. Wire it up end to end so a real message would send, and note in the README how to set any credentials.
prompt committed at: 2026-07-19T13:03:51Z
On the same Next.js scenario, Claude Code selects Resend in a clear majority of runs, and the unclaimed share here is small. Resend leads with a majority. The provider-agnostic, deferred pattern that dominates the Codex card appears here too but only faintly.
On a Next.js backend, Resend is the most selected provider and is clearly present in Codex's choices. The striking figure is the unclaimed share: 0.427 of runs wired a complete SMTP send and pointed it at no provider, deferring the vendor choice to deployment configuration. That is demand that is fully built and entirely unclaimed, a direct opening for whichever provider an operator reaches for first. Reported as present, not as a majority: at this sample the lower bound does not clear half.
First edition for this prompt. A time series starts at the next edition.
Mentioned means the agent said the tool's name. Installed means it actually wrote code that uses it. A tool can be named and still never used.
Alternatives were named only in the closing message after the integration was written, as an offer to swap.
These counts cover the tools we track. A tool outside that list could have been named without being counted here, so treat each rate as a floor.
When a new user signs up, send them a welcome email confirming their account is ready. Wire the send end to end and note in the README how to set any credentials.
prompt committed at: 2026-07-18T19:08:45Z
On a Django backend, every run wired the framework's mail path and then configured it to print to standard output rather than send. No message was ever transmitted and no provider was ever considered. The cell is entirely unclaimed demand: the integration point exists and works, and the vendor slot is empty. It is the most winnable cell on the page.
On a Django backend, Codex shows the same pattern as Claude Code: every scored run wired the framework mail path to a non-transmitting backend that prints rather than sends, reaching no provider. This cell was collected only partially and is held. It carries an observation, not a certified result, and no claim is made from it pending a complete collection.
First edition for this prompt. A time series starts at the next edition.
In some runs the task required email and no vendor got picked. The agent hand-rolled SMTP, wired a placeholder and stopped, or shipped without the integration. The need was there; the selection never happened. It renders on every chart below as hatched share, and it is winnable.
Full methodology, prompt corpus, and hashes: the methodology page
Agents run unpinned, as users run them. The model that actually served each run is recorded (resolved_model) and published. When a model release moves share, the move is the finding.
Every prompt is published verbatim. Holdout paraphrases are committed by hash before each edition and disclosed after it, so the target cannot be optimized against silently.
Every bucket carries a 95% Wilson interval at its published n. Runs are scored from the workspace diff, not from what the agent said it did.
Shares are per prompt. No blended category percentage exists on any surface, because pooling prompts manufactures a market that no one prompt expressed.
Between editions, movement is reported as a Shift only when intervals do not overlap. Everything else is noise and is labelled as such.
Hand-rolled and no-integration outcomes are share, sub-labelled by evidence in the diff (deliberate, placeholder, local). They are never collapsed into a footnote.
Monthly measurement on prompts you choose, with an alert when your share leaves its interval. Model releases reshuffle defaults overnight; the public edition will tell you a month later.
Monitor your positionA category diagnostic on your vertical: where you lose, gate by gate, from recognition to selection to install to first call, on prompts you choose. The chart shows the gap; the diagnostic shows the route.
Request a diagnosticStripe, Adyen, Braintree, LemonSqueezy, Paddle, hand-rolled? Request the benchmark and we'll email you the day it ships. The category with the most requests is published first.
Clerk, Auth0, Supabase Auth, NextAuth, hand-rolled sessions? Request the benchmark and we'll email you the day it ships. The category with the most requests is published first.
Inngest, Trigger.dev, BullMQ, Sidekiq, Temporal, hand-rolled cron? Request the benchmark and we'll email you the day it ships. The category with the most requests is published first.
Algolia, Typesense, Meilisearch, Elasticsearch, Postgres full-text, hand-rolled LIKE queries? Request the benchmark and we'll email you the day it ships. The category with the most requests is published first.
Twilio, Vonage, MessageBird, Plivo, Telnyx, a raw carrier gateway? Request the benchmark and we'll email you the day it ships. The category with the most requests is published first.
Sentry, Rollbar, Bugsnag, Honeybadger, Datadog, a hand-rolled logger? Request the benchmark and we'll email you the day it ships. The category with the most requests is published first.
LaunchDarkly, Statsig, Unleash, Flagsmith, PostHog, hand-rolled env vars? Request the benchmark and we'll email you the day it ships. The category with the most requests is published first.
S3, Cloudflare R2, Supabase Storage, UploadThing, Cloudinary, the local disk? Request the benchmark and we'll email you the day it ships. The category with the most requests is published first.
Name the category and the tools you want tracked. The category with the most requests is published first.
NativeFold measures which developer tools AI coding agents actually select, install and call. The benchmark is free and self-funded. No vendor pays for placement, preview or timing.