Results
What the benchmark measured
Two agent benchmarks, two AI coding agents, 6 search providers and 609 graded test runs, plus a replication of the search gap bench on Claude Code and a study that called four search APIs directly. Every number on this page is generated from the published report data.
Read this before comparing providers
- Small samples18 to 42 test runs per provider. The 95% ranges overlap widely, so neighbouring positions are not rankings, and no provider beat answering from memory on the documented-knowledge benchmark by a statistically clear margin.
- One connection per providerIn the two agent benchmarks, each provider ran through its own official MCP server at a pinned version and default settings; other search modes and tools were not tested there. The search API head-to-head called the APIs directly with tuned parameters, and its results do not transfer to agents, or back.
- The agent mattersRankings changed between Claude Code and Codex. A result on one agent does not transfer to another.
- Built-in search rerunIn the original runs Claude Code's built-in search could not open pages. After the fix it was rerun on 2026-10-06, and its rows here come from that rerun. The original rows stay on record in the reports.
- Cost coverageToken counts are reported; provider dollar spend was only partly metered, so dollar comparisons are omitted.
- Who gradedAutomatic checks where possible; otherwise model graders that compared answers with the captured sources without seeing which provider produced them.
Results at a glance
reports/*/summary.json.Leaderboards
Ordered by observed pass rate. The bar is the 95% range; the tick is the observed rate. Machine-readable data: leaderboard.json.
Search gap bench: job outcomes on Claude Code (claude-opus-5-5)
7 admitted post-cutoff brief tasks × 3 repetitions = 21 runs per setup. Pass = right decision, weighted key-fact recall ≥ 0.7, unsupported claims ≤ 0.25. Interval: Wilson 95% (published). The dashed line marks the lower end of the top setup's range; badges describe overlap of individual ranges only. Overlap does not establish statistical equivalence or test a difference between setups.
| Search setup | Passed | 95% range on a 0–100% scale | Gap closure | Tokens per success | Notes |
|---|---|---|---|---|---|
| Parallel | 20/21 · 95% 77–99% | 1.00 (0.86–1.17) | 64k | overlaps top interval | |
| Firecrawl | 19/21 · 90% 71–97% | 0.95 (0.83–1.00) | 116k | overlaps top interval | |
| Brave | 17/21 · 81% 60–92% | 0.85 (0.62–1.00) | 93k | overlaps top interval | |
| Built-in | 16/21 · 76% 55–89% | 0.80 (0.53–1.00) | 61k | overlaps top interval Rerun 2026-10-06 with page opening fixed | |
| Perplexity | 16/21 · 76% 55–89% | 0.80 (0.50–1.00) | 74k | overlaps top interval | |
| Exa | 15/21 · 71% 50–86% | 0.75 (0.47–0.95) | 89k | overlaps top interval | |
| Tavily | 15/21 · 71% 50–86% | 0.75 (0.43–1.00) | 83k | overlaps top interval |
floor (no search): 0% · ceiling (answer excerpt): 95%
GAP tokens per success counts all agent tokens in the setup, failed runs included, divided by passing runs.
Source: Search gap bench.
Search gap bench: job outcomes on codex (gpt-6.1-sol)
6 admitted post-cutoff brief tasks × 3 repetitions = 18 runs per setup. Pass = right decision, weighted key-fact recall ≥ 0.7, unsupported claims ≤ 0.25. Interval: Wilson 95% (published). The dashed line marks the lower end of the top setup's range; badges describe overlap of individual ranges only. Overlap does not establish statistical equivalence or test a difference between setups.
| Search setup | Passed | 95% range on a 0–100% scale | Gap closure | Tokens per success | Notes |
|---|---|---|---|---|---|
| Brave | 17/18 · 94% 74–99% | 1.06 (0.88–1.29) | 111k | overlaps top interval | |
| Firecrawl | 16/18 · 89% 67–97% | 1.00 (1.00–1.00) | 166k | overlaps top interval | |
| Tavily | 16/18 · 89% 67–97% | 1.00 (0.71–1.29) | 129k | overlaps top interval | |
| Built-in | 14/18 · 78% 55–91% | 0.88 (0.47–1.20) | 117k | overlaps top interval | |
| Exa | 14/18 · 78% 55–91% | 0.88 (0.47–1.29) | 146k | overlaps top interval | |
| Parallel | 13/18 · 72% 49–88% | 0.81 (0.41–1.14) | 157k | overlaps top interval | |
| Perplexity | 13/18 · 72% 49–88% | 0.81 (0.47–1.12) | 117k | overlaps top interval |
floor (no search): 0% (0–18%) · ceiling (answer excerpt): 89% (67–97%)
GAP tokens per success counts all agent tokens in the setup, failed runs included, divided by passing runs.
Source: Search gap bench.
Web search bakeoff: competitive research tasks on Claude Code (claude-opus-5-5)
14 competitive tasks × 3 repetitions = 42 runs per setup, regraded 2026-09-30. Mostly stable, documented knowledge; search adds less here than on post-cutoff jobs. Interval: Wilson 95% (computed from the published counts). The dashed line marks the lower end of the top setup's range; badges describe overlap of individual ranges only. Overlap does not establish statistical equivalence or test a difference between setups.
| Search setup | Passed | 95% range on a 0–100% scale | Tokens per success | Notes |
|---|---|---|---|---|
| Firecrawl | 38/42 · 90% 78–96% | 61.5k | overlaps top interval | |
| Perplexity | 37/42 · 88% 75–95% | 47.7k | overlaps top interval | |
| Tavily | 36/42 · 86% 72–93% | 61.3k | overlaps top interval | |
| Built-in | 35/42 · 83% 69–92% | 39.7k | overlaps top interval Rerun 2026-10-06 with page opening fixed | |
| Exa | 35/42 · 83% 69–92% | 65.0k | overlaps top interval | |
| Parallel | 33/42 · 79% 64–88% | 69.8k | overlaps top interval | |
| No search | 32/42 · 76% 61–87% | 20.9k | reference | |
| Brave | 31/42 · 74% 59–85% | 89.2k | overlaps top interval |
WSB tokens per success uses measured tokens from graded runs divided by measured successes. Ungraded and unmeasured runs are excluded; per-setup coverage counts were not retained, so complete accounting cannot be established.
Source: Web search bakeoff.
Search APIs called directly
A separate study (Search API head-to-head, 2026-09-26 to 2026-10-06) sent the same pre-registered research questions to four search APIs directly, with each vendor's best-practice parameters, and had blind model judges verify every counted item on its source page. It measures what each API adds to a knowledge base already built with built-in search, not what an agent does with it.
| Measure | Exa | Tavily | Parallel | Firecrawl |
|---|---|---|---|---|
| New or correcting claims, 40 open questions | 25 | 11 | 19 | 10 |
| New dated facts, two weeks of company news | 20 | 12 | 3 | 2 |
| Pages other tools could not read, recovered (of 15) | 5 | 10 | 12 | 14 |
| Schema values right on ground truth (of 133) | 105 | — | 113 | 94 |
| Entities new to the knowledge base, three lists | 28 | — | 7 | 19 |
| New tier-1 claims on four big vendors (no lead significant) | 16 | 13 | 14 | 9 |
Counts of verified items from one run per question, transcribed from the report. No intervals were published; differences of a few items are within noise. A dash means not run: Tavily sells no list-building or structured-extraction product. An Exa-only Monitors probe made 49 daily runs on 5 competitors' pricing and changelog pages and reported 20 changes, 14 of them confirmed on the page; 3 of 5 monitors surfaced a real, dated pricing or product change.
Context: the study was run while building a go-to-market knowledge base about Exa, one of the four providers, and its questions come from that knowledge base. Full report: Search API head-to-head.
Full reports
The recorded reports, transcribed and checked by scripts/check_reports.py, with plain-language labels
for search setups, test runs and agents.
Web search bakeoff (2026-09-29)
Full recorded report
WSB live battery — initial results (2026-09-29)
Update (2026-10-06): the native setup was rerun with page fetching working. In the original battery the setup exposed WebSearch and WebFetch but pre-approved only WebSearch, and headless Claude Code refused WebFetch in 40 of 54 native runs (110 calls). The configuration is fixed and pinned by a test. All 54 native runs (18 tasks × 3) were rerun on 2026-10-06 with the same seed, per-task budgets and model, and graded against the historical regrade catalog by a blinded Claude judge; no call was refused. In the regrade table the native row now comes from the rerun: 35/42 competitive runs (83%) at 39.7k tokens per success, against 33/42 (79%) and 36.3k recorded originally. The rerun's 12 expected-fail runs passed 0/12. The initial-grading tables below are the original record and still carry the original native runs. Provider setups were unaffected. Details: agent search behavior.
Pack: web-search-bakeoff (WSB) on the Search Evaluation Workbench (SEW). Agent / model: claude-code, Claude Opus 5.5 (claude-opus-5-5[1m]), authenticated execution on the original host. Status: historical initial grading; headline and Results tables are superseded by the 2026-09-30 regrade, where Firecrawl leads at 90%. The numbers below are graded with SEWBENCH-01 applied to the stored deliverables. The same deliverables graded without those fixes are summarised in Failure dive and were materially wrong.
Headline
Historical figures under the initial grader: rankings and cost lower bounds here apply only to that grading, not the later operator-ruled catalog.
On the 14 competitive production tasks (42 runs per setup), every search setup except Brave scores above answering without search:
- Firecrawl and Perplexity lead at 83%. Exa is at 81%, Tavily 79%, native web search and Parallel 76%, no-search 71% and Brave 69%.
- The 95% intervals overlap. At n=42 per setup the ordering is directional, not conclusive: no setup's difference from no-search or from native has an interval that excludes zero.
- Search matters most on tasks that need live data. On the cloud egress pricing table only Firecrawl and Perplexity went 3/3 (Brave 2/3, the rest 0/3). The AWS pricing pages need JavaScript rendering, and most setups reported the intra-AZ run as unavailable, honestly.
- Among fully priced setups, native search costs least per success ($0.167), followed by Tavily ($0.206). Perplexity (at least $0.148; 37/42 runs priced) and Firecrawl (at least $0.159; 39/42 priced) are lower bounds, excluded from the cost ranking; Perplexity's
askspend is unmetered. Brave and no-search are also partially priced (at least $0.28 each); no-search had a few long runaway answers. - Brave is last, failing
unanswerable-nonexistent-postgres-gucandupstream-diagnosis-node-openssl3-md4in all three repetitions.
What ran
| Run | Setups | Runs | Window (UTC) |
|---|---|---|---|
| first setup group | no-search, native, brave, tavily | 216 | 2026-09-29 03:29 – 13:26 |
| second setup group | exa, parallel-web, firecrawl, perplexity | 216 | 2026-09-29 13:39 – 22:49 |
- Tasks: all 18 production tasks (production task catalog), × 3 repetitions.
- 14 are
competitive. - 4 are
expected_failby design: no retrieval surface can answer them, and the honest deliverable is a partial answer or a decline.
- 14 are
- Setups:
- Provider setups expose one vendor MCP server each, and nothing else.
- The native setup uses the agent's own web search.
- The no-search setup has no network tools.
- Pinned MCP servers:
- @brave/brave-search-mcp-server version 2.1.4
tavily-mcp@0.2.22exa-mcp-server@3.4.1firecrawl-mcp@3.26.0- @perplexity-ai/mcp-server version 1.3.0
- Parallel's hosted Search MCP (
https://search.parallel.ai/mcp-oauth, authenticated) viamcp-remote@0.14.3
- Grading:
- Deterministic validators where a task has one.
- Otherwise a blinded claude-code judge scoring the task's rubric.
bakeoff grade --regradere-scores the stored deliverables. Nothing was re-run.
- Cost: model tokens are priced at the published Opus 5.5 rates. Vendor spend is metered per MCP call where a tariff exists. It is unknown for Exa, Parallel and Perplexity's
asktool, which report no spend over MCP. In the tables,partialis the priced runs' spend divided by all of the setup's successes. Unpriced runs can only add spend, so eachpartialfigure is a lower bound on the full-setup $/success, and these setups are excluded from the cost ranking.bakeoff reportdivides by the successes among priced runs instead, which gives $0.172 for Perplexity, $0.163 for Firecrawl and $0.29 for Brave on the competitive set, and no figure for Exa or Parallel on the expected-fail set, where every success is in an unpriced run.
Results
All Results tables below use the historical initial grading. The ungraded column counts delivered answers that could not be graded. Runs that exhausted their budget before delivering are counted under budget-exhausted and as failures, rather than under ungraded.
Competitive tasks (14 tasks × 3 repetitions)
| setup | pass | 95% CI | $/success | priced | tokens/run | provider calls | budget-exhausted | ungraded |
|---|---|---|---|---|---|---|---|---|
| firecrawl | 35/42 (83%) | 69–92% | partial $0.159 (39/42 priced) | 39/42 | 55k | 219 | 2 | 0 |
| perplexity | 35/42 (83%) | 69–92% | partial $0.148 (37/42 priced) | 37/42 | 42k | 145 | 0 | 0 |
| exa | 34/42 (81%) | 67–90% | unpriced | 0/42 | 54k | 105 | 0 | 0 |
| tavily | 33/42 (79%) | 64–88% | $0.206 | 42/42 | 52k | 119 | 0 | 0 |
| native | 32/42 (76%) | 61–87% | $0.167 | 42/42 | 28k | 256 | 0 | 0 |
| parallel-web | 32/42 (76%) | 61–87% | unpriced | 0/42 | 54k | 90 | 1 | 0 |
| no-search | 30/42 (71%) | 56–83% | partial $0.285 (41/42 priced) | 41/42 | 15k | 0 | 0 | 0 |
| brave | 29/42 (69%) | 54–81% | partial $0.280 (41/42 priced) | 41/42 | 65k | 154 | 0 | 0 |
Competitive deltas against controls
Each setup against each control over the same 14 tasks × 3 repetitions. Intervals are 95% Newcombe (hybrid Wilson score) intervals for a difference of two independent proportions. They ignore the task pairing, so they are conservative. Token ratios compare mean tokens per run. Provider setups and their controls come from two runs about ten hours apart (see Caveats).
| setup | vs no-search: pass Δ (95% CI) | vs no-search: tokens | vs native: pass Δ (95% CI) | vs native: tokens |
|---|---|---|---|---|
| firecrawl | +12 pp (−6 to +29) | 3.49× | +7 pp (−10 to +24) | 1.95× |
| perplexity | +12 pp (−6 to +29) | 2.63× | +7 pp (−10 to +24) | 1.47× |
| exa | +10 pp (−9 to +27) | 3.39× | +5 pp (−13 to +22) | 1.90× |
| tavily | +7 pp (−11 to +25) | 3.29× | +2 pp (−15 to +20) | 1.84× |
| native | +5 pp (−14 to +23) | 1.79× | — | — |
| parallel-web | +5 pp (−14 to +23) | 3.44× | +0 pp (−18 to +18) | 1.92× |
| no-search | — | — | −5 pp (−23 to +14) | 0.56× |
| brave | −2 pp (−21 to +17) | 4.13× | −7 pp (−25 to +12) | 2.31× |
Expected-fail tasks (4 tasks × 3 repetitions)
Passing here means an honest partial answer or decline, which the validator can confirm.
| setup | pass | 95% CI | $/success | priced | tokens/run | provider calls | budget-exhausted | ungraded |
|---|---|---|---|---|---|---|---|---|
| exa | 5/12 (42%) | 19–68% | partial $0.046 (6/12 priced) | 6/12 | 67k | 30 | 0 | 0 |
| no-search | 3/12 (25%) | 9–53% | partial $2.827 (9/12 priced) | 9/12 | 43k | 0 | 0 | 0 |
| native | 3/12 (25%) | 9–53% | $0.273 | 12/12 | 16k | 28 | 0 | 0 |
| brave | 3/12 (25%) | 9–53% | partial $0.763 (11/12 priced) | 11/12 | 66k | 45 | 1 | 0 |
| firecrawl | 3/12 (25%) | 9–53% | partial $0.672 (10/12 priced) | 10/12 | 75k | 104 | 2 | 0 |
| perplexity | 3/12 (25%) | 9–53% | partial $0.197 (9/12 priced) | 9/12 | 36k | 30 | 0 | 0 |
| parallel-web | 2/12 (17%) | 5–45% | partial $0.116 (6/12 priced) | 6/12 | 60k | 34 | 1 | 0 |
| tavily | 1/12 (8%) | 1–35% | partial $1.923 (11/12 priced) | 11/12 | 64k | 49 | 1 | 0 |
Per task
| task | outcome | no-search | native | brave | tavily | exa | parallel-web | firecrawl | perplexity |
|---|---|---|---|---|---|---|---|---|---|
change-detection-kubernetes-dockershim-removal | competitive | 2/3 | 2/3 | 1/3 | 3/3 | 2/3 | 3/3 | 3/3 | 3/3 |
change-detection-openapi-30-to-31 | competitive | 1/3 | 1/3 | 0/3 | 0/3 | 2/3 | 2/3 | 0/3 | 1/3 |
competitive-table-cloud-egress-pricing | competitive | 0/3 | 0/3 | 2/3 | 0/3 | 0/3 | 0/3 | 3/3 | 3/3 |
competitive-table-copyleft-obligations | competitive | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
list-build-kubernetes-122-api-removals | competitive | 3/3 | 3/3 | 3/3 | 3/3 | 2/3 | 2/3 | 3/3 | 3/3 |
list-build-python313-pep594-removals | competitive | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
list-build-sca-tool-shortlist | competitive | 0/3 | 2/3 | 2/3 | 1/3 | 3/3 | 3/3 | 1/3 | 3/3 |
multi-hop-cve-to-fixed-release | competitive | 3/3 | 1/3 | 3/3 | 3/3 | 3/3 | 2/3 | 2/3 | 1/3 |
multi-hop-python-feature-peps | competitive | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
unanswerable-nonexistent-postgres-guc | competitive | 2/3 | 3/3 | 0/3 | 3/3 | 2/3 | 2/3 | 3/3 | 1/3 |
unanswerable-private-company-audited-arr | competitive | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 1/3 | 3/3 | 3/3 |
unanswerable-python4-ga-date | competitive | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
upstream-diagnosis-node-openssl3-md4 | competitive | 1/3 | 2/3 | 0/3 | 2/3 | 2/3 | 2/3 | 2/3 | 2/3 |
upstream-diagnosis-urllib3-libressl-import-error | competitive | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
change-detection-unversioned-runtime-drift | expected_fail | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 |
competitive-table-unpublished-enterprise-pricing | expected_fail | 3/3 | 3/3 | 3/3 | 1/3 | 2/3 | 1/3 | 2/3 | 3/3 |
multi-hop-npm-semver-dependents | expected_fail | 0/3 | 0/3 | 0/3 | 0/3 | 3/3 | 1/3 | 1/3 | 0/3 |
upstream-diagnosis-unindexed-ordering-regression | expected_fail | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 |
Findings
1. Search earns its cost on live or long-tail data, not on stable knowledge.
- Tasks whose answers are stable and well known were 3/3 in every setup, including no-search: copyleft obligations, PEP 594 removals, Python feature PEPs and the urllib3 diagnosis.
- The separation comes from the tasks that need current data:
- the egress pricing table;
- the SCA tool shortlist (no-search 0/3; Exa, Parallel and Perplexity 3/3);
- the dockershim change-detection task.
2. Rendering matters for vendor pricing pages. Firecrawl (which renders JavaScript) and Perplexity were the only setups to read AWS's intra-AZ transfer pricing. The others honestly reported it unavailable, and the validator fails an unavailable run. 3. Declining is not the default for search setups. unanswerable-nonexistent-postgres-guc separates the setups: Brave 0/3 and Perplexity 1/3 against 3/3 for native, Tavily and Firecrawl. 4. Exa handles impossible tasks best (5/12 honest partials on the expected-fail set). It is the only setup to go 3/3 on multi-hop-npm-semver-dependents, where the honest move is to state the method and its limits rather than invent a ranking. 5. Provider setups use more reported tokens per run. The MCP provider setups report 42k–65k tokens per run against 15k for no-search. Native search reports 28k. These totals do not separate input from output tokens, which have different prices, and some vendor spend is unknown; the totals alone cannot establish what dominates cost per success.
Failure dive
Two tasks were initially at 0/24 across all setups; a third mostly failed, and the judged SCA task was mostly ungraded. All four exposed bench bugs. Each was confirmed by reading the deliverables, and each is fixed in the bench fix:
| Task | As first graded | Graded with SEWBENCH-01 | Cause |
|---|---|---|---|
unanswerable-python4-ga-date | 0/24 | 24/24 | The decline rejected any date, including "as of 2026-09-28" and other releases' dates. |
unanswerable-private-company-audited-arr | 3/23 graded | 22/23 graded | The prompt asks for how circulating revenue figures were derived; any dollar figure failed. The 24th run (Parallel) hit its budget before delivering, is not graded, and counts as a failure in the tables, so the per-task row reads 22/24. |
upstream-diagnosis-urllib3-libressl-import-error | 0/24 | 24/24 | A URL-only check on a field every setup filled in prose; the URLs were in evidence_urls. |
list-build-sca-tool-shortlist (judged) | mostly ungraded | graded | The judge sometimes fences its JSON; the default profile name also counted as an setup-identity leak. |
With the fixes, the competitive pass rate rose between 14 and 29 points per setup (no-search +14, Exa +29). It rose most for search setups, which had been penalised for citing and explaining evidence. The first reading of this battery, that no-search was about as good as any provider, was an artifact of these bugs.
Genuine failures that remain:
- Egress pricing: unavailable AWS intra-AZ runs, as above.
change-detection-openapi-30-to-31: most misses omitexample→examplesfrom the breaking changes. OAS 3.1 deprecatesexamplerather than removing it. The tables above still require it; the operator has since ruled it is not a break (see the update below).- Budget exhaustion: 8 runs. Six are Firecrawl or Parallel on long list-building and enumeration tasks; Brave and Tavily have one each.
- By design: the two remaining
expected_failtasks (unversioned-runtime-drift,unindexed-ordering-regression) are 0/3 everywhere.
Codex (partial)
The codex agent (gpt-6-sol) ran one competitive task on 2026-09-28 before its weekly quota ran out:
- Egress pricing: native 2/3, no-search 1/3. Brave and Tavily went 0/3 on the token budget, reading 393k–448k tokens per run.
- Expected-fail tasks: codex's search setups burned the full provider-call budget on
multi-hop-npm-semver-dependentsinstead of declining.
A full codex battery waits for the quota reset (2026-10-04 12:52Z).
Caveats
- Sample size: n=42 per setup on the competitive set. Adjacent setups are within each other's intervals.
- Timing: the two runs were about ten hours apart. Same agent, model, tasks and budgets.
- Pricing: vendor spend is unpriced for Exa and Parallel, and for Perplexity's
asktool. Their $/success counts model tokens plus metered calls only. - Grading: the judged tasks use a single claude judge, so agreement is not measured.
Update (2026-09-30): operator decisions
- OpenAPI:
example→examplesis a deprecation, not a break, so the task no longer requires it. The task also stops requiring a construct labelled as the JSON Schema dialect: the dialect has its own field, and complete answers list its consequences instead. Both changes are in the bench fix. On replay, the task goes from 7/24 to 22/24. The two remaining misses cite hosts outside the allowlist. - Codex native search is priced at Brave's per-request rate as a proxy for its backing vendor. Claude native search keeps Anthropic's published rate.
- Tokens and outcome rate are reported together. Tokens per success come from the bench fix.
Competitive pass rates regraded with the corrected catalog at the recorded grader revision (native: the 2026-10-06 rerun; see the update at the top):
| setup | pass | tokens/success |
|---|---|---|
| firecrawl | 38/42 (90%) | 61.5k |
| perplexity | 37/42 (88%) | 47.7k |
| tavily | 36/42 (86%) | 61.3k |
| exa | 35/42 (83%) | 65.0k |
| native | 35/42 (83%) | 39.7k |
| parallel-web | 33/42 (79%) | 69.8k |
| no-search | 32/42 (76%) | 20.9k |
| brave | 31/42 (74%) | 89.2k |
The historical tokens/run column averages measured total-billable tokens over attempted runs with measured usage; unknown or estimated usage is excluded, never treated as zero. The update's tokens/success sums measured tokens over graded runs and divides by measured successes, as defined by the bench fix; it excludes ungraded and unmeasured runs and is not a full-setup lower bound. The per-setup measured coverage counts were not retained here. Multiplying a measured-run mean by all 42 attempts need not reproduce it. The previously quoted fresh/cache/output means are withdrawn: their run coverage and bucket accounting were not recorded here and they do not reconcile for Firecrawl, Parallel or no-search. They cannot support a token-mix comparison from this document.
The output-reduction claim is also withdrawn. The quoted no-search mean included long runaway answers; without a median or an outlier-excluded comparison, this battery does not establish a typical reduction or a causal effect of search on output length. Among the six MCP providers, a higher pass rate goes with fewer tokens per success: Perplexity is best on tokens, and Brave is worst on both. Spearman ρ ≈ −0.83 (pass rate versus tokens/success, n=6) is suggestive, not established, and is not significant at α=0.05. No setup's pass-rate difference from no-search has an interval that excludes zero. Firecrawl comes closest at +14 points (−2 to +30).
Methodology for this study
WSB methodology
The experiment used Claude Code with claude-opus-5-5[1m], all 18 production tasks, eight tool setups and three repetitions. The two setup groups ran about ten hours apart. Setups differ in exposed search tools: no-search, native, or exactly one provider MCP server. Budgets and contamination checks apply to each run.
The 14 competitive tasks and four expected-fail tasks have separate denominators. Expected-fail success means an honest partial answer or decline. Deterministic validators grade structured tasks; a blinded Claude judge grades rubric tasks. Single-judge agreement was not measured. Pass-rate intervals are Wilson 95%; control deltas use independent-proportion Newcombe intervals, ignoring pairing. Adjacent rankings are not statistically established.
The original bench corrections fixed overly strict date/dollar decline checks, a prose-versus-URL field mismatch and judge JSON parsing/blinding. The later regrade relaxed incorrect OpenAPI breaking-change requirements. Both grade the same stored answers; neither reran agents. REPORT.md retains initial and updated tables with their original precision.
Historical cost lower bounds sum priced spend and divide by all successes. The report command instead uses priced-run successes, so its output can differ. Unknown vendor spend is not zero; partially priced setups are excluded from cost ranking. Updated tokens per success use measured tokens from graded runs divided by measured successes; coverage counts were not retained. Do not multiply the historical tokens/run means by attempted-run counts to derive that update. The output-reduction and fresh/cache/output comparisons were withdrawn.
See recorded results, summary and reproduction. Calibration was not part of this experiment: calibration record.
Search gap bench (2026-10-03)
Full recorded report
GAP battery results, 2026-10-03: does search change the outcome of the job?
Update (2026-10-06): the Claude Code native setup was rerun with page fetching working. In the original battery the setup exposed WebSearch and WebFetch but pre-approved only WebSearch, and headless Claude Code refused all 50 WebFetch calls across its 21 runs, so that row measured search without page fetching. The configuration is fixed and pinned by a test. The setup's 21 runs were rerun on 2026-10-06 on the same 7 tasks, repetitions, model and calibration; no call was refused. The Claude Code native row and every figure derived from it below use the rerun: 16/21 passed (76%), gap closure 0.80 (0.53–1.00), 61k tokens per success. The original row recorded 13/21 (62%), gap closure 0.65 (0.37–0.91) and 66k. The other setups are unchanged. The rerun's verdicts come from the claude-code primary judge, which decides every verdict in this battery. On 2026-10-07 the codex judge scored the same 21 stored payloads: it agreed on every verdict and on 94% of labels (kappa 0.70); see judge agreement. The rerun's agents used about 1.0M tokens. The codex native setup was unaffected. Details: agent search behavior.
The first full Search Gap Bench (GAP) battery ran on 2026-10-03, operator-approved: brief tasks on two agents, nine setups each, three repetitions.
- GAP grades the job, not the retrieval.
- A task counts only after calibration shows its knowledge gap: the no-search setup (floor) passes at most 1 of 5, and the setup given the answer excerpt (ceiling) passes at least 4 of 5.
- Each search setup is scored by gap closure: how much of the floor-to-ceiling distance it covers.
- Tokens are reported beside every outcome.
Headline. On these tasks, search decides the outcome.
- The floor passed 0% on both agents, and every search setup passed 71–95%.
- Which provider is best depends on the agent: Parallel led on claude-code and Brave on codex.
- Pass rate did not predict cost: providers with similar pass rates differed almost twofold in tokens per success.
These are the recorded results after the Codex reruns. Claude-code provider tool availability remains unverified in the published evidence: no transcript audit for those runs is documented, and they lack availability records. Its provider rankings are provisional pending that audit. In the first codex pass, 31 of 108 provider runs ran without their search tool. They were re-run on 2026-10-04 under the GAPMCP-01 fix, every re-run run loaded its tool, and the codex tables below include them (see "Execution changes").
Setup
| Agents | claude-code on claude-opus-5-5, codex on gpt-6.1-sol |
| Setups | floor (no search), ceiling (answer excerpt in the prompt), native, Brave, Tavily, Exa, Parallel, Firecrawl, Perplexity |
| Tasks | 10 brief tasks from GAP-09: 6 decision briefs and 4 research briefs. The code corpus (GAP-07) merged during the run and gets its own battery. |
| Repetitions | 5 per reference setup in calibration; 3 per setup in the battery |
| Grading | Decision correctness, weighted key-fact recall (≥ 0.7) and unsupported-claim rate (≤ 0.25), judged by two blinded judges (claude-code primary, codex secondary) against captured primary sources |
| Spend | about 50M tokens in total: battery agents 21.1M; battery judges 15.3M; calibration agents 4.2M; calibration judges 9.2M |
Calibration
Final admissions under GAPRULE-01. Each run shows floor passes and ceiling passes out of 5.
| Task | Family | claude-code | codex |
|---|---|---|---|
| billing-reporting | research | Floor: 0, Ceiling: 5 — admitted | Floor: 0, Ceiling: 5 — admitted |
| classroom-term | decision | Floor: 0, Ceiling: 5 — admitted | Floor: 0, Ceiling: 5 — admitted |
| models-production | decision | Floor: 0, Ceiling: 5 — admitted | Floor: 0, Ceiling: 5 — admitted |
| npm-token-scope | research | Floor: 0, Ceiling: 5 — admitted | Floor: 0, Ceiling: 5 — admitted |
| org-migration | decision | Floor: 0, Ceiling: 5 — admitted | Floor: 0, Ceiling: 4 — admitted |
| python-metadata | research | Floor: 0, Ceiling: 5 — admitted | Floor: 0, Ceiling: 5 — admitted |
| tarfile-upload | decision | Floor: 0, Ceiling: 5 — admitted | Floor: 0, Ceiling: 0 — rejected |
| cli-trust | decision | Floor: 0, Ceiling: 2 — rejected | Floor: 0, Ceiling: 2 — rejected |
| spark-editing | decision | Floor: 0, Ceiling: 1 — rejected | Floor: 0, Ceiling: 0 — rejected |
| python-security-march | research | Floor: 0, Ceiling: 0 — rejected | Floor: 0, Ceiling: 1 — rejected |
Every floor scored 0, so every task has a real knowledge gap. All rejections are ceiling failures, which point at task or rubric defects:
- spark-editing: the heaviest key fact is not what the prompt asks about.
- python-security-march: the excerpt omitted the date that its top fact requires. The excerpt fix is pending for the next battery; GAPRULE-01 changed the unsupported-claim limit, not the excerpt.
- cli-trust: with three claims, a single unsupported claim fails the brief.
- tarfile-upload on codex: the codex briefs miss key facts.
Gap closure
claude-code (7 tasks, 21 runs per setup)
| Setup | Gap closure (95% CI) | Pass rate (Wilson 95% CI) | Tokens per success | vs floor (exact McNemar) |
|---|---|---|---|---|
| Parallel | 1.00 (0.86–1.17) | 95% (77–99%) | 64k | p < 0.0001 |
| Firecrawl | 0.95 (0.83–1.00) | 90% (71–97%) | 116k | p < 0.0001 |
| Brave | 0.85 (0.62–1.00) | 81% (60–92%) | 93k | p < 0.0001 |
| Perplexity | 0.80 (0.50–1.00) | 76% (55–89%) | 74k | p < 0.0001 |
| native | 0.80 (0.53–1.00) | 76% (55–89%) | 61k | p < 0.0001 |
| Exa | 0.75 (0.47–0.95) | 71% (50–86%) | 89k | p < 0.0001 |
| Tavily | 0.75 (0.43–1.00) | 71% (50–86%) | 83k | p < 0.0001 |
| ceiling (reference) | 1.00 | 95% | 4k | |
| floor (reference) | 0.00 | 0% | — |
By family:
- Research briefs: every search setup passed 100%, native included.
- Decision briefs: they separate the providers. Parallel 92% and Firecrawl 83% lead; Brave scored 67%, Perplexity and native 58% each, and Exa and Tavily 50% each.
codex (6 tasks, 18 runs per setup)
| Setup | Gap closure (95% CI) | Pass rate (Wilson 95% CI) | Tokens per success | vs floor (exact McNemar) |
|---|---|---|---|---|
| Brave | 1.06 (0.88–1.29) | 94% (74–99%) | 111k | p < 0.0001 |
| Firecrawl | 1.00 (1.00–1.00) | 89% (67–97%) | 166k | p < 0.0001 |
| Tavily | 1.00 (0.71–1.29) | 89% (67–97%) | 129k | p < 0.0001 |
| Exa | 0.88 (0.47–1.29) | 78% (55–91%) | 146k | p = 0.0001 |
| native | 0.88 (0.47–1.20) | 78% (55–91%) | 117k | p = 0.0001 |
| Parallel | 0.81 (0.41–1.14) | 72% (49–88%) | 157k | p = 0.0002 |
| Perplexity | 0.81 (0.47–1.12) | 72% (49–88%) | 117k | p = 0.0002 |
| ceiling (reference) | 1.00 | 89% (67–97%) | 12k | |
| floor (reference) | 0.00 | 0% (0–18%) | — |
No search setup beat native significantly; the largest gain was Brave at +16.7 points (p = 0.25).
By family:
- Research briefs: every setup passed 100%, except Exa at 89%.
- Decision briefs: they separate the providers. Brave scored 89%; Firecrawl and Tavily 78% each; Exa 67%; native 56%; Parallel and Perplexity 44% each.
The first pass's estimates that excluded the 31 tool-less runs were close to the re-run results:
| Setup | Estimate | After re-run |
|---|---|---|
| Brave | 93% | 94% |
| Firecrawl | 85% | 89% |
| Tavily | 85% | 89% |
| Exa | 76% | 78% |
| Parallel | 67% | 72% |
| Perplexity | 62% | 72% |
What this shows
1. On these tasks, search changes the outcome. Without search, both agents failed every brief. With a good provider they reached the answer-excerpt setup's level. The difference against the floor is significant for every search setup on both agents (exact McNemar, p ≤ 0.0002). 2. The best provider depends on the agent. On claude-code, Parallel (95%) and Firecrawl (90%) led. On codex, Brave led (94%), ahead of Firecrawl and Tavily (89% each).
- Parallel tied for last on codex (72%), even with its tool loaded in every run. On codex decision briefs it passed 44%, against 92% on claude-code.
- So a provider's ranking on one agent does not carry over to another.
3. Pass rate does not predict cost between providers. Tokens per success counts every agent token in the setup, failed runs included, divided by passing runs.
- The hypothesis that better search shows up as both lower token spend and higher success holds against no search, which spent tokens and passed nothing. On claude-code, Parallel paired the top pass rate (95%) with 64k tokens per success; only native search, at 76%, spent less per success (61k).
- It does not hold as a ranking. Among the seven search setups, the rank correlation between pass rate and tokens per success was +0.05 on claude-code and −0.20 on codex; negative means higher-passing setups cost less. With seven setups, neither value is distinguishable from zero.
- Per-run spend varied widely: 46k (native) to 105k (Firecrawl) on claude-code, and 84k (Perplexity) to 148k (Firecrawl) on codex. Firecrawl passed 90% at 116k tokens per success, against Parallel's 95% at 64k.
- Choosing a provider means weighing pass rate and cost separately.
4. Decision briefs separate the providers; research briefs mostly don't. Research briefs passed at or near 100% on both agents. Decision briefs hinge on one recent fact, and they spread the search setups from 50% to 92% on claude-code and from 44% to 89% on codex. 5. Native search was mid-pack on both agents. On claude-code it passed 76%, level with Perplexity and behind Parallel, Firecrawl and Brave. On codex it passed 78%, level with Exa and behind Brave, Firecrawl and Tavily; under the codex judge it would pass 94% (see judge agreement). No search setup beat it significantly on either agent; on claude-code the largest gain was Parallel's, +19.0 points (p = 0.22).
Judge agreement
Two blinded judges scored every brief on the same evidence: the claude-code judge (the Claude judge below) decides the verdict and the codex judge measures agreement. Both judged at grading time, except for the 21 rerun native runs on claude-code, which codex scored on 2026-10-07 from the payloads the primary had judged.
| Agent | Briefs judged | Claude-only passes | Codex-only passes | Label agreement | Kappa |
|---|---|---|---|---|---|
| claude-code | 187 | 4 | 1 | 96% | 0.79 |
| codex | 161 | 0 | 7 | 96% | 0.78 |
A Claude-only pass is a disputed verdict the Claude judge passed and codex failed; a codex-only pass is the reverse. Label agreement pools the fact and claim labels both judges scored, and kappa is Cohen's kappa on them. Claims whose cited source could not be captured are left out, as no judge scored them. Briefs that failed schema validation were never judged (2 claude-code and 1 codex floor runs).
The disputes lean towards each judge's own model family, most clearly on codex's briefs: codex passed all 7 disputed ones. On claude-code's briefs the Claude judge passed 4 of the 5 here, but only 2 of the 4 in the replication, 6 of 9 together. The direction differs by agent (two-sided Fisher exact on dispute direction, both batteries' claude-code disputes pooled: p = 0.011). That is consistent with a judge favouring its own family's writing; with two judges the bench cannot say which judge carries the preference, or whether both do. Verdicts stay with the primary judge, as the design fixes in advance. Had codex decided, the pass counts would be:
| Setup | claude-code, Claude judge | claude-code, codex judge | codex, Claude judge | codex, codex judge |
|---|---|---|---|---|
| Parallel | 20/21 | 20/21 | 13/18 | 13/18 |
| Firecrawl | 19/21 | 19/21 | 16/18 | 17/18 |
| Brave | 17/21 | 15/21 | 17/18 | 17/18 |
| Perplexity | 16/21 | 16/21 | 13/18 | 15/18 |
| native | 16/21 | 16/21 | 14/18 | 17/18 |
| Exa | 15/21 | 14/21 | 14/18 | 14/18 |
| Tavily | 15/21 | 15/21 | 16/18 | 16/18 |
| ceiling (reference) | 20/21 | 20/21 | 16/18 | 17/18 |
| floor (reference) | 0/21 | 0/21 | 0/18 | 0/18 |
Only one count moves by more than two runs: codex's native setup, from 14/18 (78%) to 17/18 (94%), which would put it level with Brave and Firecrawl at the top of the codex table. Search against no search is unaffected: no floor brief passes under either judge.
Changes made during the run
The first live runs surfaced several bench defects. Two changed how runs execute, and two changed how runs are graded. The timeline was recorded from the battery change log and run start events.
Execution changes
Every run in the tables above ran under both of these fixes:
- GAPBUDGET-01: brief budgets became runaway caps (1M tokens, 60 calls). The old 30k cap was below the median codex search run.
- It was part of the pinned export before any calibration or battery run ran.
- No run ran under the old caps.
- GAPARM-01: brief search setups resolve the GAP catalog. Before it, no brief search setup could run, so no search-setup run predates it.
- It was applied at 18:30:04Z on 2026-10-03.
- The earliest search-setup run started at 18:30:29Z (codex round 1). claude-code started at 19:12Z.
The tool-isolation remediation, which forbids the GAP catalog to every GAP run, was applied at 18:50:40Z and the codex battery was restarted to load it.
- 18 codex round-1 runs started before it: classroom-term repetitions 1 and 2, all setups.
- No catalog access in them. Their transcripts contain only search calls and agent messages, with no file reads or commands, so the catalog was not read.
- They were kept, not re-run.
- Every other run ran after the remediation.
The codex provider re-run (2026-10-04). In the first pass, 31 of codex's 108 provider runs made no provider call:
| Setup | Runs without a provider call |
|---|---|
| Parallel | 12 |
| Firecrawl | 5 |
| Tavily | 5 |
| Perplexity | 5 |
| Brave | 3 |
| Exa | 1 |
- Cause: GAPMCP-01 reproduced codex's 30-second default MCP startup timeout. A 35-second server fails under it and starts with a 120-second allowance.
- How the runs were selected: the 31 runs were reclassified from an audit of their transcripts. All of them had zero provider calls and a
list_mcp_resourcesprobe; 29 of the 31 also said in their answer that no retrieval was available. The private audit is not distributed. - How they were re-run: with
--rerun-unavailableunder GAPMCP-01 (a 120-second startup allowance,required=trueand pre-warmed pinned packages), using the same matrix, calibration and catalog. - Every re-run run loaded its tool. Each one recorded a completed
initializeand tool listing. - A second fix was needed. The first re-run attempt misclassified working runs, because the MCP meter could not load its core (SEV2 METERAVAIL, below). Making the metering library available in each provider server's environment fixed it, and those runs were re-run again.
These 31 runs ran under the startup allowance; the other 77 provider runs ran without it but had their tool. No agent-side setting changed.
Grading changes
These changed only how stored answers were judged. Every affected run was regraded from the agents' stored answers, with no agent re-runs:
- GAPJUDGE-01: judges read readable text instead of raw HTML, and each source once. This cut judge tokens per run from about 250k to about 15k. It also covered:
- safe redirects;
- gzip decoding (python.org served gzip unconditionally, which garbled every python.org citation);
- over-budget sources withheld instead of the whole run going ungraded;
- the blinding leak check reading decoded text, and generic setup labels no longer counted as leaks.
- GAPRULE-01, operator-approved: the unsupported-claim limit was raised from 0.10 to 0.25. Ceiling agents see a short excerpt, and their hedges about it ("the notice does not specify X") were contradicted by the full pages the judges read. All stored grades were recomputed from their saved judge labels; 49 outcomes flipped, all from fail to pass.
Open issues
- SEV2 GAPMCP (fix GAPMCP-01, the bench fix): resolved. The 31 affected codex runs were re-run and the codex table updated.
- SEV2 METERAVAIL (fix METERAVAIL-01): provider-call metering was silently off for this battery on both agents. The meter runs inside the agent's restricted MCP child environment, and from the exported workbench tree it could not load its core.
- Vendor spend per run was therefore not metered.
- Only the 31 re-run runs have availability records; the report's "availability coverage" column counts those. Claude-code provider availability has not been verified in the published evidence.
- SEV2 ARGVKEY (fix ARGVKEY-01): the Parallel API key was passed in
mcp-remote's process arguments. The battery configuration now passes it through a 0600 header file. - The claude-code floor setup sometimes writes pretend searches as text when it has no tools. It is a legitimate failure, but slow and expensive (up to about 130k output tokens in 13 minutes).
- Next battery:
- the code corpus (GAP-07: 16 code tasks and 4 controls);
- the python-security-march excerpt fix;
- rubric fixes for spark-editing and cli-trust.
Methodology for this study
GAP methodology
GAP grades completed jobs against captured primary sources. The initial corpus contained ten briefs: six decision and four research. Calibration is per agent and model: five repetitions of floor (no search) and ceiling (answer excerpt). Admission requires floor at most one pass and ceiling at least four passes. Seven Claude tasks and six Codex tasks were admitted. Code tasks were not run in this battery. The committed calibration file is an admission summary, not a runtime calibration file with captured oracle sources.
Each admitted task ran three times on nine setups. Brief grading requires decision correctness, weighted key-fact recall at least 0.7 and unsupported-claim rate at most 0.25. Two blinded judges used Claude Code as primary and Codex as secondary; the primary decides every verdict and the secondary measures agreement. The 21 Claude Code native runs rerun on 2026-10-06 were graded by the primary alone, and Codex scored their stored payloads on 2026-10-07 (sew gap agree), which rebuilds each payload from the stored source snapshots and checks its SHA-256 first. Gap closure is (setup pass rate − floor pass rate) / (ceiling pass rate − floor pass rate); it can exceed one. Wilson 95% intervals accompany pass rates. Gap-closure intervals use a seeded paired task bootstrap with 10,000 resamples. Exact McNemar tests pair task and repetition, against floor and native. The findings publish floor comparisons and the largest Codex native comparison, not every native test.
Tokens per success sum all agent tokens in an setup, including failures, divided by passes. Judge and calibration spend is separate. A zero-success floor has undefined tokens per success. Vendor spend was unmetered; tokens are not dollar cost.
GAPJUDGE-01 changed readable source extraction, redirects, gzip, source budget handling and blinding. GAPRULE-01 raised the unsupported-claim limit from 0.10 to 0.25; recomputing saved judge labels flipped 49 failures to passes. These grading changes did not rerun agents. Execution changes and the later rerun of 31 Codex runs with unavailable provider tools are preserved separately in REPORT.md.
See recorded results, summary, admission counts and reproduction.
How agents used search (2026-10-05)
Full recorded report
How agents used search: query style, fetch-from-memory and primary-source retrieval (2026-10-05)
This report analyses the stored transcripts of the two published batteries, the web search bakeoff and the search gap bench. No agent was rerun. It asks how agents used the search tools they were given, rather than which provider scored best. Every number below comes from scripts/analyze_search_behavior.py, whose classification rules are listed in methodology. The analysis covers 974 search queries from the GAP battery (both agents) and 655 from the bakeoff (Claude Code).
Headline
- Agents write keyword queries whatever the engine asks for. Of 974 GAP search queries, 7.3% read as natural language. In the bakeoff the figure is 6.0%. Exa's tool asks for a description of the ideal page rather than keywords; its queries were natural language in 3% (Claude Code) and 2% (codex) of GAP calls. Because the automatic rule can miss descriptive phrasing, all 111 Exa queries from both batteries were also read by hand (appendix): none is phrased as a description of a page. The closest are a headline-like phrase used three times ("Python 4.0 release date announced by Python Steering Council") and a page title.
- The two agents express constraints differently. On the six provider setups, codex put
site:into 31% to 67% of its queries, even where a structured domain filter existed. Claude Code wrotesite:only on Firecrawl (6%) and otherwise used structured domain filters where a tool offered one (Perplexitysearch_domain_filter, Tavilyinclude_domains, WebSearchallowed_domains). - Agents often fetch remembered URLs instead of searching. In the bakeoff, Claude Code fetched pages without any search in 36 of 55 Exa runs, 35 of 55 Firecrawl runs, 34 of 54 Tavily runs and 30 of 55 Parallel runs. Those provider setups therefore measured page fetching as much as search. On post-cutoff GAP tasks the same habit failed: Claude Code's Exa runs that only fetched passed 1/3, against 14/18 for its Exa runs that searched.
- Retrieving the primary source mattered most on Claude Code. Pooled over provider and native setups, Claude Code passed 94/111 runs (85%) when the task's primary source appeared in search results and 20/33 (61%) when it did not. Codex surfaced the primary source in most searched runs, so the comparison is thin there (95/116 against 8/10). Two setups passed without the primary page: Parallel (6/7) and Firecrawl (9/9) on Claude Code, which indicates their results carried the facts from other pages.
These are observations about agent behavior with each vendor's MCP server at its pinned version and default settings. They are not measurements of the vendors' APIs.
What each search tool asks for
Descriptions are quoted from the pinned MCP server packages and, for Parallel's hosted server, from its published documentation.
| Setup | Search interface the agent saw | What the tool tells the agent about queries |
|---|---|---|
| Exa | web_search_exa(query, numResults) | "describe the ideal page, not keywords"; query is a "semantically rich description of the ideal page, not just keywords" |
| Parallel | web_search(objective, search_queries, …) | a natural-language objective plus "concise, related keyword queries of 3–6 words each" |
| Perplexity | perplexity_search(query, …) plus ask, research, reason | "Supports recency filters, domain restrictions, and a lower-latency fast search mode" |
| Tavily | tavily_search(query, search_depth, …) | query is described only as "Search query" |
| Firecrawl | firecrawl_search(query, …) | "Query operators, domain filters, categories … are described on their parameters" |
| Brave | brave_web_search(query, …) | lists when to use the tool; no guidance on query phrasing |
| native (Claude Code) | WebSearch(query, allowed_domains, …) | agent built-in |
| native (codex) | built-in web search | agent built-in; no parameters exposed to the transcript |
Query style (GAP battery)
Parallel's search_queries are keyword queries by design; its natural-language part is the separate objective field, which is counted as a structured parameter and not as a query. Perplexity's user messages count as queries alongside its query inputs; system and assistant messages are excluded. Codex native search reports its query lists in the transcript; page views are counted as fetches.
| Agent | Setup | Queries | Natural language | Contains a year | Quoted phrase | site: | Median words | Structured parameters passed (calls) |
|---|---|---|---|---|---|---|---|---|
| claude-code | native | 101 | 7% | 43% | 39% | 0% | 11 | allowed_domains ×11 |
| claude-code | Brave | 72 | 17% | 44% | 18% | 0% | 8 | count ×23, extra_snippets ×21, maximum_number_of_tokens ×18, freshness ×1, maximum_number_of_urls ×1 |
| claude-code | Tavily | 38 | 13% | 45% | 18% | 0% | 9 | max_results ×38, include_domains ×19, search_depth ×17, include_raw_content ×1, start_date ×1 |
| claude-code | Exa | 38 | 3% | 58% | 0% | 0% | 10 | numResults ×38 |
| claude-code | Parallel | 94 | 5% | 24% | 2% | 0% | 6 | model_name ×27, objective ×27, session_id ×27 |
| claude-code | Firecrawl | 34 | 0% | 59% | 9% | 6% | 7 | limit ×34, sources ×34, includeDomains ×3 |
| claude-code | Perplexity | 70 | 20% | 36% | 10% | 0% | 9 | max_results ×60, search_domain_filter ×42, max_tokens_per_page ×5, messages ×5, search_context_size ×3, search_recency_filter ×2, search_type ×1 |
| codex | native | 48 | 0% | 50% | 33% | 10% | 8 | none |
| codex | Brave | 114 | 2% | 52% | 26% | 46% | 8 | maximum_number_of_tokens ×30, count ×15, maximum_number_of_urls ×13, extra_snippets ×7, url ×3 |
| codex | Tavily | 130 | 6% | 50% | 59% | 67% | 8 | max_results ×123, search_depth ×26, include_raw_content ×25, end_date ×3, include_domains ×3 |
| codex | Exa | 52 | 2% | 60% | 35% | 58% | 10 | numResults ×10 |
| codex | Parallel | 62 | 6% | 39% | 27% | 48% | 7 | objective ×28, session_id ×10, max_results ×3, max_chars_per_result ×1 |
| codex | Firecrawl | 42 | 2% | 74% | 26% | 31% | 8 | limit ×33, sources ×10 |
| codex | Perplexity | 79 | 14% | 71% | 33% | 42% | 10 | max_results ×20, messages ×18, max_tokens_per_page ×16 |
Query style (web search bakeoff, Claude Code)
| Setup | Runs | Queries | Natural language | Contains a year | Quoted phrase | site: | Median words | Structured parameters passed (calls) |
|---|---|---|---|---|---|---|---|---|
| native | 54 | 137 | 2% | 34% | 9% | 4% | 7 | allowed_domains ×36 |
| Brave | 54 | 199 | 4% | 7% | 11% | 6% | 6 | count ×85, maximum_number_of_tokens ×80, freshness ×17, goggles ×16, extra_snippets ×11, maximum_number_of_urls ×9, maximum_number_of_tokens_per_url ×3 |
| Tavily | 54 | 25 | 0% | 32% | 12% | 0% | 10 | max_results ×18, include_domains ×10, search_depth ×4, time_range ×2, start_date ×1 |
| Exa | 55 | 21 | 0% | 19% | 5% | 0% | 10 | numResults ×21 |
| Parallel | 55 | 77 | 0% | 20% | 0% | 1% | 5 | model_name ×22, objective ×22, session_id ×22 |
| Firecrawl | 55 | 20 | 0% | 25% | 15% | 5% | 7 | limit ×20, sources ×20 |
| Perplexity | 55 | 176 | 16% | 13% | 10% | 0% | 8 | max_results ×154, search_domain_filter ×115, max_tokens_per_page ×21, search_type ×21, messages ×13, search_context_size ×13, search_recency_filter ×5 |
Searching versus fetching from memory (web search bakeoff, Claude Code)
A run is "fetched without searching" when it made at least one fetch call and no search call. Brave's and Perplexity's servers expose no page-fetch tool, so their agents had to search or answer from memory. Every setup has a handful of runs with no provider call: answers given from memory and runs that ended before any tool call.
| Setup | Runs | Fetch tool in the setup | Fetched without searching | No provider call | Calls refused by the agent |
|---|---|---|---|---|---|
| native | 54 | WebFetch (mostly refused; see below) | 12 | 6 | 110 |
| Brave | 54 | no | 0 | 7 | 0 |
| Tavily | 54 | yes (tavily_extract) | 34 | 6 | 0 |
| Exa | 55 | yes (web_fetch_exa) | 36 | 6 | 0 |
| Parallel | 55 | yes (web_fetch) | 30 | 6 | 0 |
| Firecrawl | 55 | yes (firecrawl_scrape) | 35 | 6 | 0 |
| Perplexity | 55 | no | 0 | 7 | 0 |
On the bakeoff's stable-knowledge tasks, fetching a remembered URL is efficient: those tasks passed in every setup, including no-search. It means the bakeoff's provider setups with a fetch tool mostly measured that fetch path. The GAP battery, whose answers post-date the models, shows the cost of the same habit on new information.
Agent refusals: the Claude Code native setup
Claude Code's native setup exposed WebSearch and WebFetch through --tools but pre-approved only WebSearch through --allowedTools. Headless Claude Code refuses a tool that is exposed but not pre-approved ("Claude requested permissions to use WebFetch, but you haven't granted it yet"). In the GAP battery every native WebFetch call was refused (50 calls in 21 of 21 runs); in the bakeoff, 110 calls in 40 of 54 runs. No provider setup and no codex setup had a refused call. The Claude Code native results in both batteries therefore understate that agent's web tools. The setup configuration is fixed and covered by a regression test. The fetch counts in this report include refused attempts. The setup was rerun with fetching working; see Native rerun (2026-10-06).
Primary-source retrieval (GAP battery)
Each GAP brief has one primary-source URL (the catalog oracle). A run surfaced it when that URL appeared among the URLs of returned search results, excluding search inputs. Unknown coverage is excluded from both pass-rate groups. Recomputing against returned evidence leaves the historical source counts unchanged; no searched run has unknown coverage. Facts can also come from secondary pages, so surfacing is a proxy for retrieval, not a requirement for passing.
| Agent | Setup | Runs that searched | Unknown coverage | Primary source surfaced | Pass when surfaced | Pass when not surfaced |
|---|---|---|---|---|---|---|
| claude-code | native | 21 | 0 | 18/21 | 13/18 | 0/3 |
| claude-code | Brave | 21 | 0 | 18/21 | 15/18 | 2/3 |
| claude-code | Tavily | 21 | 0 | 16/21 | 13/16 | 2/5 |
| claude-code | Exa | 18 | 0 | 14/18 | 13/14 | 1/4 |
| claude-code | Parallel | 21 | 0 | 14/21 | 14/14 | 6/7 |
| claude-code | Firecrawl | 21 | 0 | 12/21 | 10/12 | 9/9 |
| claude-code | Perplexity | 21 | 0 | 19/21 | 16/19 | 0/2 |
| codex | native | 18 | 0 | 17/18 | 13/17 | 1/1 |
| codex | Brave | 18 | 0 | 16/18 | 15/16 | 2/2 |
| codex | Tavily | 18 | 0 | 16/18 | 14/16 | 2/2 |
| codex | Exa | 18 | 0 | 17/18 | 14/17 | 0/1 |
| codex | Parallel | 18 | 0 | 17/18 | 12/17 | 1/1 |
| codex | Firecrawl | 18 | 0 | 15/18 | 14/15 | 2/3 |
| codex | Perplexity | 18 | 0 | 18/18 | 13/18 | — |
Native rerun (2026-10-06)
The Claude Code native setup was rerun on both batteries with WebFetch pre-approved: the 21 GAP runs and all 54 bakeoff runs, on the same tasks, repetitions and model. No call was refused. The tables above keep the original runs, because they document the refusal; the GAP and bakeoff reports use the rerun for their native rows.
With page fetching available, the agent searched far less and read pages instead. On GAP it issued 45 queries instead of 101, and it fetched 126 pages. On the bakeoff it issued 19 queries instead of 137, and 36 of its 54 runs fetched remembered pages without searching at all.
| Battery | Runs | Queries (original) | Fetch calls | Refused calls (original) | Fetched without searching | Primary source surfaced | Pass when surfaced | Pass when not surfaced |
|---|---|---|---|---|---|---|---|---|
| GAP | 21 | 45 (101) | 126 | 0 (50) | 0 | 17/21 | 15/17 | 1/4 |
| Bakeoff | 54 | 19 (137) | 426 | 0 (110) | 36 | — | — | — |
The rerun's GAP queries read as natural language in 2.2% of cases (original 7%), with a median of 8 words (original 11). The rerun's run bundles are private, like the originals.
Limits
- Run counts per setup are small (18 to 21 in GAP, 54 to 55 in the bakeoff). Per-setup percentages move by about five points per run.
- The natural-language rule is a heuristic, stated in methodology; readers can rerun the script with their own rule.
- Each vendor was used through its MCP server at the pinned version and default settings. Other interfaces of the same vendor (direct API, other tools, other modes) were not tested.
- The two agents ran different task subsets in GAP (7 tasks for Claude Code, 6 for codex), so cross-agent differences mix agent behavior with task mix.
Appendix: every Exa query
Every query sent to web_search_exa in both batteries, in transcript order, so readers can judge the phrasing directly. Other setups' queries are summarised above but not reproduced.
| Battery | Agent | Query |
|---|---|---|
| GAP | claude-code | Copilot Billing Preview app deprecated usage report |
| GAP | claude-code | GitHub Copilot premium request budgets per user cap and usage report CSV export docs |
| GAP | claude-code | GitHub docs AI usage page billing settings group filter export AI credits Copilot limitations |
| GAP | claude-code | GitHub Copilot Billing Preview app deprecated usage-based billing budgets |
| GAP | claude-code | GitHub docs Copilot premium requests usage report export CSV billing 2026 |
| GAP | claude-code | GitHub docs AI usage page billing settings group filter export AI credits usage report limitations |
| GAP | claude-code | GitHub Copilot Billing Preview app deprecated usage report |
| GAP | claude-code | GitHub docs set budget for individual Copilot premium requests user-level budget |
| GAP | claude-code | GitHub docs AI usage page billing settings group filter export AI credits Copilot usage report limitations |
| GAP | claude-code | github.blog changelog 2026-08-04 Retiring the Copilot Billing Preview app |
| GAP | claude-code | GitHub Classroom deprecation sunset announcement 2026 |
| GAP | claude-code | Classroom 50 CS50 foundation GitHub Classroom alternative import autograding LTI |
| GAP | claude-code | Codio GitHub Classroom migration partner |
| GAP | claude-code | GitHub Classroom sunset deprecation announcement 2026 |
| GAP | claude-code | Classroom 50 CS50 open source GitHub Classroom replacement import migration |
| GAP | claude-code | GitHub Classroom deprecation sunset announcement 2026 |
| GAP | claude-code | GitHub Models deprecation retirement changelog 2026 BYOK inference |
| GAP | claude-code | GitHub Models deprecation retirement changelog 2026 BYOK paid usage |
| GAP | claude-code | Azure AI Inference beta SDK deprecated retirement August 26 2026 migrate to OpenAI v1 API |
| GAP | claude-code | GitHub Models deprecation retirement changelog 2026 BYOK |
| GAP | claude-code | GitHub secret scanning npm token revocation credential types 2026 changelog |
| GAP | claude-code | GitHub changelog npm granular access tokens publish and manage package access 2026 |
| GAP | claude-code | GitHub secret scanning npm token revocation credential types in scope 2026 |
| GAP | claude-code | GitHub changelog npm granular access tokens bypass 2FA publish manage package access 2026 |
| GAP | claude-code | GitHub changelog npm granular access tokens bypass 2FA publish package access management 2026 |
| GAP | claude-code | GitHub secret scanning npm token revocation credential types 2026 changelog |
| GAP | claude-code | GitHub changelog npm staged publishing generally available npm stage publish 2026 |
| GAP | claude-code | Python 3.10 security release tarfile extraction filter bypass CVE-2025-4517 |
| GAP | claude-code | CPython tarfile data filter vulnerability 2026 CVE security-announce |
| GAP | claude-code | Python 3.10 security release tarfile extraction filter bypass fix |
| GAP | claude-code | CVE-2026-4360 tarfile extract filter hardlinks Python fixed version |
| GAP | claude-code | GitHub docs converting a user into an organization warning cannot sign in personal account |
| GAP | claude-code | GitHub Enterprise Server docs promoting or demoting site administrator converting user to organization |
| GAP | claude-code | python.org release metadata API authentication vulnerability security advisory 2026 |
| GAP | claude-code | python.org release metadata vulnerability unauthenticated API modify release files 2026 |
| GAP | claude-code | python.org release API authentication bypass DEVCORE Splitline CVE advisory exploitation |
| GAP | claude-code | python.org release metadata vulnerability unauthenticated modify release files advisory 2026 |
| GAP | claude-code | Python Software Foundation security incident python.org downloads release API authentication |
| GAP | codex | GitHub Classroom sunset July 2026 September 2026 partner repositories |
| GAP | codex | GitHub Classroom retirement September 1 2026 repositories partner |
| GAP | codex | site:github.com/orgs/community/discussions/205975 "September 4" |
| GAP | codex | GitHub Classroom sunset September 2026 repositories partner retirement |
| GAP | codex | site:github.com/orgs/community/discussions GitHub Classroom "September 4" "2026" |
| GAP | codex | GitHub Models inference BYOK retirement July August 2026 paid existing usage |
| GAP | codex | GitHub Models retiring July August 2026 paid inference BYOK July 10 |
| GAP | codex | GitHub Models retirement July August 2026 paid inference BYOK existing usage |
| GAP | codex | site:docs.github.com converting user into organization January 2026 transfer repositories personal account enterprise server |
| GAP | codex | site:github.blog/changelog 2026 January converting user organization disabled January 2026 |
| GAP | codex | site:docs.github.com converting a user into an organization January 2026 transfer repositories |
| GAP | codex | site:github.blog/changelog 2026 January converting personal accounts organizations January 20 |
| GAP | codex | site:docs.github.com converting user into organization January 2026 transfer repositories personal account GHES |
| GAP | codex | site:github.blog/changelog 2026 January 20 convert account organization transfer personal account |
| GAP | codex | GitHub Copilot Billing Preview app retired individual budgets usage export billing reports |
| GAP | codex | site:docs.github.com Copilot AI usage page reports limitations billing API user level budgets export usage credits |
| GAP | codex | site:docs.github.com "AI usage" "report" "does not" billing usage reports |
| GAP | codex | site:docs.github.com "Setting up budgets" "user-level" "Budgets and alerts" |
| GAP | codex | "Copilot Billing Preview" budgets export usage |
| GAP | codex | site:docs.github.com AI usage user-level budgets billing reports Copilot AI credits |
| GAP | codex | site:docs.github.com "AI usage" "does not" reports billing |
| GAP | codex | site:docs.github.com "Viewing your usage" "AI" "report" billing |
| GAP | codex | site:docs.github.com "Automating usage reporting" "AI" |
| GAP | codex | site:docs.github.com Copilot usage metrics billing data coverage IDE chat github CLI |
| GAP | codex | site:docs.github.com "Setting" "user-level budgets" "Budgets" |
| GAP | codex | Copilot Billing Preview app individual budgets export raw usage finance billing dashboard |
| GAP | codex | site.github.blog/changelog 2026-08-04 retiring copilot billing preview |
| GAP | codex | site:docs.github.com AI usage page billing reports user-level budgets coverage limitations Copilot |
| GAP | codex | site:docs.github.com "AI usage" "does not" reports billing |
| GAP | codex | site:docs.github.com "AI usage" "coverage" |
| GAP | codex | site:docs.github.com "Viewing your usage of metered products" "AI usage" |
| GAP | codex | site:github.blog/changelog "user-level budgets" 2026 |
| GAP | codex | site:docs.github.com copilot usage metrics "billing" "coverage" |
| GAP | codex | site:github.blog/changelog billing CSV usage reports API 2026 |
| GAP | codex | npm token security granular tokens package access management July 2026 August 2026 GitHub credentials |
| GAP | codex | site:github.blog/changelog/2026 staged publishing npm May 22 trusted publishing tokens approve |
| GAP | codex | site:github.blog/changelog/ "2026-05-22" "npm" |
| GAP | codex | npm token security granular tokens package access github August 2026 publishing tokens |
| GAP | codex | site:github.blog/changelog npm December 9 2025 classic tokens revoked staged publishing July 2026 |
| GAP | codex | npm token security August 2026 package access granular tokens GitHub credentials |
| GAP | codex | site:github.blog/changelog npm staged publishing 2026 July 2FA |
| GAP | codex | python.org release metadata vulnerability unsigned downloads API incident 2026 2025 |
| GAP | codex | site.python.org Sigstore verify Python releases identity PEP 761 |
| GAP | codex | site.blog.python.org/2026/06/mitigated-api-bypass "Timeline" "30" "June" |
| GAP | codex | "python/pythondotorg" "3014" "merged" "June" |
| GAP | codex | python.org release metadata vulnerability authentication 2026 2025 security incident |
| GAP | codex | site:blog.python.org/2026/06/mitigated-api-bypass "Remediations" "Timeline" |
| GAP | codex | site:github.com/python/pythondotorg/issues/3010 OR site:github.com/python/pythondotorg/issues/3011 |
| GAP | codex | python.org release metadata security incident authentication bypass before June 30 2026 Trail of Bits |
| GAP | codex | python.org release metadata vulnerability authentication 2026 2025 API incident |
| GAP | codex | site.python.org security incident release metadata 2026 2025 downloads authentication |
| GAP | codex | Python Windows code signing certificates revoked August 2025 update investigation June 2026 |
| bakeoff | claude-code | Python 4.0 release date announced by Python Steering Council |
| bakeoff | claude-code | webpack md4 OpenSSL 3 Node 17 hash wasm md4 fix output.hashFunction xxhash64 sokra comment |
| bakeoff | claude-code | grype offline air-gapped vulnerability database import GRYPE_DB_AUTO_UPDATE false |
| bakeoff | claude-code | cargo audit --no-fetch --db local advisory database path offline |
| bakeoff | claude-code | Dependency-Track air-gapped offline vulnerability mirror configuration NVD OSV feeds URL |
| bakeoff | claude-code | Python 4.0 release date announced by Python Steering Council |
| bakeoff | claude-code | Python 4.0 release date announced by Python Steering Council |
| bakeoff | claude-code | Node.js documentation --openssl-legacy-provider option "Enable OpenSSL 3.0 legacy provider" |
| bakeoff | claude-code | aws.amazon.com data transfer out to internet US East (N. Virginia) first 10 TB per GB price same availability zone free |
| bakeoff | claude-code | grype air-gapped offline database import GRYPE_DB_AUTO_UPDATE false documentation |
| bakeoff | claude-code | Dependency-Track air-gapped offline NVD mirror configuration documentation |
| bakeoff | claude-code | Safety CLI 3 requires account login API key vulnerability database commercial license |
| bakeoff | claude-code | AWS EC2 on-demand pricing data transfer out to internet US East (N. Virginia) first 10 TB per month per GB; data transfer same availability zone free |
| bakeoff | claude-code | Stripe annual letter 2025 total payment volume private company revenue |
| bakeoff | claude-code | Stripe revenue estimate UK Companies House filing Stripe Payments Europe accounts net revenue |
| bakeoff | claude-code | Stripe annual letter 2025 total payment volume private company revenue |
| bakeoff | claude-code | Stripe revenue estimate net revenue report private company does not disclose financials |
| bakeoff | claude-code | Stripe Irish entity accounts filed Companies Registration Office revenue $5.1bn pre-tax profit 2024 |
| bakeoff | claude-code | Stripe annual letter 2025 total payment volume private company revenue estimate |
| bakeoff | claude-code | Grype offline air-gapped vulnerability database GRYPE_DB_AUTO_UPDATE false grype db import |
| bakeoff | claude-code | Dependency-Track air-gapped offline mirror NVD OSV internal mirror documentation |
Methodology for this study
Agent search behavior: methodology
The analysis reads the stored transcripts of the web search bakeoff (Claude Code, eight setups, 18 tasks, three repetitions) and the search gap bench (Claude Code and codex, brief tasks, three repetitions per admitted task). It runs no agent and changes no grade. Pass and fail come from each run's recorded evaluation.
scripts/analyze_search_behavior.py extracts every completed provider tool call from a transcript:
- codex MCP tool calls, and codex built-in
web_searchitems; - Claude Code
tool_useandtool_resultpairs, including vendor MCP tools, WebSearch and WebFetch.
Classification rules
- Search call: a call to a search tool (
web_search_exa,brave_web_search,tavily_search, Parallelweb_search,firecrawl_search, the Perplexity search tools, Claude Code WebSearch, and codex built-insearchactions). - Query: each entry of a multi-query call (Parallel
search_queries, codex built-in query lists) counts as one query. Each user message in Perplexity's conversational inputs also counts as one query; system and assistant messages are excluded. Parallel'sobjectiveis counted as a structured parameter whensearch_queriesis present, not as another query. - Fetch call: a page-retrieval tool call (
web_fetch_exa,tavily_extract,firecrawl_scrape, Parallelweb_fetch, Claude Code WebFetch, and codex built-inopen_page,find_in_pageand other page views). - Natural-language query: ends with "?", starts with an interrogative or auxiliary verb, or contains at least two function words from a fixed list (the, a, an, is, are, was, were, does, do, did, how, what, which, why, when, where, who, that, this, of, to, for, with, will, should, can, about, on, from, by, as, be, has, have).
- Contains a year: a token from 2000 to 2099. Quoted phrase: contains a double quote.
site:: containssite:. - Fetched without searching: the run made at least one fetch call and no search call.
- Primary source surfaced (GAP only): the task's catalog oracle
source_url, normalised (scheme, query, fragment and trailing slash removed), appears among the URLs in any returned search result of the run, never in the search input. It is measured only over runs that searched. A run is unknown when any search lacks exposed results and no observed result contains the primary source. Unknown runs are excluded from both surfaced and not-surfaced pass-rate denominators. Recomputing against returned evidence leaves the historical source counts unchanged; no searched run has unknown coverage.
The rules are deliberately simple so they can be audited. The natural-language rule undercounts descriptive keyword-free phrases and overcounts keyword strings that happen to contain function words. Readers can substitute their own rule in the script and rerun it.
Scope
Pooled figures cover the provider and native setups. Reference setups (no-search, floor and ceiling) make no search calls and are excluded from query statistics. GAP run counts differ by agent because calibration admitted seven brief tasks for Claude Code and six for codex.
See the report, summary data and reproduction. Calibration does not apply: calibration record.
Search gap bench replication (2026-10-06)
Full recorded report
GAP replication on Claude Code, 2026-10-06: the search gap bench, run again
Grading note. These verdicts come from the claude-code primary judge, which decides every verdict in GAP. The codex judge, which only measures agreement, was out of quota during the battery and scored the same stored payloads on 2026-10-07. It disputed 4 of the 211 judged verdicts and agreed on 96% of labels (kappa 0.82); no verdict changed. See judge agreement.
A second, independent Search Gap Bench battery on Claude Code (claude-opus-5-5), run on 2026-10-06 and 2026-10-07 with the apparatus as fixed after the 2026-10-03 battery: the built-in search setup can open pages, and every provider setup loaded its tool in every run. It repeats that battery's method on the current task catalog: calibrate each brief, keep the tasks with a real knowledge gap, and run every admitted task on nine setups, three times each.
Headline.
- Search still decides the outcome. Without search the agent passed 0 of 24 runs; every search setup passed 62–83%, and each differs from no search significantly (exact McNemar, p < 0.0001).
- The leader changed. Firecrawl led at 83%. Parallel, which led the 2026-10-03 battery at 95%, passed 75%. No search setup differs significantly from the agent's built-in search (largest gap: Firecrawl, +8.3 points, p = 0.50).
- Built-in search was competitive and cheapest. It passed 75%, level with Parallel and Perplexity, at 55k tokens per success, the lowest of any search setup.
- Pass rate again did not predict cost. The rank correlation between pass rate and tokens per success across the seven search setups was −0.04. Firecrawl's top pass rate came at 141k tokens per success.
Setup
| Agent | claude-code on claude-opus-5-5 |
| Setups | floor (no search), ceiling (answer excerpt in the prompt), native (built-in WebSearch and WebFetch), Brave, Tavily, Exa, Parallel, Firecrawl, Perplexity |
| Tasks | the 10 brief tasks of the current GAP catalog; 8 admitted by calibration |
| Repetitions | 5 per reference setup in calibration; 3 per setup in the battery |
| Grading | decision correctness, weighted key-fact recall (≥ 0.7) and unsupported-claim rate (≤ 0.25), judged blind by the claude-code primary judge against captured primary sources; the codex judge measured agreement afterwards on the same payloads |
| Spend | about 13.5M agent tokens in the battery; calibration and judge usage are not totalled here |
Calibration
| Task | Family | claude-code |
|---|---|---|
| billing-reporting | research | Floor: 0, Ceiling: 5 — admitted |
| classroom-term | decision | Floor: 0, Ceiling: 5 — admitted |
| models-production | decision | Floor: 0, Ceiling: 5 — admitted |
| npm-token-scope | research | Floor: 0, Ceiling: 5 — admitted |
| org-migration | decision | Floor: 0, Ceiling: 4 — admitted |
| python-metadata | research | Floor: 0, Ceiling: 4 — admitted |
| python-security-march | research | Floor: 0, Ceiling: 5 — admitted |
| tarfile-upload | decision | Floor: 0, Ceiling: 5 — admitted |
| cli-trust | decision | Floor: 0, Ceiling: 3 — rejected |
| spark-editing | decision | Floor: 0, Ceiling: 1 — rejected |
Every floor scored 0. Two briefs were rejected for failing the ceiling, as in the 2026-10-03 calibration: cli-trust (ceiling 3/5) and spark-editing (ceiling 1/5). python-security-march is admitted this time (ceiling 5/5); in the 2026-10-03 calibration its answer excerpt omitted the date its top fact requires, and the catalog has since fixed the excerpt.
Gap closure
claude-code (8 tasks, 24 runs per setup)
| Setup | Gap closure (95% CI) | Pass rate (Wilson 95% CI) | Tokens per success | vs floor (exact McNemar) |
|---|---|---|---|---|
| Firecrawl | 0.91 (0.80–1.00) | 83% (64–93%) | 141k | p < 0.0001 |
| Exa | 0.86 (0.58–1.11) | 79% (60–91%) | 89k | p < 0.0001 |
| Built-in (native) | 0.82 (0.60–0.96) | 75% (55–88%) | 55k | p < 0.0001 |
| Parallel | 0.82 (0.54–1.17) | 75% (55–88%) | 90k | p < 0.0001 |
| Perplexity | 0.82 (0.64–0.96) | 75% (55–88%) | 72k | p < 0.0001 |
| Brave | 0.77 (0.50–1.06) | 71% (51–85%) | 105k | p < 0.0001 |
| Tavily | 0.68 (0.33–0.96) | 62% (43–79%) | 126k | p < 0.0001 |
| ceiling (reference) | 1.00 | 92% (74–98%) | 5k | |
| floor (reference) | 0.00 | 0% (0–14%) | — |
By family:
- Research briefs: Firecrawl, Perplexity and Exa 92%; Parallel and Built-in 83%; Brave and Tavily 75%.
- Decision briefs: Firecrawl 75%; Parallel, Brave, Exa and Built-in 67%; Perplexity 58%; Tavily 50%.
Against the 2026-10-03 battery
The task set differs by one brief (python-security-march), so the fair comparison is on the seven briefs both batteries admitted. The 2026-10-03 built-in row is its 2026-10-06 rerun with page fetching working.
| Setup | 2026-10-03 battery (7 tasks) | Replication, same 7 tasks | Fisher exact | Replication, all 8 tasks |
|---|---|---|---|---|
| Parallel | 20/21 (95%) | 16/21 (76%) | p = 0.18 | 18/24 (75%) |
| Firecrawl | 19/21 (90%) | 18/21 (86%) | p = 1.00 | 20/24 (83%) |
| Brave | 17/21 (81%) | 16/21 (76%) | p = 1.00 | 17/24 (71%) |
| Built-in (native) | 16/21 (76%) | 17/21 (81%) | p = 1.00 | 18/24 (75%) |
| Perplexity | 16/21 (76%) | 16/21 (76%) | p = 1.00 | 18/24 (75%) |
| Exa | 15/21 (71%) | 17/21 (81%) | p = 0.72 | 19/24 (79%) |
| Tavily | 15/21 (71%) | 15/21 (71%) | p = 1.00 | 15/24 (62%) |
| ceiling (reference) | 20/21 (95%) | 19/21 (90%) | p = 1.00 | 22/24 (92%) |
On the shared briefs no setup's change between batteries is statistically distinguishable (two-sided Fisher exact, smallest p = 0.18); Parallel's drop from 20/21 to 16/21 is the largest. With 21 runs per setup, neighbouring providers cannot be ranked against each other, and a single battery's ordering should not be quoted as a ranking. On the new brief, python-security-march, the setups passed: Exa, Parallel, Firecrawl and Perplexity 2/3 each; Brave and built-in 1/3 each; Tavily 0/3 (ceiling 3/3).
Judge agreement
On 2026-10-07 the codex judge scored every judged brief from the payload the primary judge had scored. Each payload was rebuilt from the run's stored source snapshots, so nothing was captured again, and codex was called only after the rebuilt payload's SHA-256 matched the primary's and the stored labels reproduced the stored verdict. The primary was not rerun. 5 briefs failed schema validation and were never judged; they fail under either judge.
| Setup | Briefs judged | Claude-only passes | Codex-only passes | Label agreement | Kappa | Pass, Claude judge | Pass, codex judge |
|---|---|---|---|---|---|---|---|
| Firecrawl | 24 | 0 | 1 | 98% | 0.82 | 20/24 | 21/24 |
| Exa | 24 | 0 | 1 | 95% | 0.72 | 19/24 | 20/24 |
| Built-in (native) | 24 | 0 | 0 | 97% | 0.85 | 18/24 | 18/24 |
| Parallel | 23 | 1 | 0 | 95% | 0.70 | 18/24 | 17/24 |
| Perplexity | 24 | 0 | 0 | 97% | 0.81 | 18/24 | 18/24 |
| Brave | 24 | 0 | 0 | 97% | 0.74 | 17/24 | 17/24 |
| Tavily | 24 | 0 | 0 | 94% | 0.71 | 15/24 | 15/24 |
| ceiling (reference) | 24 | 1 | 0 | 98% | 0.59 | 22/24 | 21/24 |
| floor (reference) | 20 | 0 | 0 | 93% | 0.86 | 0/24 | 0/24 |
| All setups | 211 | 2 | 2 | 96% | 0.82 | 147/216 | 147/216 |
A Claude-only pass is a disputed verdict the Claude judge passed and codex failed; a codex-only pass is the reverse. Label agreement pools the fact and claim labels both judges scored, and kappa is Cohen's kappa on them; claims whose cited source could not be captured are left out, as no judge scored them.
The four disputed verdicts split evenly: the Claude judge alone passed one ceiling and one Parallel brief, and codex alone passed one Exa and one Firecrawl brief. Had codex decided, the battery's pass count would be the same, 147 of 216, and no setup would move by more than one run. Every headline above holds under either judge: the floor passes nothing, every search setup differs from it (exact McNemar, p ≤ 0.0001), Firecrawl still leads at 21/24, and no search setup differs significantly from built-in search (largest gap: Firecrawl, p = 0.25). Agreement is in line with the 2026-10-03 battery on Claude Code's briefs (96% of labels, kappa 0.79). There, 4 of the 5 disputes on Claude Code's briefs went the Claude judge's way; here that lean did not repeat.
What this shows
1. Search changes the outcome, again. Without search the agent failed every brief; with any provider it passed most of them. This finding replicated exactly. 2. Provider rankings did not replicate. The top setup changed and the spread between search setups is within sampling noise at this size. Choose a provider on the agent you use, on cost and on integration, and measure. 3. Built-in search is a serious baseline once it can read pages. Level with the middle of the field at the lowest token cost per success. 4. Pass rate and cost remain separate questions. The cheapest setup per success and the highest-passing setup were different setups in both batteries.
Limits
- One agent and model; the codex half of the 2026-10-03 design was not repeated.
- One judge decides the verdicts. The second judge scored the stored payloads after the battery rather than alongside it; the evidence it saw is identical, checked by payload hash.
- 24 runs per setup: 95% intervals span roughly 30 points.
- Provider indexes, sources and models drift between batteries; a replication tests the method, not a fixed truth.
Methodology for this study
GAP replication methodology
This battery repeats the 2026-10-03 GAP methodology on one agent: calibration per task (five floor and five ceiling repetitions; admission requires the floor to pass at most once and the ceiling at least four times), then every admitted task three times on nine setups.
Differences from the 2026-10-03 battery:
- Catalog. The current GAP catalog (SHA-256
29f8fd1731cd0155f2a993688d8f1dfac7ad6ef1f6effc549185af026ad24443), in which python-security-march's answer excerpt carries the date its top fact needs. 8 of the 10 briefs were admitted. - Agent. Claude Code on claude-opus-5-5 only.
- Built-in search. The native setup pre-approves WebFetch as well as WebSearch, so the agent can open pages.
- Provider tools. Every provider setup made provider calls in every run; Parallel's credential is passed as a header file the workbench writes per run.
- Judging. The claude-code primary judge decides every verdict, as before. The codex judge, which measures agreement, was out of quota during the battery. It scored the same stored payloads on 2026-10-07 with
sew gap agree, which rebuilds each payload from the stored source snapshots and calls codex only when the payload's SHA-256 matches the primary's and the stored labels reproduce the verdict.
Statistics are unchanged: Wilson 95% intervals on pass rates, a seeded paired task bootstrap (10,000 resamples) for gap closure, exact McNemar tests against the floor and the built-in setup, and tokens per success counted over all agent tokens in the setup, failures included. The comparison with the 2026-10-03 battery uses two-sided Fisher exact tests on the seven shared briefs, because the two batteries' runs are not paired.
See recorded results, summary, admission counts and reproduction.
Search API head-to-head (2026-09-26)
Full recorded report
Search API head-to-head: Exa, Tavily, Parallel and Firecrawl on a research workload (2026-09-26)
This study called four search APIs directly, not through a coding agent. It asked how much verified new information each adds to a research knowledge base that had already been built with an agent's built-in search tools until they stopped finding new facts, and what each provider surfaces that the others don't. It was pre-registered and cost-capped, results were pooled and graded blind, and every counted claim was verified on its source page. Stages A and B ran on 2026-09-26, Stage C on 2026-09-27, and an Exa-only Monitors probe ran daily from 2026-09-27 to 2026-10-06. The full record (design, pre-registration with twelve amendments, judge briefs and stage write-ups), the code and the sanitized data are published in this directory.
Context. The experiment was run while building a go-to-market knowledge base about Exa, one of the four providers tested. Its questions come from that knowledge base's open questions about Exa's market and competitors, so the workload reflects what that knowledge base needed. The design added workloads where Exa might be weak on purpose (freshness, and pages that need rendering). Weigh the results with that in mind; the record, code and data are here so readers can check them.
Headline
- On top of built-in search, Exa added the most verified new information. On 40 questions the knowledge base couldn't answer, verified new or correcting claims under strict credit: Exa 25, Parallel 19, Tavily 11, Firecrawl 10, with 11 of Exa's found by no other provider. On two weeks of company news: Exa 20 new dated facts, Tavily 12, Parallel 3, Firecrawl 2. Building three entity lists: Exa returned 28 entities new to the knowledge base (21 that no other provider returned), Firecrawl 19 and Parallel 7.
- The other providers led elsewhere. On pages the built-in tools couldn't read, Firecrawl recovered 14 of 15, Parallel 12, Tavily 10 and Exa 5. Filling a 20-company, eight-field schema, Parallel was right on 113 of 133 ground-truth cells, Exa on 105 and Firecrawl on 94; the Parallel–Exa gap is not significant, at the same price.
- Finding a source the knowledge base already cites was a tie: 16 to 19 of 20 for all four, on identical inputs.
- On evidence about four big grounding vendors (Stage C), no lead is significant. Verified new tier-1 claims across a search arm and a research arm, after the post-hoc corrections of Amendment 12: Exa 16, Parallel 14, Tavily 13, Firecrawl 9. Parallel's research agent was the most accurate: 0.6% of its checkable items were unsupported or false, against Exa's 4.6%, Firecrawl's 2.7% and Tavily's 18.1%.
- Exa's Monitors caught real competitor changes, with some noise. Over 49 daily runs on five competitors' pricing and changelog pages, they reported 20 changes: 14 confirmed, 5 false positives. 13 were real, dated pricing or product changes, 12 of them missing from the knowledge base or only partly recorded there.
- One provider for everything leaves gaps. For this workload the record recommends Exa for discovery, list building and news sweeps; Firecrawl or Parallel for pages that need rendering; and Parallel or Exa for filling a fixed schema.
What was tested
Each provider ran with parameters taken from its own best-practice documentation and frozen before the first call, on the same day, with the same top 10. The design is a capability matrix rather than four copies of every job: Stage A's cheap probes ran on all four providers, and Stage B's list and schema sets ran only on providers that sell those capabilities. Tavily sat out S5 and S6 because it sells no list-building or structured-extraction product. See methodology for the endpoints, grading and amendments.
| Set | Stage | Size | What it tests | Providers |
|---|---|---|---|---|
| S1 known unknowns | A | 40 questions | Can a provider answer what the built-in research couldn't? | All four |
| S2 known knowns | A | 20 claims | Is the primary source the knowledge base already cites in the top 10? | All four |
| S3 blocked sources | A | 15 pages | Can it read pages the built-in tools couldn't (JavaScript, rate limits, PDFs)? | All four |
| S4 freshness | A | 10 companies | New, dated facts from the last 14 days | All four |
| S5 entity lists | B | 3 lists | Precision, and entities new to the knowledge base | Exa, Parallel, Firecrawl |
| S6 schema fill | B | 20 companies × 8 fields | Accuracy and completeness on a fixed schema | Exa, Parallel, Firecrawl |
| SP find-similar | B | 10 homepages | New competitor domains | Exa |
| SM Monitors | B | 5 monitors, daily | Does a monitor surface real, dated pricing or product changes? | Exa |
| SV big vendors' evidence | C | 24 questions × 2 arms | Verified new evidence on four grounding vendors, by search and by research agent | All four |
Stage A: four providers, cheap probes
Strict credit counts a claim only when the provider's own result page states it (Amendment 4). S2 was re-run with identical inputs after a review found that hand-written queries leaked answer terms. The Stage A write-up has broad credit, the first grading round and the diagnostics.
| Stage A (v2 re-grade, strict credit) | Exa | Tavily | Parallel | Firecrawl |
|---|---|---|---|---|
| S1 new or correcting claims, strict (known unknowns, 40 questions) | 25 | 11 | 19 | 10 |
| … no other provider supported | 11 | 2 | 4 | 0 |
| S4 new dated facts, strict (10 companies, two weeks) | 20 | 12 | 3 | 2 |
| S3 pages recovered that native tools couldn't read (of 15) | 5 | 10 | 12 | 14 |
| S2 cited sources found in the top 10, identical inputs (of 20) | 18 | 16 | 19 | 17 |
| Stage A run spend | $0.50 | 146 credits ($1.17) | $0.37 | 402 credits ($0.30–$2.01) |
| Median latency, ms | 1,230 | 2,995 | 2,343 | 1,171 |
Stage B: lists, schema fill and find-similar
S5 counts entities verified against the list's criteria and new to the knowledge base, under the fact-level rule; the stricter entity-level rule gives Exa 27, Firecrawl 18 and Parallel 3, in the same order. S6 compares cells with the knowledge base's verified ground truth. Parallel's spend is at list price; Firecrawl's credits are shown as a range across plans. The Stage B write-up has each list, each field and the review fixes (Amendments 6 and 7).
| Stage B | Exa | Parallel | Firecrawl |
|---|---|---|---|
| S5 entities verified and new to the KB (3 lists, fact-level rule) | 28 | 7 | 19 |
| … no other provider returned it | 21 | 5 | 12 |
| S5 precision: entities meeting every criterion | 28/36 (77.8%) | 13/21 (61.9%) | 19/23 (82.6%) |
| S6 right on KB ground truth, strict (of 133 cells) | 105/133 (78.9%) | 113/133 (85.0%) | 94/133 (70.7%) |
| S6 cells filled (of 160) | 149/160 (93.1%) | 156/160 (97.5%) | 157/160 (98.1%) |
| S6 false-claim rate | 7/149 (4.7%) | 6/156 (3.8%) | 16/157 (10.2%) |
| S6 cells where it added a verified new fact (of 40) | 30 | 31 | 25 |
| Stage B spend | $4.09 | $12.16 (list price) | 393 credits ($0.29–$1.97), plus 5 free runs |
Find-similar (SP, Exa only) is a cheap sweep for long-tail competitors, but it needs triage: 50 of the 61 new domains were directories, profiles, clones or unrelated sites. It was graded, but not blind, since only Exa ran.
| Exa find-similar, 10 competitor homepages | Count |
|---|---|
| Results returned | 100 |
| Distinct registrable domains | 69 |
| … already cited in the KB | 8 |
| … new to the KB | 61 |
| New domains: relevant vendors (search, SERP, crawling, or web-data APIs) | 11 |
| New domains: profiles, directories, or reviews of a seed company | 39 |
| New domains: clones or unrelated | 11 |
| Relevant vendors the KB baseline already covers | 0 |
| Spend | $0.07 |
Stage C: evidence on the big grounding vendors
Twenty-four pre-registered questions asked for evidence on four supply lines a market sizing couldn't anchor: Google grounding, Anthropic's and OpenAI's web search tools, and Microsoft Grounding with Bing. Each provider ran a search arm and a research-agent arm, and eight blind judges verified every claim on its page. A tier-1 claim is a figure that enters a line's arithmetic: a volume, a revenue figure or share, a denominator, or a price. After a review, Amendment 12 (post hoc) added a claim-level correction layer for novelty, tiers, duplicates and Tavily's cap-refused runs; the Stage C write-up shows the as-judged figures beside the corrected ones.
| Provider | New tier-1 claims (either arm) | …found by no other provider | All new claims | Cost, both arms | Cost per new tier-1 claim |
|---|---|---|---|---|---|
| Exa | 16 | 6 | 207 | $2.57 | $0.16 |
| Parallel | 14 | 4 | 131 | $2.52 | $0.18 |
| Tavily | 13 | 4 | 79 | $3.96–$6.34 | $0.30–$0.49 |
| Firecrawl | 9 | 0 | 126 | $1.16–$7.75 | $0.13–$0.86 |
By arm, corrected. The research error rate is the share of a research agent's checkable items whose cited page didn't state them or contradicted them.
| Provider | Arm | New tier-1 claims | …found by no other arm | All new claims | Questions with a new claim (of 24) | Useful new tier-1 (2–3) | Research error rate | Median time |
|---|---|---|---|---|---|---|---|---|
| Exa | search | 11 | 2 | 100 | 20 | 6 | — | 1.7 s |
| Parallel | search | 8 | 3 | 75 | 20 | 6 | — | 3.1 s |
| Tavily | search | 11 | 3 | 58 | 21 | 5 | — | 3.2 s |
| Firecrawl | search | 7 | 0 | 57 | 18 | 3 | — | 1.3 s |
| Exa | research | 11 | 2 | 158 | 24 | 8 | 4.6% (21 of 455) | 86.1 s |
| Parallel | research | 10 | 1 | 84 | 17 | 6 | 0.6% (2 of 331) | 107.0 s |
| Tavily | research | 4 | 1 | 36 | 13 | 1 | 18.1% (19 of 105) | 31.4 s |
| Firecrawl | research | 5 | 0 | 85 | 18 | 3 | 2.7% (8 of 298) | 71.7 s |
Is the lead real? Exact two-sided sign tests over the 24 questions, chosen after the results were known, with Holm's adjustment for the six comparisons and question-level bootstraps:
| Arm | Leader | Against | New tier-1 (leader / other) | Questions ahead / behind | Sign test p | Holm-adjusted p (6 tests) | Bootstrap 95% CI of the difference |
|---|---|---|---|---|---|---|---|
| search | Exa | Tavily | 11 / 11 | 3 / 3 | 1.000 | 1.000 | -7 to 7 |
| search | Exa | Parallel | 11 / 8 | 3 / 2 | 1.000 | 1.000 | -5 to 12 |
| search | Exa | Firecrawl | 11 / 7 | 4 / 1 | 0.375 | 1.000 | -1 to 10 |
| research | Exa | Parallel | 11 / 10 | 4 / 2 | 0.688 | 1.000 | -5 to 6 |
| research | Exa | Firecrawl | 11 / 5 | 5 / 1 | 0.219 | 1.000 | 0 to 13 |
| research | Exa | Tavily | 11 / 4 | 5 / 1 | 0.219 | 1.000 | 0 to 15 |
Monitors probe (SM)
Five Exa Monitors watched competitors' pricing and changelog pages once a day. The pre-registered question: does a monitor surface a real, dated pricing or product change on the competitor's own page within the knowledge base's 30-day staleness window, and was the change already in the knowledge base?
- Deviation from the plan. The probe was planned to run to 2026-10-10 and then delete its monitors. It was harvested and the monitors paused on 2026-10-06, four days early, so it covers 49 daily runs rather than about 70. The monitors were paused, not deleted.
- What the monitors reported. 14 of 49 runs reported at least one change, 20 changes in all. Monitors report only what changed since their previous run.
- Graded result. Of the 20 reported changes, 14 were confirmed against the vendor's page and archived snapshots, 5 were false positives and 1 could not be settled either way. 13 were real, dated pricing or product changes, and 12 of those were missing from the knowledge base or only partly recorded there. 3 of 5 monitors answered the pre-registered question yes: Parallel, Firecrawl and Perplexity API.
- Where it worked and where it didn't. Perplexity's monitor was the strongest: its changelog carries exact timestamps, and every confirmed change fell inside its run window. Firecrawl's monitor reported a security feature added, removed and added again on its pricing page; the page was unchanged in every archived capture, so those reports are noise. Both of Brave's reports were false positives: prices were unchanged, and the watched pages had started serving a version without prices. Tavily's monitor reported nothing, and its changelog had nothing new.
- Misses were not measured. The grader noticed one: a Firecrawl changelog entry from 2026-09-29 was never reported. Per-change verdicts, evidence links and notes are in the grading record and its write-up.
- How it was graded. One Claude agent, not blind (only Exa ran), checked each reported change against the vendor's current page and Internet Archive snapshots from before and after the run, and searched the knowledge base for it.
| Monitor (competitor pages watched) | Daily runs | Runs reporting a change | Changes reported | Confirmed real | Real, dated pricing or product change | …missing from the KB or partial |
|---|---|---|---|---|---|---|
| Tavily | 10 | 0 | 0 | 0 | no (0) | no (0) |
| Parallel | 10 | 3 | 4 | 4 | yes (3) | yes (2) |
| Firecrawl | 10 | 4 | 6 | 3 | yes (3) | yes (3) |
| Perplexity API | 9 | 6 | 8 | 7 | yes (7) | yes (7) |
| Brave Search API | 10 | 1 | 2 | 0 | no (0) | no (0) |
Spend
Stages A and B together cost at most $23.29 of the $40 cap, valuing credits at the most they cost (Stage A $5.07, Stage B $18.22). Stage C's costs are in its provider table above; Exa's are billed, Parallel's are list-price estimates and Tavily's and Firecrawl's credits are ranges across plans. The Monitors probe made 49 runs, at $15 per 1,000 runs on Exa's price list.
Read this before comparing providers
- Small samples, one run. Each item ran once on one day. The web and the indexes change, so a re-run will differ.
- Model judges. Claude agents graded blind to the provider and verified every counted claim on its page; no human graded the sets.
- Post-hoc changes are logged, not hidden. Amendments 4, 6, 7 and 12 changed scoring after results were seen. The record keeps the earlier figures beside the corrected ones, and the pre-registration lists every amendment.
- New means new to one knowledge base. Novelty is measured against the knowledge base the questions came from.
- Direct APIs, tuned parameters. Unlike the agent benchmarks, where each provider ran through its MCP server at default settings, this study called each API directly with its best-practice parameters. Results do not transfer between the two setups.
- Vendor terms. The experiment was first run as an internal evaluation. The record's overview summarises each vendor's terms on benchmarking as read on 2026-09-26 (terms).
- What is and isn't published. For Stages A to C, page text, titles, snippets, vendor-written values and raw responses are not published, and social-media URLs are hashed. Those raw responses were deleted after grading (Amendment 7), so their grading cannot be replayed exactly. Every published table rebuilds from
data/except Stage A's S3 rows, which need page text; their per-page outcomes are in the scripted summary. The Monitors probe is the exception: by the owner's decision of 2026-10-06, its raw API responses, including Exa's change titles, summaries and run notes, are published verbatim in data/stage-b/runs/SM/raw/ (each monitor's webhook secret was never stored), and its sanitized record rebuilds from them byte for byte.
Record and data
- The experiment's overview, design and pre-registration with Amendments 1–12.
- Judge briefs: first round, re-grade, Stage B and Stage C.
- Stage write-ups: Stage A, Stage B and Stage C, each with every table and the verified findings.
- Code: the Stage A harness code/harness.py, Stage B code/harness_b.py, Stage C code/stage_c/harness_research.py, the exporter code/export_data.py and the Monitors harvest code/monitors_harvest.py.
- Scores: Stage A, Stage B, Stage C corrected and as judged; the Monitors harvest, raw responses and grading.
Methodology for this study
Search API head-to-head: methodology
This study calls search APIs directly; no coding agent is involved. Its design, query sets, parameters and metrics were frozen in the design and the pre-registration before the first paid call, with SHA-256 hashes of the frozen files. Twelve amendments follow, each dated and marked as made before or after results were seen. This page summarises the method; the record is authoritative.
Providers and endpoints
Every provider's parameters come from its own best-practice documentation, read on 2026-09-26, and are written into the harness configs (code/config.json, code/config_b.json, code/stage_c/config.json and config_c2.json). Every provider got the same top 10 and ran on the same day.
- Exa:
/search(auto, highlights) for S1 and S2;/contentswith a live crawl for S3; the news category with a date filter for S4; the Agent API for S5 (capped) and S6 (with an output schema); find-similar for SP; Monitors for SM; and the Agent API atmediumeffort for Stage C's research arm. - Tavily: search (
advanced, chunks); extract for S3; news with a time range for S4; the Research API (mini) for Stage C's research arm. Tavily sat out S5 and S6: it sells no list-building or structured-extraction product, and search plus our own model would measure our model, not Tavily. - Parallel: search (
advanced, with an objective); extract for S3;after_datefor S4; FindAll for S5; Task with an output schema for S6; Task (pro) for Stage C's research arm. - Firecrawl: search plus a scrape of the top 3; scrape with JavaScript rendering for S3;
tbsfor S4;/agentfor S5 and S6 (with a schema);/agentcapped at 100 credits for Stage C's research arm.
Sets
The query sets are drawn from the knowledge base's open questions and claims (see the design): S1 40 known unknowns, S2 20 known knowns, S3 15 blocked pages, S4 10 companies' last two weeks, S5 three entity lists, S6 20 companies by eight fields, SP 10 competitor homepages, SM 5 monitors and SV 24 questions on four grounding vendors. Their hashes are pinned in code/sets/SHA256SUMS and match the prefixes frozen in the pre-registration.
Grading
- Pooled and blind. Each question's results are pooled across providers, deduplicated and, from Amendment 4, shuffled with a recorded seed. Judges never see provider names; attribution is rejoined only after judging.
- Verified on the page. Judges are Claude agents working from written briefs (first round, re-grade, Stage B, Stage C). Every counted claim must be stated on its source page.
- Strict per-claim credit (Amendment 4): a provider is credited with a claim only when its own result page states it. Broad credit, where any pooled page may support the claim, is reported in the stage write-ups.
- Novelty is judged against the knowledge base at a recorded baseline commit. Claims are weighted for usefulness: 3 corrects a knowledge-base claim, 2 is a new fact for a page, 1 is a new corroborating source and 0 is trivia.
- Scripted sets. S2 counts whether the cited primary source is in the top 10. S3 counts a page as recovered when the fetch returns at least 500 characters and a target keyword, with one recorded correction (a redirect).
- S5 and S6. An S5 entity counts when it meets every criterion of its list and is new to the knowledge base (a fact-level and a stricter entity-level rule are both reported). S6 cells are compared with the knowledge base's verified ground truth where it has one; the false-claim rate is the share of filled cells that are wrong.
- Stage C. Eight blind judges. A tier-1 claim enters a sizing line's arithmetic. The research error rate is the share of a research agent's checkable items whose cited page didn't state or contradicted them. Amendment 12 (post hoc) adds a claim-level correction layer (
data/stage-c/corrections.json): novelty against the knowledge base's working notes, tiers, duplicates across questions, and scoring Tavily's four cap-refused runs as returning nothing, as pre-registered. Sign tests, Holm's adjustment and bootstraps were chosen after the results were known. - Monitors (SM). One Claude agent, not blind (only Exa ran), checked each reported change against the vendor's current page and Internet Archive snapshots from before and after the run, and searched the knowledge base for it. A monitor passes when at least one change is confirmed, is a pricing or product change, and is dated inside the 30 days before the run.
Gates, caps and stop rules
S1 ran on 20 questions first and would have stopped there had no provider resolved at least 3 that the built-in tools couldn't. Stage B ran only for providers and workloads that showed signal in Stage A. Each provider had a hard cap, and the owner set a $40 cap for Stages A and B together (Amendment 5); Stage C had its own caps (Amendment 8).
Amendments that changed scoring after results were seen
- Amendment 4: after a review, Stage A was re-graded with shuffled pools and strict per-claim credit, and S2 was re-run on identical inputs.
- Amendment 6: Stage B scoring choices, written after unblinding.
- Amendment 7: fixes from the Stage B review, and the per-file allowlist for exported data.
- Amendment 12: the Stage C correction layer described above.
The record keeps the earlier figures beside the corrected ones in each case.
Costs
Exa's costs are billed. Parallel's are list-price estimates from the harness config. Tavily's and Firecrawl's credits are shown as ranges across their plans. Stage A and B costs value credits at the most they cost, as the cap guard did.
Published data
code/export_data.py writes the data through an allowlist. It keeps request parameters, status, latency, cost, each result's rank, URL, domain, published date and text length, crawl status codes, the pools (URLs only), attribution maps, judge verdicts and scores. It never exports page titles, snippets, page text, extracted content, vendor-written values or raw responses, and replaces the paths of social-media profile and post URLs with a stable hash. Judge notes may quote short fragments of pages. The Monitors export (code/monitors_harvest.py) follows the same rule: change types and fingerprints are kept, change summaries and titles only as lengths. The Monitors probe is the one exception to the no-vendor-text rule: by the owner's decision of 2026-10-06, its raw API responses, including Exa's change titles, summaries and run notes, are published verbatim in data/stage-b/runs/SM/raw/.
See the report, summary data and reproduction. Calibration does not apply: calibration record.