Results to date
Two benchmarks, two coding-agent harnesses, eight retrieval arms,
567 graded provider and native cells in the
leaderboard below, and a transcript analysis of how the agents searched.
Every number on this page is generated from the published report data in reports/.
scripts/build_site.py from reports/*/summary.json. The same image is
site/infographic.svg.Read this before comparing vendors
- Small samples18 to 42 cells per arm. Intervals overlap widely; adjacent positions are not rankings, and no bakeoff arm differs from answering without search at 95% confidence.
- One interface per vendorEach vendor ran through its own MCP server at a pinned version with default settings. Direct APIs, other search modes and other tools were not tested.
- The agent mattersRankings changed between Claude Code and codex. Parallel led on one harness and tied last on the other. A result on one agent does not transfer to another.
- Native arm defectClaude Code's native arm could not fetch pages: the harness refused WebFetch because only WebSearch was pre-approved. Its rows are flagged and not comparable. Fixed for future runs.
- Cost coverageGAP tokens per success counts all agent tokens, including failures. WSB excludes ungraded and unmeasured cells and lacks per-arm coverage counts. Vendor dollar spend was metered only partly (off on the gap bench), so dollar comparisons are omitted here.
- Who gradedDeterministic validators where possible; otherwise blinded model judges against captured sources (Claude primary, codex second judge on the gap bench). Judges never saw which arm produced an answer.
Leaderboard
Ordered by observed pass rate. The bar is the 95% interval; the tick is the observed rate. Token accounting and coverage are
defined beside each benchmark's table. The leaderboard updates automatically when a new report lands in
reports/; machine-readable data is in leaderboard.json.
Search gap bench: job outcomes on Claude Code (claude-opus-5-5)
7 admitted post-cutoff brief tasks × 3 repetitions = 21 cells per arm. Pass = right decision, weighted key-fact recall ≥ 0.7, unsupported claims ≤ 0.25. Interval: Wilson 95% (published). The dashed line marks the lower end of the top arm's interval; badges describe overlap of individual intervals only. Overlap does not establish statistical equivalence or test a difference between arms.
| Arm | Passed | 95% interval on a 0–100% scale | Gap closure | Tokens per success | Notes |
|---|---|---|---|---|---|
| Parallel | 20/21 · 95% 77–99% | 1.00 (0.86–1.17) | 64k | overlaps top interval | |
| Firecrawl | 19/21 · 90% 71–97% | 0.95 (0.83–1.00) | 116k | overlaps top interval | |
| Brave | 17/21 · 81% 60–92% | 0.85 (0.62–1.00) | 93k | overlaps top interval | |
| Perplexity | 16/21 · 76% 55–89% | 0.80 (0.50–1.00) | 74k | overlaps top interval | |
| Exa | 15/21 · 71% 50–86% | 0.75 (0.47–0.95) | 89k | overlaps top interval | |
| Tavily | 15/21 · 71% 50–86% | 0.75 (0.43–1.00) | 83k | overlaps top interval | |
| native | 13/21 · 62% 41–79% | 0.65 (0.37–0.91) | 66k | overlaps top interval WebFetch refused by the harness (search only); not comparable |
floor (no search): 0% · ceiling (answer excerpt): 95%
GAP tokens per success counts all agent tokens in the arm, failed cells included, divided by passing cells.
Source: Search gap bench.
Search gap bench: job outcomes on codex (gpt-6.1-sol)
6 admitted post-cutoff brief tasks × 3 repetitions = 18 cells per arm. Pass = right decision, weighted key-fact recall ≥ 0.7, unsupported claims ≤ 0.25. Interval: Wilson 95% (published). The dashed line marks the lower end of the top arm's interval; badges describe overlap of individual intervals only. Overlap does not establish statistical equivalence or test a difference between arms.
| Arm | Passed | 95% interval on a 0–100% scale | Gap closure | Tokens per success | Notes |
|---|---|---|---|---|---|
| Brave | 17/18 · 94% 74–99% | 1.06 (0.88–1.29) | 111k | overlaps top interval | |
| Firecrawl | 16/18 · 89% 67–97% | 1.00 (1.00–1.00) | 166k | overlaps top interval | |
| Tavily | 16/18 · 89% 67–97% | 1.00 (0.71–1.29) | 129k | overlaps top interval | |
| Exa | 14/18 · 78% 55–91% | 0.88 (0.47–1.29) | 146k | overlaps top interval | |
| native | 14/18 · 78% 55–91% | 0.88 (0.47–1.20) | 117k | overlaps top interval | |
| Parallel | 13/18 · 72% 49–88% | 0.81 (0.41–1.14) | 157k | overlaps top interval | |
| Perplexity | 13/18 · 72% 49–88% | 0.81 (0.47–1.12) | 117k | overlaps top interval |
floor (no search): 0% (0–18%) · ceiling (answer excerpt): 89% (67–97%)
GAP tokens per success counts all agent tokens in the arm, failed cells included, divided by passing cells.
Source: Search gap bench.
Web search bakeoff: competitive research tasks on Claude Code (claude-opus-5-5)
14 competitive tasks × 3 repetitions = 42 cells per arm, regraded 2026-09-30. Mostly stable, documented knowledge; search adds less here than on post-cutoff jobs. Interval: Wilson 95% (computed from the published counts). The dashed line marks the lower end of the top arm's interval; badges describe overlap of individual intervals only. Overlap does not establish statistical equivalence or test a difference between arms.
| Arm | Passed | 95% interval on a 0–100% scale | Tokens per success | Notes |
|---|---|---|---|---|
| Firecrawl | 38/42 · 90% 78–96% | 61.5k | overlaps top interval | |
| Perplexity | 37/42 · 88% 75–95% | 47.7k | overlaps top interval | |
| Tavily | 36/42 · 86% 72–93% | 61.3k | overlaps top interval | |
| Exa | 35/42 · 83% 69–92% | 65.0k | overlaps top interval | |
| Parallel | 33/42 · 79% 64–88% | 69.8k | overlaps top interval | |
| native | 33/42 · 79% 64–88% | 36.3k | overlaps top interval WebFetch refused in 40 of 54 cells | |
| no-search | 32/42 · 76% 61–87% | 20.9k | reference | |
| Brave | 31/42 · 74% 59–85% | 89.2k | overlaps top interval |
WSB tokens per success uses measured tokens from graded cells divided by measured successes. Ungraded and unmeasured cells are excluded; per-arm coverage counts were not retained, so complete accounting cannot be established.
Source: Web search bakeoff.
Test apparatus
Each cell is one task, one arm and one repetition. The runner gives the agent a fresh scratch workspace, starts the harness with exactly the tools its arm allows, records every tool call and the final answer into an evidence bundle, and grades the answer without telling the grader which arm produced it.
From catalog to report
flowchart LR
subgraph catalog[Task catalogs]
T1[production tasks<br/>web search bakeoff]
T2[gap briefs<br/>post-cutoff events]
end
subgraph cell[One cell = task x arm x repetition]
W[fresh scratch<br/>workspace]
H[harness<br/>Claude Code or codex]
M[metering proxy]
V[(vendor MCP server<br/>pinned version)]
end
subgraph grade[Grading]
D[deterministic<br/>validators]
J[blinded judges<br/>vs captured sources]
end
T1 & T2 --> W --> H
H <-->|tool calls| M <--> V
H --> B[(evidence bundle<br/>transcript, final answer,<br/>arm audit, metrics)]
B --> D & J --> R[reports/*/summary.json] --> S[this site]
What each arm can touch
An arm differs from another only in its retrieval surface. Claude Code is started with --strict-mcp-config, so the
only MCP server it can reach is the arm's vendor server, and --tools limits its built-ins. Codex runs read-only with its
shell, browser, app and plugin features disabled. After each cell an arm audit compares every observed tool call with the arm
contract; a call outside the contract marks the cell contaminated.
flowchart TB A[arm contract] --> CC[Claude Code spawn<br/>--strict-mcp-config<br/>--tools: the arm's built-ins<br/>--allowedTools: the same set] A --> CX[codex spawn<br/>exec --sandbox read-only<br/>--disable shell_tool, browser_use,<br/>apps, plugins, ...] CC --> P1[provider arm:<br/>one vendor MCP server<br/>+ ToolSearch + Read] CC --> N1[native arm:<br/>WebSearch + WebFetch + Read] CC --> F1[no-search / floor / ceiling:<br/>no tools] CX --> P2[provider arm:<br/>one vendor MCP server] CX --> N2[native arm:<br/>codex --search] CX --> F2[no-search / floor / ceiling:<br/>no tools] P1 & N1 & F1 & P2 & N2 & F2 --> AU[arm audit:<br/>every observed tool call<br/>checked against the contract]
Claude Code invocation
claude --print --output-format stream-json --verbose --no-session-persistence \ --model claude-opus-5-5 --mcp-config <cell>/claude-mcp.json --strict-mcp-config \ --tools ToolSearch,Read --allowedTools 'mcp__<vendor>__*' # provider arm --tools WebSearch,WebFetch,Read --allowedTools WebSearch,WebFetch # native arm (fixed) --tools '' --allowedTools '' # no-search, floor, ceiling
codex invocation
codex [--search] exec --json --ephemeral --skip-git-repo-check --sandbox read-only \ --model gpt-6.1-sol --disable shell_tool --disable unified_exec \ --disable browser_use --disable browser_use_external --disable computer_use \ --disable in_app_browser --disable apps --disable plugins --disable remote_plugin ... # --search only on the native arm; provider arms add one MCP server via the cell config
Per-cell limits
| Limit | Gap bench | Web search bakeoff |
|---|---|---|
| Wall clock per cell | 1,200 s | 300 to 1,200 s (set per task) |
| Total tokens per cell | 1,000,000 | 120,000 to 350,000 (set per task) |
| Provider calls per cell | 60 | 15 to 50 (set per task) |
| Harness boot timeout | 120 s | 120 s |
| Repetitions | 3 (battery), 5 (calibration) | 3 |
How a gap-bench task is admitted
A task counts only if it is genuinely beyond the model: with no search the agent must fail, and with the answer excerpt handed over it must succeed. Calibration is repeated for each harness and model.
flowchart LR
T[candidate brief<br/>event after 2026-01-01] --> FL[floor: no search<br/>5 runs]
T --> CE[ceiling: answer excerpt<br/>in the prompt, 5 runs]
FL --> Q{floor passes ≤ 1<br/>and ceiling passes ≥ 4?}
CE --> Q
Q -- yes --> AD[admitted for this<br/>harness and model]
Q -- no --> RJ[rejected]
AD --> BAT[battery: every arm x 3]
BAT --> GC[gap closure =<br/>arm - floor / ceiling - floor]
gap closure = (arm pass rate − floor pass rate) / (ceiling pass rate − floor pass rate)
floor arm ceiling
0% ●─────────────────●──────────────────● 95%
└──── closed ────┘└──── remaining ────┘
How answers are graded
sequenceDiagram participant A as Agent answer participant V as Schema validator participant S as Source capture participant J1 as Judge 1 (Claude Code) participant J2 as Judge 2 (codex, gap bench) A->>V: JSON deliverable (brief, claims, citations) V-->>A: reject malformed answers A->>S: cited URLs S->>J1: readable source text + rubric (blinded to arm) S->>J2: same inputs J1-->>A: key-fact recall, unsupported claims, decision J2-->>A: independent labels, agreement recorded Note over A,J2: pass = right decision AND recall >= 0.7 AND unsupported <= 0.25
Arm configurations
Pinned in config/provider-arm-tools.yaml. Every tool a server advertised was available to the agent; no vendor tool was hidden.
| Arm | Interface | Version / endpoint | Tools exposed to the agent |
|---|---|---|---|
| Brave | stdio MCP server (npx) | @brave/brave-search-mcp-server@2.1.4 | 8: brave_image_search, brave_llm_context, brave_local_search, brave_news_search, brave_place_search, brave_summarizer, brave_video_search, brave_web_search |
| Exa | stdio MCP server (npx) | exa-mcp-server@3.4.1 | 2: web_fetch_exa, web_search_exa |
| Firecrawl | stdio MCP server (npx) | firecrawl-mcp@3.26.0 | 29: firecrawl_agent, firecrawl_agent_status, firecrawl_check_crawl_status, firecrawl_crawl, firecrawl_credit_usage, firecrawl_developer_search, firecrawl_extract, firecrawl_feedback, firecrawl_find_tools, firecrawl_interact, firecrawl_interact_stop, firecrawl_map, firecrawl_monitor_check, firecrawl_monitor_checks, firecrawl_monitor_create, firecrawl_monitor_delete, firecrawl_monitor_get, firecrawl_monitor_list, firecrawl_monitor_run, firecrawl_monitor_update, firecrawl_parse, firecrawl_research_inspect_paper, firecrawl_research_read_paper, firecrawl_research_related_papers, firecrawl_research_search_github, firecrawl_research_search_papers, firecrawl_scrape, firecrawl_search, firecrawl_search_feedback |
| Parallel | hosted MCP via mcp-remote | mcp-remote@0.14.3 → https://search.parallel.ai/mcp-oauth | 2: web_fetch, web_search |
| Perplexity | stdio MCP server (npx) | @perplexity-ai/mcp-server@1.3.0 | 4: perplexity_ask, perplexity_reason, perplexity_research, perplexity_search |
| Tavily | stdio MCP server (npx) | tavily-mcp@0.2.22 | 5: tavily_crawl, tavily_extract, tavily_map, tavily_research, tavily_search |
| native (Claude Code) | harness built-in | Claude Code 2.1.282 | WebSearch, WebFetch (refused before 2026-10-06), Read |
| native (codex) | harness built-in | codex --search | built-in web search and page views |
Runs
Each run is the recorded report, transcribed into summary.json and checked by scripts/check_reports.py.
Web search bakeoff (2026-09-29)
Full recorded report
WSB live battery — initial results (2026-09-29)
Correction (2026-10-06): the native arm's page fetches were mostly refused. The arm exposed WebSearch and WebFetch but pre-approved only WebSearch, and headless Claude Code refuses tools that are not pre-approved. WebFetch was refused in 40 of 54 native cells (110 calls). Provider arms were unaffected. The native rows therefore understate Claude Code's own web tools. The configuration is fixed and pinned by a test; the recorded numbers below are unchanged. Details: agent search behavior.
Pack: web-search-bakeoff (WSB) on the Search Evaluation Workbench (SEW). Harness / model: claude-code, Claude Opus 5.5 (claude-opus-5-5[1m]), authenticated execution on the original host. Status: historical initial grading; headline and Results tables are superseded by the 2026-09-30 regrade, where Firecrawl leads at 90%. The numbers below are graded with SEWBENCH-01 applied to the stored deliverables. The same deliverables graded without those fixes are summarised in Failure dive and were materially wrong.
Headline
Historical figures under the initial grader: rankings and cost lower bounds here apply only to that grading, not the later operator-ruled catalog.
On the 14 competitive production tasks (42 cells per arm), every search arm except Brave scores above answering without search:
- Firecrawl and Perplexity lead at 83%. Exa is at 81%, Tavily 79%, native web search and Parallel 76%, no-search 71% and Brave 69%.
- The 95% intervals overlap. At n=42 per arm the ordering is directional, not conclusive: no arm's difference from no-search or from native has an interval that excludes zero.
- Search matters most on tasks that need live data. On the cloud egress pricing table only Firecrawl and Perplexity went 3/3 (Brave 2/3, the rest 0/3). The AWS pricing pages need JavaScript rendering, and most arms reported the intra-AZ cell as unavailable, honestly.
- Among fully priced arms, native search costs least per success ($0.167), followed by Tavily ($0.206). Perplexity (at least $0.148; 37/42 cells priced) and Firecrawl (at least $0.159; 39/42 priced) are lower bounds, excluded from the cost ranking; Perplexity's
askspend is unmetered. Brave and no-search are also partially priced (at least $0.28 each); no-search had a few long runaway answers. - Brave is last, failing
unanswerable-nonexistent-postgres-gucandupstream-diagnosis-node-openssl3-md4in all three repetitions.
What ran
| Run | Arms | Cells | Window (UTC) |
|---|---|---|---|
| first arm group | no-search, native, brave, tavily | 216 | 2026-09-29 03:29 – 13:26 |
| second arm group | exa, parallel-web, firecrawl, perplexity | 216 | 2026-09-29 13:39 – 22:49 |
- Tasks: all 18 production tasks (production task catalog), × 3 repetitions.
- 14 are
competitive. - 4 are
expected_failby design: no retrieval surface can answer them, and the honest deliverable is a partial answer or a decline.
- 14 are
- Arms:
- Provider arms expose one vendor MCP server each, and nothing else.
- The native arm uses the harness's own web search.
- The no-search arm has no network tools.
- Pinned MCP servers:
- @brave/brave-search-mcp-server version 2.1.4
tavily-mcp@0.2.22exa-mcp-server@3.4.1firecrawl-mcp@3.26.0- @perplexity-ai/mcp-server version 1.3.0
- Parallel's hosted Search MCP (
https://search.parallel.ai/mcp-oauth, authenticated) viamcp-remote@0.14.3
- Grading:
- Deterministic validators where a task has one.
- Otherwise a blinded claude-code judge scoring the task's rubric.
bakeoff grade --regradere-scores the stored deliverables. Nothing was re-run.
- Cost: model tokens are priced at the published Opus 5.5 rates. Vendor spend is metered per MCP call where a tariff exists. It is unknown for Exa, Parallel and Perplexity's
asktool, which report no spend over MCP. In the tables,partialis the priced cells' spend divided by all of the arm's successes. Unpriced cells can only add spend, so eachpartialfigure is a lower bound on the full-arm $/success, and these arms are excluded from the cost ranking.bakeoff reportdivides by the successes among priced cells instead, which gives $0.172 for Perplexity, $0.163 for Firecrawl and $0.29 for Brave on the competitive set, and no figure for Exa or Parallel on the expected-fail set, where every success is in an unpriced cell.
Results
All Results tables below use the historical initial grading. The ungraded column counts delivered answers that could not be graded. Cells that exhausted their budget before delivering are counted under budget-exhausted and as failures, rather than under ungraded.
Competitive tasks (14 tasks × 3 repetitions)
| arm | pass | 95% CI | $/success | priced | tokens/cell | provider calls | budget-exhausted | ungraded |
|---|---|---|---|---|---|---|---|---|
| firecrawl | 35/42 (83%) | 69–92% | partial $0.159 (39/42 priced) | 39/42 | 55k | 219 | 2 | 0 |
| perplexity | 35/42 (83%) | 69–92% | partial $0.148 (37/42 priced) | 37/42 | 42k | 145 | 0 | 0 |
| exa | 34/42 (81%) | 67–90% | unpriced | 0/42 | 54k | 105 | 0 | 0 |
| tavily | 33/42 (79%) | 64–88% | $0.206 | 42/42 | 52k | 119 | 0 | 0 |
| native | 32/42 (76%) | 61–87% | $0.167 | 42/42 | 28k | 256 | 0 | 0 |
| parallel-web | 32/42 (76%) | 61–87% | unpriced | 0/42 | 54k | 90 | 1 | 0 |
| no-search | 30/42 (71%) | 56–83% | partial $0.285 (41/42 priced) | 41/42 | 15k | 0 | 0 | 0 |
| brave | 29/42 (69%) | 54–81% | partial $0.280 (41/42 priced) | 41/42 | 65k | 154 | 0 | 0 |
Competitive deltas against controls
Each arm against each control over the same 14 tasks × 3 repetitions. Intervals are 95% Newcombe (hybrid Wilson score) intervals for a difference of two independent proportions. They ignore the task pairing, so they are conservative. Token ratios compare mean tokens per cell. Provider arms and their controls come from two runs about ten hours apart (see Caveats).
| arm | vs no-search: pass Δ (95% CI) | vs no-search: tokens | vs native: pass Δ (95% CI) | vs native: tokens |
|---|---|---|---|---|
| firecrawl | +12 pp (−6 to +29) | 3.49× | +7 pp (−10 to +24) | 1.95× |
| perplexity | +12 pp (−6 to +29) | 2.63× | +7 pp (−10 to +24) | 1.47× |
| exa | +10 pp (−9 to +27) | 3.39× | +5 pp (−13 to +22) | 1.90× |
| tavily | +7 pp (−11 to +25) | 3.29× | +2 pp (−15 to +20) | 1.84× |
| native | +5 pp (−14 to +23) | 1.79× | — | — |
| parallel-web | +5 pp (−14 to +23) | 3.44× | +0 pp (−18 to +18) | 1.92× |
| no-search | — | — | −5 pp (−23 to +14) | 0.56× |
| brave | −2 pp (−21 to +17) | 4.13× | −7 pp (−25 to +12) | 2.31× |
Expected-fail tasks (4 tasks × 3 repetitions)
Passing here means an honest partial answer or decline, which the validator can confirm.
| arm | pass | 95% CI | $/success | priced | tokens/cell | provider calls | budget-exhausted | ungraded |
|---|---|---|---|---|---|---|---|---|
| exa | 5/12 (42%) | 19–68% | partial $0.046 (6/12 priced) | 6/12 | 67k | 30 | 0 | 0 |
| no-search | 3/12 (25%) | 9–53% | partial $2.827 (9/12 priced) | 9/12 | 43k | 0 | 0 | 0 |
| native | 3/12 (25%) | 9–53% | $0.273 | 12/12 | 16k | 28 | 0 | 0 |
| brave | 3/12 (25%) | 9–53% | partial $0.763 (11/12 priced) | 11/12 | 66k | 45 | 1 | 0 |
| firecrawl | 3/12 (25%) | 9–53% | partial $0.672 (10/12 priced) | 10/12 | 75k | 104 | 2 | 0 |
| perplexity | 3/12 (25%) | 9–53% | partial $0.197 (9/12 priced) | 9/12 | 36k | 30 | 0 | 0 |
| parallel-web | 2/12 (17%) | 5–45% | partial $0.116 (6/12 priced) | 6/12 | 60k | 34 | 1 | 0 |
| tavily | 1/12 (8%) | 1–35% | partial $1.923 (11/12 priced) | 11/12 | 64k | 49 | 1 | 0 |
Per task
| task | outcome | no-search | native | brave | tavily | exa | parallel-web | firecrawl | perplexity |
|---|---|---|---|---|---|---|---|---|---|
change-detection-kubernetes-dockershim-removal | competitive | 2/3 | 2/3 | 1/3 | 3/3 | 2/3 | 3/3 | 3/3 | 3/3 |
change-detection-openapi-30-to-31 | competitive | 1/3 | 1/3 | 0/3 | 0/3 | 2/3 | 2/3 | 0/3 | 1/3 |
competitive-table-cloud-egress-pricing | competitive | 0/3 | 0/3 | 2/3 | 0/3 | 0/3 | 0/3 | 3/3 | 3/3 |
competitive-table-copyleft-obligations | competitive | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
list-build-kubernetes-122-api-removals | competitive | 3/3 | 3/3 | 3/3 | 3/3 | 2/3 | 2/3 | 3/3 | 3/3 |
list-build-python313-pep594-removals | competitive | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
list-build-sca-tool-shortlist | competitive | 0/3 | 2/3 | 2/3 | 1/3 | 3/3 | 3/3 | 1/3 | 3/3 |
multi-hop-cve-to-fixed-release | competitive | 3/3 | 1/3 | 3/3 | 3/3 | 3/3 | 2/3 | 2/3 | 1/3 |
multi-hop-python-feature-peps | competitive | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
unanswerable-nonexistent-postgres-guc | competitive | 2/3 | 3/3 | 0/3 | 3/3 | 2/3 | 2/3 | 3/3 | 1/3 |
unanswerable-private-company-audited-arr | competitive | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 1/3 | 3/3 | 3/3 |
unanswerable-python4-ga-date | competitive | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
upstream-diagnosis-node-openssl3-md4 | competitive | 1/3 | 2/3 | 0/3 | 2/3 | 2/3 | 2/3 | 2/3 | 2/3 |
upstream-diagnosis-urllib3-libressl-import-error | competitive | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
change-detection-unversioned-runtime-drift | expected_fail | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 |
competitive-table-unpublished-enterprise-pricing | expected_fail | 3/3 | 3/3 | 3/3 | 1/3 | 2/3 | 1/3 | 2/3 | 3/3 |
multi-hop-npm-semver-dependents | expected_fail | 0/3 | 0/3 | 0/3 | 0/3 | 3/3 | 1/3 | 1/3 | 0/3 |
upstream-diagnosis-unindexed-ordering-regression | expected_fail | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 | 0/3 |
Findings
1. Search earns its cost on live or long-tail data, not on stable knowledge.
- Tasks whose answers are stable and well known were 3/3 in every arm, including no-search: copyleft obligations, PEP 594 removals, Python feature PEPs and the urllib3 diagnosis.
- The separation comes from the tasks that need current data:
- the egress pricing table;
- the SCA tool shortlist (no-search 0/3; Exa, Parallel and Perplexity 3/3);
- the dockershim change-detection task.
2. Rendering matters for vendor pricing pages. Firecrawl (which renders JavaScript) and Perplexity were the only arms to read AWS's intra-AZ transfer pricing. The others honestly reported it unavailable, and the validator fails an unavailable cell. 3. Declining is not the default for search arms. unanswerable-nonexistent-postgres-guc separates the arms: Brave 0/3 and Perplexity 1/3 against 3/3 for native, Tavily and Firecrawl. 4. Exa handles impossible tasks best (5/12 honest partials on the expected-fail set). It is the only arm to go 3/3 on multi-hop-npm-semver-dependents, where the honest move is to state the method and its limits rather than invent a ranking. 5. Provider arms use more reported tokens per cell. The MCP provider arms report 42k–65k tokens per cell against 15k for no-search. Native search reports 28k. These totals do not separate input from output tokens, which have different prices, and some vendor spend is unknown; the totals alone cannot establish what dominates cost per success.
Failure dive
Two tasks were initially at 0/24 across all arms; a third mostly failed, and the judged SCA task was mostly ungraded. All four exposed bench bugs. Each was confirmed by reading the deliverables, and each is fixed in the bench fix:
| Task | As first graded | Graded with SEWBENCH-01 | Cause |
|---|---|---|---|
unanswerable-python4-ga-date | 0/24 | 24/24 | The decline rejected any date, including "as of 2026-09-28" and other releases' dates. |
unanswerable-private-company-audited-arr | 3/23 graded | 22/23 graded | The prompt asks for how circulating revenue figures were derived; any dollar figure failed. The 24th cell (Parallel) hit its budget before delivering, is not graded, and counts as a failure in the tables, so the per-task row reads 22/24. |
upstream-diagnosis-urllib3-libressl-import-error | 0/24 | 24/24 | A URL-only check on a field every arm filled in prose; the URLs were in evidence_urls. |
list-build-sca-tool-shortlist (judged) | mostly ungraded | graded | The judge sometimes fences its JSON; the default profile name also counted as an arm-identity leak. |
With the fixes, the competitive pass rate rose between 14 and 29 points per arm (no-search +14, Exa +29). It rose most for search arms, which had been penalised for citing and explaining evidence. The first reading of this battery, that no-search was about as good as any provider, was an artifact of these bugs.
Genuine failures that remain:
- Egress pricing: unavailable AWS intra-AZ cells, as above.
change-detection-openapi-30-to-31: most misses omitexample→examplesfrom the breaking changes. OAS 3.1 deprecatesexamplerather than removing it. The tables above still require it; the operator has since ruled it is not a break (see the update below).- Budget exhaustion: 8 cells. Six are Firecrawl or Parallel on long list-building and enumeration tasks; Brave and Tavily have one each.
- By design: the two remaining
expected_failtasks (unversioned-runtime-drift,unindexed-ordering-regression) are 0/3 everywhere.
Codex (partial)
The codex harness (gpt-6-sol) ran one competitive task on 2026-09-28 before its weekly quota ran out:
- Egress pricing: native 2/3, no-search 1/3. Brave and Tavily went 0/3 on the token budget, reading 393k–448k tokens per cell.
- Expected-fail tasks: codex's search arms burned the full provider-call budget on
multi-hop-npm-semver-dependentsinstead of declining.
A full codex battery waits for the quota reset (2026-10-04 12:52Z).
Caveats
- Sample size: n=42 per arm on the competitive set. Adjacent arms are within each other's intervals.
- Timing: the two runs were about ten hours apart. Same harness, model, tasks and budgets.
- Pricing: vendor spend is unpriced for Exa and Parallel, and for Perplexity's
asktool. Their $/success counts model tokens plus metered calls only. - Grading: the judged tasks use a single claude judge, so agreement is not measured.
Update (2026-09-30): operator decisions
- OpenAPI:
example→examplesis a deprecation, not a break, so the task no longer requires it. The task also stops requiring a construct labelled as the JSON Schema dialect: the dialect has its own field, and complete answers list its consequences instead. Both changes are in the bench fix. On replay, the task goes from 7/24 to 22/24. The two remaining misses cite hosts outside the allowlist. - Codex native search is priced at Brave's per-request rate as a proxy for its backing vendor. Claude native search keeps Anthropic's published rate.
- Tokens and outcome rate are reported together. Tokens per success come from the bench fix.
Competitive pass rates regraded with the corrected catalog at the recorded grader revision:
| arm | pass | tokens/success |
|---|---|---|
| firecrawl | 38/42 (90%) | 61.5k |
| perplexity | 37/42 (88%) | 47.7k |
| tavily | 36/42 (86%) | 61.3k |
| exa | 35/42 (83%) | 65.0k |
| native | 33/42 (79%) | 36.3k |
| parallel-web | 33/42 (79%) | 69.8k |
| no-search | 32/42 (76%) | 20.9k |
| brave | 31/42 (74%) | 89.2k |
The historical tokens/cell column averages measured total-billable tokens over attempted cells with measured usage; unknown or estimated usage is excluded, never treated as zero. The update's tokens/success sums measured tokens over graded cells and divides by measured successes, as defined by the bench fix; it excludes ungraded and unmeasured cells and is not a full-arm lower bound. The per-arm measured coverage counts were not retained here. Multiplying a measured-cell mean by all 42 attempts need not reproduce it. The previously quoted fresh/cache/output means are withdrawn: their cell coverage and bucket accounting were not recorded here and they do not reconcile for Firecrawl, Parallel or no-search. They cannot support a token-mix comparison from this document.
The output-reduction claim is also withdrawn. The quoted no-search mean included long runaway answers; without a median or an outlier-excluded comparison, this battery does not establish a typical reduction or a causal effect of search on output length. Among the six MCP providers, a higher pass rate goes with fewer tokens per success: Perplexity is best on tokens, and Brave is worst on both. Spearman ρ ≈ −0.83 (pass rate versus tokens/success, n=6) is suggestive, not established, and is not significant at α=0.05. No arm's pass-rate difference from no-search has an interval that excludes zero. Firecrawl comes closest at +14 points (−2 to +30).
Methodology
WSB methodology
The experiment used Claude Code with claude-opus-5-5[1m], all 18 production tasks, eight tool arms and three repetitions. The two arm groups ran about ten hours apart. Arms differ in exposed search tools: no-search, native, or exactly one provider MCP server. Budgets and contamination checks apply to each cell.
The 14 competitive tasks and four expected-fail tasks have separate denominators. Expected-fail success means an honest partial answer or decline. Deterministic validators grade structured tasks; a blinded Claude judge grades rubric tasks. Single-judge agreement was not measured. Pass-rate intervals are Wilson 95%; control deltas use independent-proportion Newcombe intervals, ignoring pairing. Adjacent rankings are not statistically established.
The original bench corrections fixed overly strict date/dollar decline checks, a prose-versus-URL field mismatch and judge JSON parsing/blinding. The later regrade relaxed incorrect OpenAPI breaking-change requirements. Both grade the same stored answers; neither reran agents. REPORT.md retains initial and updated tables with their original precision.
Historical cost lower bounds sum priced spend and divide by all successes. The report command instead uses priced-cell successes, so its output can differ. Unknown vendor spend is not zero; partially priced arms are excluded from cost ranking. Updated tokens per success use measured tokens from graded cells divided by measured successes; coverage counts were not retained. Do not multiply the historical tokens/cell means by attempted-cell counts to derive that update. The output-reduction and fresh/cache/output comparisons were withdrawn.
See recorded results, summary and reproduction. Calibration was not part of this experiment: calibration record.
Search gap bench (2026-10-03)
Full recorded report
GAP battery results, 2026-10-03: does search change the outcome of the job?
Correction (2026-10-06): the Claude Code native arm could not fetch pages. The arm exposed WebSearch and WebFetch but pre-approved only WebSearch, and headless Claude Code refuses tools that are not pre-approved. Every WebFetch call in the native arm was refused (50 calls across all 21 native cells). Each provider arm could use its vendor's fetch tool, so the Claude Code native row measures search without page fetching and is not comparable to the provider rows. The codex native arm was unaffected. The configuration is fixed and pinned by a test; the recorded numbers below are unchanged. Details: agent search behavior.
The first full Search Gap Bench (GAP) battery ran on 2026-10-03, operator-approved: brief tasks on two harnesses, nine arms each, three repetitions.
- GAP grades the job, not the retrieval.
- A task counts only after calibration shows its knowledge gap: the no-search arm (floor) passes at most 1 of 5, and the arm given the answer excerpt (ceiling) passes at least 4 of 5.
- Each search arm is scored by gap closure: how much of the floor-to-ceiling distance it covers.
- Tokens are reported beside every outcome.
Headline. On these tasks, search decides the outcome.
- The floor passed 0% on both harnesses, and every search arm passed 62–95%.
- Which provider is best depends on the harness: Parallel led on claude-code and Brave on codex.
- Pass rate did not predict cost: providers with similar pass rates differed almost twofold in tokens per success.
These are the recorded results after the Codex reruns. Claude-code provider tool availability remains unverified in the published evidence: no transcript audit for those cells is documented, and they lack availability records. Its provider rankings are provisional pending that audit. In the first codex pass, 31 of 108 provider cells ran without their search tool. They were re-run on 2026-10-04 under the GAPMCP-01 fix, every re-run cell loaded its tool, and the codex tables below include them (see "Execution changes").
Setup
| Harnesses | claude-code on claude-opus-5-5, codex on gpt-6.1-sol |
| Arms | floor (no search), ceiling (answer excerpt in the prompt), native, Brave, Tavily, Exa, Parallel, Firecrawl, Perplexity |
| Tasks | 10 brief tasks from GAP-09: 6 decision briefs and 4 research briefs. The code corpus (GAP-07) merged during the run and gets its own battery. |
| Repetitions | 5 per reference arm in calibration; 3 per arm in the battery |
| Grading | Decision correctness, weighted key-fact recall (≥ 0.7) and unsupported-claim rate (≤ 0.25), judged by two blinded judges (claude-code primary, codex secondary) against captured primary sources |
| Spend | about 50M tokens in total: battery agents 21.1M; battery judges 15.3M; calibration agents 4.2M; calibration judges 9.2M |
Calibration
Final admissions under GAPRULE-01. Each cell shows floor passes and ceiling passes out of 5.
| Task | Family | claude-code | codex |
|---|---|---|---|
| billing-reporting | research | Floor: 0, Ceiling: 5 — admitted | Floor: 0, Ceiling: 5 — admitted |
| classroom-term | decision | Floor: 0, Ceiling: 5 — admitted | Floor: 0, Ceiling: 5 — admitted |
| models-production | decision | Floor: 0, Ceiling: 5 — admitted | Floor: 0, Ceiling: 5 — admitted |
| npm-token-scope | research | Floor: 0, Ceiling: 5 — admitted | Floor: 0, Ceiling: 5 — admitted |
| org-migration | decision | Floor: 0, Ceiling: 5 — admitted | Floor: 0, Ceiling: 4 — admitted |
| python-metadata | research | Floor: 0, Ceiling: 5 — admitted | Floor: 0, Ceiling: 5 — admitted |
| tarfile-upload | decision | Floor: 0, Ceiling: 5 — admitted | Floor: 0, Ceiling: 0 — rejected |
| cli-trust | decision | Floor: 0, Ceiling: 2 — rejected | Floor: 0, Ceiling: 2 — rejected |
| spark-editing | decision | Floor: 0, Ceiling: 1 — rejected | Floor: 0, Ceiling: 0 — rejected |
| python-security-march | research | Floor: 0, Ceiling: 0 — rejected | Floor: 0, Ceiling: 1 — rejected |
Every floor scored 0, so every task has a real knowledge gap. All rejections are ceiling failures, which point at task or rubric defects:
- spark-editing: the heaviest key fact is not what the prompt asks about.
- python-security-march: the excerpt omitted the date that its top fact requires. The excerpt fix is pending for the next battery; GAPRULE-01 changed the unsupported-claim limit, not the excerpt.
- cli-trust: with three claims, a single unsupported claim fails the brief.
- tarfile-upload on codex: the codex briefs miss key facts.
Gap closure
claude-code (7 tasks, 21 cells per arm)
| Arm | Gap closure (95% CI) | Pass rate (Wilson 95% CI) | Tokens per success | vs floor (exact McNemar) |
|---|---|---|---|---|
| Parallel | 1.00 (0.86–1.17) | 95% (77–99%) | 64k | p < 0.0001 |
| Firecrawl | 0.95 (0.83–1.00) | 90% (71–97%) | 116k | p < 0.0001 |
| Brave | 0.85 (0.62–1.00) | 81% (60–92%) | 93k | p < 0.0001 |
| Perplexity | 0.80 (0.50–1.00) | 76% (55–89%) | 74k | p < 0.0001 |
| Exa | 0.75 (0.47–0.95) | 71% (50–86%) | 89k | p < 0.0001 |
| Tavily | 0.75 (0.43–1.00) | 71% (50–86%) | 83k | p < 0.0001 |
| native | 0.65 (0.37–0.91) | 62% (41–79%) | 66k | p = 0.0002 |
| ceiling (reference) | 1.00 | 95% | 4k | |
| floor (reference) | 0.00 | 0% | — |
By family:
- Research briefs: every provider arm passed 100% (native 89%).
- Decision briefs: they separate the providers. Parallel 92% and Firecrawl 83% lead; Brave scored 67%, Perplexity 58%, Exa and Tavily 50% each, and native 42%.
codex (6 tasks, 18 cells per arm)
| Arm | Gap closure (95% CI) | Pass rate (Wilson 95% CI) | Tokens per success | vs floor (exact McNemar) |
|---|---|---|---|---|
| Brave | 1.06 (0.88–1.29) | 94% (74–99%) | 111k | p < 0.0001 |
| Firecrawl | 1.00 (1.00–1.00) | 89% (67–97%) | 166k | p < 0.0001 |
| Tavily | 1.00 (0.71–1.29) | 89% (67–97%) | 129k | p < 0.0001 |
| Exa | 0.88 (0.47–1.29) | 78% (55–91%) | 146k | p = 0.0001 |
| native | 0.88 (0.47–1.20) | 78% (55–91%) | 117k | p = 0.0001 |
| Parallel | 0.81 (0.41–1.14) | 72% (49–88%) | 157k | p = 0.0002 |
| Perplexity | 0.81 (0.47–1.12) | 72% (49–88%) | 117k | p = 0.0002 |
| ceiling (reference) | 1.00 | 89% (67–97%) | 12k | |
| floor (reference) | 0.00 | 0% (0–18%) | — |
No search arm beat native significantly; the largest gain was Brave at +16.7 points (p = 0.25).
By family:
- Research briefs: every arm passed 100%, except Exa at 89%.
- Decision briefs: they separate the providers. Brave scored 89%; Firecrawl and Tavily 78% each; Exa 67%; native 56%; Parallel and Perplexity 44% each.
The first pass's estimates that excluded the 31 tool-less cells were close to the re-run results:
| Arm | Estimate | After re-run |
|---|---|---|
| Brave | 93% | 94% |
| Firecrawl | 85% | 89% |
| Tavily | 85% | 89% |
| Exa | 76% | 78% |
| Parallel | 67% | 72% |
| Perplexity | 62% | 72% |
What this shows
1. On these tasks, search changes the outcome. Without search, both harnesses failed every brief. With a good provider they reached the answer-excerpt arm's level. The difference against the floor is significant for every search arm on both harnesses (exact McNemar, p ≤ 0.0002). 2. The best provider depends on the harness. On claude-code, Parallel (95%) and Firecrawl (90%) led. On codex, Brave led (94%), ahead of Firecrawl and Tavily (89% each).
- Parallel tied for last on codex (72%), even with its tool loaded in every cell. On codex decision briefs it passed 44%, against 92% on claude-code.
- So a provider's ranking on one harness does not carry over to another.
3. Pass rate does not predict cost between providers. Tokens per success counts every agent token in the arm, failed cells included, divided by passing cells.
- The hypothesis that better search shows up as both lower token spend and higher success holds against no search, which spent tokens and passed nothing. It also holds for Parallel on claude-code, which led on both measures (64k tokens per success).
- It does not hold as a ranking. Among the seven search arms, the rank correlation between pass rate and tokens per success was +0.13 on claude-code and −0.20 on codex; negative means higher-passing arms cost less. With seven arms, neither value is distinguishable from zero.
- Per-cell spend varied widely: 41k (native) to 105k (Firecrawl) on claude-code, and 84k (Perplexity) to 148k (Firecrawl) on codex. Firecrawl passed 90% at 116k tokens per success, against Parallel's 95% at 64k.
- Choosing a provider means weighing pass rate and cost separately.
4. Decision briefs separate the providers; research briefs mostly don't. Research briefs passed at or near 100% on both harnesses. Decision briefs hinge on one recent fact, and they spread the search arms from 42% to 92% on claude-code and from 44% to 89% on codex. 5. Native search was not the best choice on either harness. It came last among search arms on claude-code (62%). On codex it was mid-pack (78%, level with Exa), behind Brave, Firecrawl and Tavily, though no arm beat it significantly there.
Changes made during the run
The first live runs surfaced several bench defects. Two changed how cells execute, and two changed how cells are graded. The timeline was recorded from the battery change log and cell start events.
Execution changes
Every cell in the tables above ran under both of these fixes:
- GAPBUDGET-01: brief budgets became runaway caps (1M tokens, 60 calls). The old 30k cap was below the median codex search cell.
- It was part of the pinned export before any calibration or battery cell ran.
- No cell ran under the old caps.
- GAPARM-01: brief search arms resolve the GAP catalog. Before it, no brief search arm could run, so no search-arm cell predates it.
- It was applied at 18:30:04Z on 2026-10-03.
- The earliest search-arm cell started at 18:30:29Z (codex round 1). claude-code started at 19:12Z.
The tool-isolation remediation, which forbids the GAP catalog to every GAP cell, was applied at 18:50:40Z and the codex battery was restarted to load it.
- 18 codex round-1 cells started before it: classroom-term repetitions 1 and 2, all arms.
- No catalog access in them. Their transcripts contain only search calls and agent messages, with no file reads or commands, so the catalog was not read.
- They were kept, not re-run.
- Every other cell ran after the remediation.
The codex provider re-run (2026-10-04). In the first pass, 31 of codex's 108 provider cells made no provider call:
| Arm | Cells without a provider call |
|---|---|
| Parallel | 12 |
| Firecrawl | 5 |
| Tavily | 5 |
| Perplexity | 5 |
| Brave | 3 |
| Exa | 1 |
- Cause: GAPMCP-01 reproduced codex's 30-second default MCP startup timeout. A 35-second server fails under it and starts with a 120-second allowance.
- How the cells were selected: the 31 cells were reclassified from an audit of their transcripts. All of them had zero provider calls and a
list_mcp_resourcesprobe; 29 of the 31 also said in their answer that no retrieval was available. The private audit is not distributed. - How they were re-run: with
--rerun-unavailableunder GAPMCP-01 (a 120-second startup allowance,required=trueand pre-warmed pinned packages), using the same matrix, calibration and catalog. - Every re-run cell loaded its tool. Each one recorded a completed
initializeand tool listing. - A second fix was needed. The first re-run attempt misclassified working cells, because the MCP meter could not load its core (SEV2 METERAVAIL, below). Making the metering library available in each provider server's environment fixed it, and those cells were re-run again.
These 31 cells ran under the startup allowance; the other 77 provider cells ran without it but had their tool. No agent-side setting changed.
Grading changes
These changed only how stored answers were judged. Every affected cell was regraded from the agents' stored answers, with no agent re-runs:
- GAPJUDGE-01: judges read readable text instead of raw HTML, and each source once. This cut judge tokens per cell from about 250k to about 15k. It also covered:
- safe redirects;
- gzip decoding (python.org served gzip unconditionally, which garbled every python.org citation);
- over-budget sources withheld instead of the whole cell going ungraded;
- the blinding leak check reading decoded text, and generic arm labels no longer counted as leaks.
- GAPRULE-01, operator-approved: the unsupported-claim limit was raised from 0.10 to 0.25. Ceiling agents see a short excerpt, and their hedges about it ("the notice does not specify X") were contradicted by the full pages the judges read. All stored grades were recomputed from their saved judge labels; 49 outcomes flipped, all from fail to pass.
Open issues
- SEV2 GAPMCP (fix GAPMCP-01, the bench fix): resolved. The 31 affected codex cells were re-run and the codex table updated.
- SEV2 METERAVAIL (fix METERAVAIL-01): provider-call metering was silently off for this battery on both harnesses. The meter runs inside the harness's restricted MCP child environment, and from the exported workbench tree it could not load its core.
- Vendor spend per cell was therefore not metered.
- Only the 31 re-run cells have availability records; the report's "availability coverage" column counts those. Claude-code provider availability has not been verified in the published evidence.
- SEV2 ARGVKEY (fix ARGVKEY-01): the Parallel API key was passed in
mcp-remote's process arguments. The battery configuration now passes it through a 0600 header file. - The claude-code floor arm sometimes writes pretend searches as text when it has no tools. It is a legitimate failure, but slow and expensive (up to about 130k output tokens in 13 minutes).
- Next battery:
- the code corpus (GAP-07: 16 code tasks and 4 controls);
- the python-security-march excerpt fix;
- rubric fixes for spark-editing and cli-trust.
Methodology
GAP methodology
GAP grades completed jobs against captured primary sources. The initial corpus contained ten briefs: six decision and four research. Calibration is per harness and model: five repetitions of floor (no search) and ceiling (answer excerpt). Admission requires floor at most one pass and ceiling at least four passes. Seven Claude tasks and six Codex tasks were admitted. Code tasks were not run in this battery. The committed calibration file is an admission summary, not a runtime calibration file with captured oracle sources.
Each admitted task ran three times on nine arms. Brief grading requires decision correctness, weighted key-fact recall at least 0.7 and unsupported-claim rate at most 0.25. Two blinded judges used Claude Code as primary and Codex as secondary. Gap closure is (arm pass rate − floor pass rate) / (ceiling pass rate − floor pass rate); it can exceed one. Wilson 95% intervals accompany pass rates. Gap-closure intervals use a seeded paired task bootstrap with 10,000 resamples. Exact McNemar tests pair task and repetition, against floor and native. The findings publish floor comparisons and the largest Codex native comparison, not every native test.
Tokens per success sum all agent tokens in an arm, including failures, divided by passes. Judge and calibration spend is separate. A zero-success floor has undefined tokens per success. Vendor spend was unmetered; tokens are not dollar cost.
GAPJUDGE-01 changed readable source extraction, redirects, gzip, source budget handling and blinding. GAPRULE-01 raised the unsupported-claim limit from 0.10 to 0.25; recomputing saved judge labels flipped 49 failures to passes. These grading changes did not rerun agents. Execution changes and the later rerun of 31 Codex cells with unavailable provider tools are preserved separately in REPORT.md.
See recorded results, summary, admission counts and reproduction.
How agents used search (2026-10-05)
Full recorded report
How agents used search: query style, fetch-from-memory and primary-source retrieval (2026-10-05)
This report analyses the stored transcripts of the two published batteries, the web search bakeoff and the search gap bench. No agent was rerun. It asks how agents used the search tools they were given, rather than which provider scored best. Every number below comes from scripts/analyze_search_behavior.py, whose classification rules are listed in methodology. The analysis covers 974 search queries from the GAP battery (both harnesses) and 655 from the bakeoff (Claude Code).
Headline
- Agents write keyword queries whatever the engine asks for. Of 974 GAP search queries, 7.3% read as natural language. In the bakeoff the figure is 6.0%. Exa's tool asks for a description of the ideal page rather than keywords; its queries were natural language in 3% (Claude Code) and 2% (codex) of GAP calls. Because the automatic rule can miss descriptive phrasing, all 111 Exa queries from both batteries were also read by hand (appendix): none is phrased as a description of a page. The closest are a headline-like phrase used three times ("Python 4.0 release date announced by Python Steering Council") and a page title.
- The two harnesses express constraints differently. On the six provider arms, codex put
site:into 31% to 67% of its queries, even where a structured domain filter existed. Claude Code wrotesite:only on Firecrawl (6%) and otherwise used structured domain filters where a tool offered one (Perplexitysearch_domain_filter, Tavilyinclude_domains, WebSearchallowed_domains). - Agents often fetch remembered URLs instead of searching. In the bakeoff, Claude Code fetched pages without any search in 36 of 55 Exa cells, 35 of 55 Firecrawl cells, 34 of 54 Tavily cells and 30 of 55 Parallel cells. Those provider arms therefore measured page fetching as much as search. On post-cutoff GAP tasks the same habit failed: Claude Code's Exa cells that only fetched passed 1/3, against 14/18 for its Exa cells that searched.
- Retrieving the primary source mattered most on Claude Code. Pooled over provider and native arms, Claude Code passed 94/111 cells (85%) when the task's primary source appeared in search results and 20/33 (61%) when it did not. Codex surfaced the primary source in most searched cells, so the comparison is thin there (95/116 against 8/10). Two arms passed without the primary page: Parallel (6/7) and Firecrawl (9/9) on Claude Code, which indicates their results carried the facts from other pages.
These are observations about agent behavior with each vendor's MCP server at its pinned version and default settings. They are not measurements of the vendors' APIs.
What each search tool asks for
Descriptions are quoted from the pinned MCP server packages and, for Parallel's hosted server, from its published documentation.
| Arm | Search interface the agent saw | What the tool tells the agent about queries |
|---|---|---|
| Exa | web_search_exa(query, numResults) | "describe the ideal page, not keywords"; query is a "semantically rich description of the ideal page, not just keywords" |
| Parallel | web_search(objective, search_queries, …) | a natural-language objective plus "concise, related keyword queries of 3–6 words each" |
| Perplexity | perplexity_search(query, …) plus ask, research, reason | "Supports recency filters, domain restrictions, and a lower-latency fast search mode" |
| Tavily | tavily_search(query, search_depth, …) | query is described only as "Search query" |
| Firecrawl | firecrawl_search(query, …) | "Query operators, domain filters, categories … are described on their parameters" |
| Brave | brave_web_search(query, …) | lists when to use the tool; no guidance on query phrasing |
| native (Claude Code) | WebSearch(query, allowed_domains, …) | harness built-in |
| native (codex) | built-in web search | harness built-in; no parameters exposed to the transcript |
Query style (GAP battery)
Parallel's search_queries are keyword queries by design; its natural-language part is the separate objective field, which is counted as a structured parameter and not as a query. Perplexity's user messages count as queries alongside its query inputs; system and assistant messages are excluded. Codex native search reports its query lists in the transcript; page views are counted as fetches.
| Harness | Arm | Queries | Natural language | Contains a year | Quoted phrase | site: | Median words | Structured parameters passed (calls) |
|---|---|---|---|---|---|---|---|---|
| claude-code | native | 101 | 7% | 43% | 39% | 0% | 11 | allowed_domains ×11 |
| claude-code | Brave | 72 | 17% | 44% | 18% | 0% | 8 | count ×23, extra_snippets ×21, maximum_number_of_tokens ×18, freshness ×1, maximum_number_of_urls ×1 |
| claude-code | Tavily | 38 | 13% | 45% | 18% | 0% | 9 | max_results ×38, include_domains ×19, search_depth ×17, include_raw_content ×1, start_date ×1 |
| claude-code | Exa | 38 | 3% | 58% | 0% | 0% | 10 | numResults ×38 |
| claude-code | Parallel | 94 | 5% | 24% | 2% | 0% | 6 | model_name ×27, objective ×27, session_id ×27 |
| claude-code | Firecrawl | 34 | 0% | 59% | 9% | 6% | 7 | limit ×34, sources ×34, includeDomains ×3 |
| claude-code | Perplexity | 70 | 20% | 36% | 10% | 0% | 9 | max_results ×60, search_domain_filter ×42, max_tokens_per_page ×5, messages ×5, search_context_size ×3, search_recency_filter ×2, search_type ×1 |
| codex | native | 48 | 0% | 50% | 33% | 10% | 8 | none |
| codex | Brave | 114 | 2% | 52% | 26% | 46% | 8 | maximum_number_of_tokens ×30, count ×15, maximum_number_of_urls ×13, extra_snippets ×7, url ×3 |
| codex | Tavily | 130 | 6% | 50% | 59% | 67% | 8 | max_results ×123, search_depth ×26, include_raw_content ×25, end_date ×3, include_domains ×3 |
| codex | Exa | 52 | 2% | 60% | 35% | 58% | 10 | numResults ×10 |
| codex | Parallel | 62 | 6% | 39% | 27% | 48% | 7 | objective ×28, session_id ×10, max_results ×3, max_chars_per_result ×1 |
| codex | Firecrawl | 42 | 2% | 74% | 26% | 31% | 8 | limit ×33, sources ×10 |
| codex | Perplexity | 79 | 14% | 71% | 33% | 42% | 10 | max_results ×20, messages ×18, max_tokens_per_page ×16 |
Query style (web search bakeoff, Claude Code)
| Arm | Cells | Queries | Natural language | Contains a year | Quoted phrase | site: | Median words | Structured parameters passed (calls) |
|---|---|---|---|---|---|---|---|---|
| native | 54 | 137 | 2% | 34% | 9% | 4% | 7 | allowed_domains ×36 |
| Brave | 54 | 199 | 4% | 7% | 11% | 6% | 6 | count ×85, maximum_number_of_tokens ×80, freshness ×17, goggles ×16, extra_snippets ×11, maximum_number_of_urls ×9, maximum_number_of_tokens_per_url ×3 |
| Tavily | 54 | 25 | 0% | 32% | 12% | 0% | 10 | max_results ×18, include_domains ×10, search_depth ×4, time_range ×2, start_date ×1 |
| Exa | 55 | 21 | 0% | 19% | 5% | 0% | 10 | numResults ×21 |
| Parallel | 55 | 77 | 0% | 20% | 0% | 1% | 5 | model_name ×22, objective ×22, session_id ×22 |
| Firecrawl | 55 | 20 | 0% | 25% | 15% | 5% | 7 | limit ×20, sources ×20 |
| Perplexity | 55 | 176 | 16% | 13% | 10% | 0% | 8 | max_results ×154, search_domain_filter ×115, max_tokens_per_page ×21, search_type ×21, messages ×13, search_context_size ×13, search_recency_filter ×5 |
Searching versus fetching from memory (web search bakeoff, Claude Code)
A cell is "fetched without searching" when it made at least one fetch call and no search call. Brave's and Perplexity's servers expose no page-fetch tool, so their agents had to search or answer from memory. Every arm has a handful of cells with no provider call: answers given from memory and cells that ended before any tool call.
| Arm | Cells | Fetch tool in the arm | Fetched without searching | No provider call | Calls refused by the harness |
|---|---|---|---|---|---|
| native | 54 | WebFetch (mostly refused; see below) | 12 | 6 | 110 |
| Brave | 54 | no | 0 | 7 | 0 |
| Tavily | 54 | yes (tavily_extract) | 34 | 6 | 0 |
| Exa | 55 | yes (web_fetch_exa) | 36 | 6 | 0 |
| Parallel | 55 | yes (web_fetch) | 30 | 6 | 0 |
| Firecrawl | 55 | yes (firecrawl_scrape) | 35 | 6 | 0 |
| Perplexity | 55 | no | 0 | 7 | 0 |
On the bakeoff's stable-knowledge tasks, fetching a remembered URL is efficient: those tasks passed in every arm, including no-search. It means the bakeoff's provider arms with a fetch tool mostly measured that fetch path. The GAP battery, whose answers post-date the models, shows the cost of the same habit on new information.
Harness refusals: the Claude Code native arm
Claude Code's native arm exposed WebSearch and WebFetch through --tools but pre-approved only WebSearch through --allowedTools. Headless Claude Code refuses a tool that is exposed but not pre-approved ("Claude requested permissions to use WebFetch, but you haven't granted it yet"). In the GAP battery every native WebFetch call was refused (50 calls in 21 of 21 cells); in the bakeoff, 110 calls in 40 of 54 cells. No provider arm and no codex arm had a refused call. The Claude Code native results in both batteries therefore understate that harness's web tools. The arm configuration is fixed and covered by a regression test. The fetch counts in this report include refused attempts.
Primary-source retrieval (GAP battery)
Each GAP brief has one primary-source URL (the catalog oracle). A cell surfaced it when that URL appeared among the URLs of returned search results, excluding search inputs. Unknown coverage is excluded from both pass-rate groups. Recomputing against returned evidence leaves the historical source counts unchanged; no searched cell has unknown coverage. Facts can also come from secondary pages, so surfacing is a proxy for retrieval, not a requirement for passing.
| Harness | Arm | Cells that searched | Unknown coverage | Primary source surfaced | Pass when surfaced | Pass when not surfaced |
|---|---|---|---|---|---|---|
| claude-code | native | 21 | 0 | 18/21 | 13/18 | 0/3 |
| claude-code | Brave | 21 | 0 | 18/21 | 15/18 | 2/3 |
| claude-code | Tavily | 21 | 0 | 16/21 | 13/16 | 2/5 |
| claude-code | Exa | 18 | 0 | 14/18 | 13/14 | 1/4 |
| claude-code | Parallel | 21 | 0 | 14/21 | 14/14 | 6/7 |
| claude-code | Firecrawl | 21 | 0 | 12/21 | 10/12 | 9/9 |
| claude-code | Perplexity | 21 | 0 | 19/21 | 16/19 | 0/2 |
| codex | native | 18 | 0 | 17/18 | 13/17 | 1/1 |
| codex | Brave | 18 | 0 | 16/18 | 15/16 | 2/2 |
| codex | Tavily | 18 | 0 | 16/18 | 14/16 | 2/2 |
| codex | Exa | 18 | 0 | 17/18 | 14/17 | 0/1 |
| codex | Parallel | 18 | 0 | 17/18 | 12/17 | 1/1 |
| codex | Firecrawl | 18 | 0 | 15/18 | 14/15 | 2/3 |
| codex | Perplexity | 18 | 0 | 18/18 | 13/18 | — |
Limits
- Cell counts per arm are small (18 to 21 in GAP, 54 to 55 in the bakeoff). Per-arm percentages move by about five points per cell.
- The natural-language rule is a heuristic, stated in methodology; readers can rerun the script with their own rule.
- Each vendor was used through its MCP server at the pinned version and default settings. Other interfaces of the same vendor (direct API, other tools, other modes) were not tested.
- The two harnesses ran different task subsets in GAP (7 tasks for Claude Code, 6 for codex), so cross-harness differences mix agent behavior with task mix.
Appendix: every Exa query
Every query sent to web_search_exa in both batteries, in transcript order, so readers can judge the phrasing directly. Other arms' queries are summarised above but not reproduced.
| Battery | Harness | Query |
|---|---|---|
| GAP | claude-code | Copilot Billing Preview app deprecated usage report |
| GAP | claude-code | GitHub Copilot premium request budgets per user cap and usage report CSV export docs |
| GAP | claude-code | GitHub docs AI usage page billing settings group filter export AI credits Copilot limitations |
| GAP | claude-code | GitHub Copilot Billing Preview app deprecated usage-based billing budgets |
| GAP | claude-code | GitHub docs Copilot premium requests usage report export CSV billing 2026 |
| GAP | claude-code | GitHub docs AI usage page billing settings group filter export AI credits usage report limitations |
| GAP | claude-code | GitHub Copilot Billing Preview app deprecated usage report |
| GAP | claude-code | GitHub docs set budget for individual Copilot premium requests user-level budget |
| GAP | claude-code | GitHub docs AI usage page billing settings group filter export AI credits Copilot usage report limitations |
| GAP | claude-code | github.blog changelog 2026-08-04 Retiring the Copilot Billing Preview app |
| GAP | claude-code | GitHub Classroom deprecation sunset announcement 2026 |
| GAP | claude-code | Classroom 50 CS50 foundation GitHub Classroom alternative import autograding LTI |
| GAP | claude-code | Codio GitHub Classroom migration partner |
| GAP | claude-code | GitHub Classroom sunset deprecation announcement 2026 |
| GAP | claude-code | Classroom 50 CS50 open source GitHub Classroom replacement import migration |
| GAP | claude-code | GitHub Classroom deprecation sunset announcement 2026 |
| GAP | claude-code | GitHub Models deprecation retirement changelog 2026 BYOK inference |
| GAP | claude-code | GitHub Models deprecation retirement changelog 2026 BYOK paid usage |
| GAP | claude-code | Azure AI Inference beta SDK deprecated retirement August 26 2026 migrate to OpenAI v1 API |
| GAP | claude-code | GitHub Models deprecation retirement changelog 2026 BYOK |
| GAP | claude-code | GitHub secret scanning npm token revocation credential types 2026 changelog |
| GAP | claude-code | GitHub changelog npm granular access tokens publish and manage package access 2026 |
| GAP | claude-code | GitHub secret scanning npm token revocation credential types in scope 2026 |
| GAP | claude-code | GitHub changelog npm granular access tokens bypass 2FA publish manage package access 2026 |
| GAP | claude-code | GitHub changelog npm granular access tokens bypass 2FA publish package access management 2026 |
| GAP | claude-code | GitHub secret scanning npm token revocation credential types 2026 changelog |
| GAP | claude-code | GitHub changelog npm staged publishing generally available npm stage publish 2026 |
| GAP | claude-code | Python 3.10 security release tarfile extraction filter bypass CVE-2025-4517 |
| GAP | claude-code | CPython tarfile data filter vulnerability 2026 CVE security-announce |
| GAP | claude-code | Python 3.10 security release tarfile extraction filter bypass fix |
| GAP | claude-code | CVE-2026-4360 tarfile extract filter hardlinks Python fixed version |
| GAP | claude-code | GitHub docs converting a user into an organization warning cannot sign in personal account |
| GAP | claude-code | GitHub Enterprise Server docs promoting or demoting site administrator converting user to organization |
| GAP | claude-code | python.org release metadata API authentication vulnerability security advisory 2026 |
| GAP | claude-code | python.org release metadata vulnerability unauthenticated API modify release files 2026 |
| GAP | claude-code | python.org release API authentication bypass DEVCORE Splitline CVE advisory exploitation |
| GAP | claude-code | python.org release metadata vulnerability unauthenticated modify release files advisory 2026 |
| GAP | claude-code | Python Software Foundation security incident python.org downloads release API authentication |
| GAP | codex | GitHub Classroom sunset July 2026 September 2026 partner repositories |
| GAP | codex | GitHub Classroom retirement September 1 2026 repositories partner |
| GAP | codex | site:github.com/orgs/community/discussions/205975 "September 4" |
| GAP | codex | GitHub Classroom sunset September 2026 repositories partner retirement |
| GAP | codex | site:github.com/orgs/community/discussions GitHub Classroom "September 4" "2026" |
| GAP | codex | GitHub Models inference BYOK retirement July August 2026 paid existing usage |
| GAP | codex | GitHub Models retiring July August 2026 paid inference BYOK July 10 |
| GAP | codex | GitHub Models retirement July August 2026 paid inference BYOK existing usage |
| GAP | codex | site:docs.github.com converting user into organization January 2026 transfer repositories personal account enterprise server |
| GAP | codex | site:github.blog/changelog 2026 January converting user organization disabled January 2026 |
| GAP | codex | site:docs.github.com converting a user into an organization January 2026 transfer repositories |
| GAP | codex | site:github.blog/changelog 2026 January converting personal accounts organizations January 20 |
| GAP | codex | site:docs.github.com converting user into organization January 2026 transfer repositories personal account GHES |
| GAP | codex | site:github.blog/changelog 2026 January 20 convert account organization transfer personal account |
| GAP | codex | GitHub Copilot Billing Preview app retired individual budgets usage export billing reports |
| GAP | codex | site:docs.github.com Copilot AI usage page reports limitations billing API user level budgets export usage credits |
| GAP | codex | site:docs.github.com "AI usage" "report" "does not" billing usage reports |
| GAP | codex | site:docs.github.com "Setting up budgets" "user-level" "Budgets and alerts" |
| GAP | codex | "Copilot Billing Preview" budgets export usage |
| GAP | codex | site:docs.github.com AI usage user-level budgets billing reports Copilot AI credits |
| GAP | codex | site:docs.github.com "AI usage" "does not" reports billing |
| GAP | codex | site:docs.github.com "Viewing your usage" "AI" "report" billing |
| GAP | codex | site:docs.github.com "Automating usage reporting" "AI" |
| GAP | codex | site:docs.github.com Copilot usage metrics billing data coverage IDE chat github CLI |
| GAP | codex | site:docs.github.com "Setting" "user-level budgets" "Budgets" |
| GAP | codex | Copilot Billing Preview app individual budgets export raw usage finance billing dashboard |
| GAP | codex | site.github.blog/changelog 2026-08-04 retiring copilot billing preview |
| GAP | codex | site:docs.github.com AI usage page billing reports user-level budgets coverage limitations Copilot |
| GAP | codex | site:docs.github.com "AI usage" "does not" reports billing |
| GAP | codex | site:docs.github.com "AI usage" "coverage" |
| GAP | codex | site:docs.github.com "Viewing your usage of metered products" "AI usage" |
| GAP | codex | site:github.blog/changelog "user-level budgets" 2026 |
| GAP | codex | site:docs.github.com copilot usage metrics "billing" "coverage" |
| GAP | codex | site:github.blog/changelog billing CSV usage reports API 2026 |
| GAP | codex | npm token security granular tokens package access management July 2026 August 2026 GitHub credentials |
| GAP | codex | site:github.blog/changelog/2026 staged publishing npm May 22 trusted publishing tokens approve |
| GAP | codex | site:github.blog/changelog/ "2026-05-22" "npm" |
| GAP | codex | npm token security granular tokens package access github August 2026 publishing tokens |
| GAP | codex | site:github.blog/changelog npm December 9 2025 classic tokens revoked staged publishing July 2026 |
| GAP | codex | npm token security August 2026 package access granular tokens GitHub credentials |
| GAP | codex | site:github.blog/changelog npm staged publishing 2026 July 2FA |
| GAP | codex | python.org release metadata vulnerability unsigned downloads API incident 2026 2025 |
| GAP | codex | site.python.org Sigstore verify Python releases identity PEP 761 |
| GAP | codex | site.blog.python.org/2026/06/mitigated-api-bypass "Timeline" "30" "June" |
| GAP | codex | "python/pythondotorg" "3014" "merged" "June" |
| GAP | codex | python.org release metadata vulnerability authentication 2026 2025 security incident |
| GAP | codex | site:blog.python.org/2026/06/mitigated-api-bypass "Remediations" "Timeline" |
| GAP | codex | site:github.com/python/pythondotorg/issues/3010 OR site:github.com/python/pythondotorg/issues/3011 |
| GAP | codex | python.org release metadata security incident authentication bypass before June 30 2026 Trail of Bits |
| GAP | codex | python.org release metadata vulnerability authentication 2026 2025 API incident |
| GAP | codex | site.python.org security incident release metadata 2026 2025 downloads authentication |
| GAP | codex | Python Windows code signing certificates revoked August 2025 update investigation June 2026 |
| bakeoff | claude-code | Python 4.0 release date announced by Python Steering Council |
| bakeoff | claude-code | webpack md4 OpenSSL 3 Node 17 hash wasm md4 fix output.hashFunction xxhash64 sokra comment |
| bakeoff | claude-code | grype offline air-gapped vulnerability database import GRYPE_DB_AUTO_UPDATE false |
| bakeoff | claude-code | cargo audit --no-fetch --db local advisory database path offline |
| bakeoff | claude-code | Dependency-Track air-gapped offline vulnerability mirror configuration NVD OSV feeds URL |
| bakeoff | claude-code | Python 4.0 release date announced by Python Steering Council |
| bakeoff | claude-code | Python 4.0 release date announced by Python Steering Council |
| bakeoff | claude-code | Node.js documentation --openssl-legacy-provider option "Enable OpenSSL 3.0 legacy provider" |
| bakeoff | claude-code | aws.amazon.com data transfer out to internet US East (N. Virginia) first 10 TB per GB price same availability zone free |
| bakeoff | claude-code | grype air-gapped offline database import GRYPE_DB_AUTO_UPDATE false documentation |
| bakeoff | claude-code | Dependency-Track air-gapped offline NVD mirror configuration documentation |
| bakeoff | claude-code | Safety CLI 3 requires account login API key vulnerability database commercial license |
| bakeoff | claude-code | AWS EC2 on-demand pricing data transfer out to internet US East (N. Virginia) first 10 TB per month per GB; data transfer same availability zone free |
| bakeoff | claude-code | Stripe annual letter 2025 total payment volume private company revenue |
| bakeoff | claude-code | Stripe revenue estimate UK Companies House filing Stripe Payments Europe accounts net revenue |
| bakeoff | claude-code | Stripe annual letter 2025 total payment volume private company revenue |
| bakeoff | claude-code | Stripe revenue estimate net revenue report private company does not disclose financials |
| bakeoff | claude-code | Stripe Irish entity accounts filed Companies Registration Office revenue $5.1bn pre-tax profit 2024 |
| bakeoff | claude-code | Stripe annual letter 2025 total payment volume private company revenue estimate |
| bakeoff | claude-code | Grype offline air-gapped vulnerability database GRYPE_DB_AUTO_UPDATE false grype db import |
| bakeoff | claude-code | Dependency-Track air-gapped offline mirror NVD OSV internal mirror documentation |
Methodology
Agent search behavior: methodology
The analysis reads the stored transcripts of the web search bakeoff (Claude Code, eight arms, 18 tasks, three repetitions) and the search gap bench (Claude Code and codex, brief tasks, three repetitions per admitted task). It runs no agent and changes no grade. Pass and fail come from each cell's recorded evaluation.
scripts/analyze_search_behavior.py extracts every completed provider tool call from a transcript:
- codex MCP tool calls, and codex built-in
web_searchitems; - Claude Code
tool_useandtool_resultpairs, including vendor MCP tools, WebSearch and WebFetch.
Classification rules
- Search call: a call to a search tool (
web_search_exa,brave_web_search,tavily_search, Parallelweb_search,firecrawl_search, the Perplexity search tools, Claude Code WebSearch, and codex built-insearchactions). - Query: each entry of a multi-query call (Parallel
search_queries, codex built-in query lists) counts as one query. Each user message in Perplexity's conversational inputs also counts as one query; system and assistant messages are excluded. Parallel'sobjectiveis counted as a structured parameter whensearch_queriesis present, not as another query. - Fetch call: a page-retrieval tool call (
web_fetch_exa,tavily_extract,firecrawl_scrape, Parallelweb_fetch, Claude Code WebFetch, and codex built-inopen_page,find_in_pageand other page views). - Natural-language query: ends with "?", starts with an interrogative or auxiliary verb, or contains at least two function words from a fixed list (the, a, an, is, are, was, were, does, do, did, how, what, which, why, when, where, who, that, this, of, to, for, with, will, should, can, about, on, from, by, as, be, has, have).
- Contains a year: a token from 2000 to 2099. Quoted phrase: contains a double quote.
site:: containssite:. - Fetched without searching: the cell made at least one fetch call and no search call.
- Primary source surfaced (GAP only): the task's catalog oracle
source_url, normalised (scheme, query, fragment and trailing slash removed), appears among the URLs in any returned search result of the cell, never in the search input. It is measured only over cells that searched. A cell is unknown when any search lacks exposed results and no observed result contains the primary source. Unknown cells are excluded from both surfaced and not-surfaced pass-rate denominators. Recomputing against returned evidence leaves the historical source counts unchanged; no searched cell has unknown coverage.
The rules are deliberately simple so they can be audited. The natural-language rule undercounts descriptive keyword-free phrases and overcounts keyword strings that happen to contain function words. Readers can substitute their own rule in the script and rerun it.
Scope
Pooled figures cover the provider and native arms. Reference arms (no-search, floor and ceiling) make no search calls and are excluded from query statistics. GAP cell counts differ by harness because calibration admitted seven brief tasks for Claude Code and six for codex.
See the report, summary data and reproduction. Calibration does not apply: calibration record.
Reproduce, correct, contribute
Install Searchlight, run sew doctor, then follow each run's reproduce.md. Raw transcripts and provider
responses are not distributed, so exact replay of the published aggregates is not possible; a new run measures the same
method. If you maintain a tested service and a configuration here misrepresents it, open an issue with the configuration you
recommend: corrections are rerun and published beside the original, never silently replaced.
pipx install 'git+https://github.com/laceyenterprises/searchlight'
sew doctor
python3 scripts/check_reports.py
python3 scripts/build_site.py --check