SEARCHLIGHT

An open benchmark of what web search does for coding agents: which search arm helps an agent finish a real job, at what token cost, and how agents actually use the tools they are given.

Results to date

Two benchmarks, two coding-agent harnesses, eight retrieval arms, 567 graded provider and native cells in the leaderboard below, and a transcript analysis of how the agents searched. Every number on this page is generated from the published report data in reports/.

Searchlight results to date How the Searchlight test works, four findings with their evidence (pass rates with 95% intervals for search vendors on two benchmarks and two coding agents, and how the agents searched), and the caveats for reading them. SEARCHLIGHT Does web search help coding agents do real work? Results to date, 2026-09-29 to 2026-10-05 · 609 graded runs · 2 agents · 8 search arms · 2 benchmarks How the test works 1 A real job Search gap bench (GAP): write a brief and make a decision about an event after the model's training cutoff. Web search bakeoff (WSB): answer a documented research question. 2 Swap only the search Same agent, model and prompt in every run. Each arm adds one vendor's MCP server, pinned, at default settings. Controls: no search, and the agent's built-in search. 3 Isolated, repeated runs A fresh sandbox per run, with every tool call logged. 18–42 runs per arm, across 3 repetitions of each task. 4 Blind grading A format check, then judges that never see which arm ran. GAP pass: right decision, at least 70% of key facts, at most 25% unsupported claims. Every result below is a pass rate with a 95% Wilson interval. Where two intervals overlap, the data cannot tell those arms apart. What we found FINDING 1 Past the training cutoff, search decides the outcome Without search, both agents failed every post-cutoff brief (0%). With any of the seven search arms, Claude Code passed 62–95% and codex 72–94%, close to the 95% and 89% reached when handed the answer excerpt. SO WHAT For work that depends on recent events, search is not a tuning knob. It is the difference between failing and passing, and every vendor tested closes most of the gap. Pass rate on post-cutoff briefs (search gap bench) No search Range across the 7 search arms Answer excerpt handed over 0% 25% 50% 75% 100% Claude Code 21 runs per arm no search 0% search 62–95% handed over 95% codex 18 runs per arm no search 0% search 72–94% handed over 89% FINDING 2 No vendor wins everywhere The top arm was Parallel on Claude Code (20/21) and Brave on codex (17/18), and most intervals overlap. With 18–21 runs per arm, the bench cannot separate the leading vendors. SO WHAT Choose on cost, latency and integration, then measure on your own agent. Rankings did not carry over from one agent to the other. Claude Code claude-opus-5-5 0% 50% 100% Parallel 20/21 · 95% Firecrawl 19/21 · 90% Brave 17/21 · 81% Perplexity 16/21 · 76% Exa 15/21 · 71% Tavily 15/21 · 71% native * 13/21 · 62% codex gpt-6.1-sol 0% 50% 100% Brave 17/18 · 94% Firecrawl 16/18 · 89% Tavily 16/18 · 89% Exa 14/18 · 78% native 14/18 · 78% Parallel 13/18 · 72% Perplexity 13/18 · 72% * Claude Code native arm: the harness refused its page fetches, so it searched without reading pages. Not comparable to the vendor rows; fixed for future runs. FINDING 3 On documented knowledge, memory does most of the work In the bakeoff, answering from memory passed 32/42 (76%). The best arm reached 90%, and no arm differed from no-search at 95% confidence, while tokens per success rose from 20.9k (no-search) to as much as 89.2k (Brave). SO WHAT Search pays off where training data runs out. On stable, well-documented questions it mostly adds tokens. Pass rate on documented-knowledge tasks (web search bakeoff, Claude Code) no-search 76% 0% 25% 50% 75% 100% Firecrawl 38/42 · 90% Perplexity 37/42 · 88% Tavily 36/42 · 86% Exa 35/42 · 83% Parallel 33/42 · 79% native * 33/42 · 79% no-search 32/42 · 76% Brave 31/42 · 74% FINDING 4 Agents still search like keyword users Only 7% of gap-bench queries read as natural language, and codex leaned on site: filters (31–67% of its vendor queries). Agents often skipped search and opened URLs they remembered. Runs did better when the primary source surfaced. SO WHAT The next gains are in how agents ask and what results surface: full-question queries, and primary sources first. These are correlations across runs, not controlled effects. 7% of gap-bench search queries read as natural language; the rest were keyword strings. 36/55 Exa bakeoff runs that fetched remembered URLs without searching (Firecrawl 35/55, Tavily 34/54, Parallel 30/55). 85% vs 61% Claude Code pass rate with the primary source in its results (94/111) vs without it (20/33). Read before comparing · 18–42 runs per arm: intervals are wide, and adjacent positions are not rankings. · Each vendor ran through its own MCP server at a pinned version and default settings. Direct APIs, other modes and other tools were not tested. · Grading used validators and blinded model judges (Claude; codex as a second judge on the gap bench). Vendor spend was only partly metered. · Search behavior comes from 1629 logged queries; its links to pass rates are correlations. github.com/laceyenterprises/searchlight · data: reports/*/summary.json · method: reports/*/methodology.md
Infographic generated by scripts/build_site.py from reports/*/summary.json. The same image is site/infographic.svg.

Read this before comparing vendors

Leaderboard

Ordered by observed pass rate. The bar is the 95% interval; the tick is the observed rate. Token accounting and coverage are defined beside each benchmark's table. The leaderboard updates automatically when a new report lands in reports/; machine-readable data is in leaderboard.json.

Search gap bench: job outcomes on Claude Code (claude-opus-5-5)

7 admitted post-cutoff brief tasks × 3 repetitions = 21 cells per arm. Pass = right decision, weighted key-fact recall ≥ 0.7, unsupported claims ≤ 0.25. Interval: Wilson 95% (published). The dashed line marks the lower end of the top arm's interval; badges describe overlap of individual intervals only. Overlap does not establish statistical equivalence or test a difference between arms.

ArmPassed95% interval on a 0–100% scaleGap closureTokens per successNotes
Parallel20/21 · 95%
77–99%
1.00 (0.86–1.17)64koverlaps top interval
Firecrawl19/21 · 90%
71–97%
0.95 (0.83–1.00)116koverlaps top interval
Brave17/21 · 81%
60–92%
0.85 (0.62–1.00)93koverlaps top interval
Perplexity16/21 · 76%
55–89%
0.80 (0.50–1.00)74koverlaps top interval
Exa15/21 · 71%
50–86%
0.75 (0.47–0.95)89koverlaps top interval
Tavily15/21 · 71%
50–86%
0.75 (0.43–1.00)83koverlaps top interval
native13/21 · 62%
41–79%
0.65 (0.37–0.91)66koverlaps top interval
WebFetch refused by the harness (search only); not comparable

floor (no search): 0% · ceiling (answer excerpt): 95%

GAP tokens per success counts all agent tokens in the arm, failed cells included, divided by passing cells.

Source: Search gap bench.

Search gap bench: job outcomes on codex (gpt-6.1-sol)

6 admitted post-cutoff brief tasks × 3 repetitions = 18 cells per arm. Pass = right decision, weighted key-fact recall ≥ 0.7, unsupported claims ≤ 0.25. Interval: Wilson 95% (published). The dashed line marks the lower end of the top arm's interval; badges describe overlap of individual intervals only. Overlap does not establish statistical equivalence or test a difference between arms.

ArmPassed95% interval on a 0–100% scaleGap closureTokens per successNotes
Brave17/18 · 94%
74–99%
1.06 (0.88–1.29)111koverlaps top interval
Firecrawl16/18 · 89%
67–97%
1.00 (1.00–1.00)166koverlaps top interval
Tavily16/18 · 89%
67–97%
1.00 (0.71–1.29)129koverlaps top interval
Exa14/18 · 78%
55–91%
0.88 (0.47–1.29)146koverlaps top interval
native14/18 · 78%
55–91%
0.88 (0.47–1.20)117koverlaps top interval
Parallel13/18 · 72%
49–88%
0.81 (0.41–1.14)157koverlaps top interval
Perplexity13/18 · 72%
49–88%
0.81 (0.47–1.12)117koverlaps top interval

floor (no search): 0% (0–18%) · ceiling (answer excerpt): 89% (67–97%)

GAP tokens per success counts all agent tokens in the arm, failed cells included, divided by passing cells.

Source: Search gap bench.

Web search bakeoff: competitive research tasks on Claude Code (claude-opus-5-5)

14 competitive tasks × 3 repetitions = 42 cells per arm, regraded 2026-09-30. Mostly stable, documented knowledge; search adds less here than on post-cutoff jobs. Interval: Wilson 95% (computed from the published counts). The dashed line marks the lower end of the top arm's interval; badges describe overlap of individual intervals only. Overlap does not establish statistical equivalence or test a difference between arms.

ArmPassed95% interval on a 0–100% scaleTokens per successNotes
Firecrawl38/42 · 90%
78–96%
61.5koverlaps top interval
Perplexity37/42 · 88%
75–95%
47.7koverlaps top interval
Tavily36/42 · 86%
72–93%
61.3koverlaps top interval
Exa35/42 · 83%
69–92%
65.0koverlaps top interval
Parallel33/42 · 79%
64–88%
69.8koverlaps top interval
native33/42 · 79%
64–88%
36.3koverlaps top interval
WebFetch refused in 40 of 54 cells
no-search32/42 · 76%
61–87%
20.9kreference
Brave31/42 · 74%
59–85%
89.2koverlaps top interval

WSB tokens per success uses measured tokens from graded cells divided by measured successes. Ungraded and unmeasured cells are excluded; per-arm coverage counts were not retained, so complete accounting cannot be established.

Source: Web search bakeoff.

Test apparatus

Each cell is one task, one arm and one repetition. The runner gives the agent a fresh scratch workspace, starts the harness with exactly the tools its arm allows, records every tool call and the final answer into an evidence bundle, and grades the answer without telling the grader which arm produced it.

From catalog to report

flowchart LR
  subgraph catalog[Task catalogs]
    T1[production tasks<br/>web search bakeoff]
    T2[gap briefs<br/>post-cutoff events]
  end
  subgraph cell[One cell = task x arm x repetition]
    W[fresh scratch<br/>workspace]
    H[harness<br/>Claude Code or codex]
    M[metering proxy]
    V[(vendor MCP server<br/>pinned version)]
  end
  subgraph grade[Grading]
    D[deterministic<br/>validators]
    J[blinded judges<br/>vs captured sources]
  end
  T1 & T2 --> W --> H
  H <-->|tool calls| M <--> V
  H --> B[(evidence bundle<br/>transcript, final answer,<br/>arm audit, metrics)]
  B --> D & J --> R[reports/*/summary.json] --> S[this site]

What each arm can touch

An arm differs from another only in its retrieval surface. Claude Code is started with --strict-mcp-config, so the only MCP server it can reach is the arm's vendor server, and --tools limits its built-ins. Codex runs read-only with its shell, browser, app and plugin features disabled. After each cell an arm audit compares every observed tool call with the arm contract; a call outside the contract marks the cell contaminated.

flowchart TB
  A[arm contract] --> CC[Claude Code spawn<br/>--strict-mcp-config<br/>--tools: the arm's built-ins<br/>--allowedTools: the same set]
  A --> CX[codex spawn<br/>exec --sandbox read-only<br/>--disable shell_tool, browser_use,<br/>apps, plugins, ...]
  CC --> P1[provider arm:<br/>one vendor MCP server<br/>+ ToolSearch + Read]
  CC --> N1[native arm:<br/>WebSearch + WebFetch + Read]
  CC --> F1[no-search / floor / ceiling:<br/>no tools]
  CX --> P2[provider arm:<br/>one vendor MCP server]
  CX --> N2[native arm:<br/>codex --search]
  CX --> F2[no-search / floor / ceiling:<br/>no tools]
  P1 & N1 & F1 & P2 & N2 & F2 --> AU[arm audit:<br/>every observed tool call<br/>checked against the contract]

Claude Code invocation

claude --print --output-format stream-json --verbose --no-session-persistence \
  --model claude-opus-5-5 --mcp-config <cell>/claude-mcp.json --strict-mcp-config \
  --tools ToolSearch,Read --allowedTools 'mcp__<vendor>__*'        # provider arm
  --tools WebSearch,WebFetch,Read --allowedTools WebSearch,WebFetch   # native arm (fixed)
  --tools '' --allowedTools ''                                       # no-search, floor, ceiling

codex invocation

codex [--search] exec --json --ephemeral --skip-git-repo-check --sandbox read-only \
  --model gpt-6.1-sol --disable shell_tool --disable unified_exec \
  --disable browser_use --disable browser_use_external --disable computer_use \
  --disable in_app_browser --disable apps --disable plugins --disable remote_plugin ...
# --search only on the native arm; provider arms add one MCP server via the cell config

Per-cell limits

LimitGap benchWeb search bakeoff
Wall clock per cell1,200 s300 to 1,200 s (set per task)
Total tokens per cell1,000,000120,000 to 350,000 (set per task)
Provider calls per cell6015 to 50 (set per task)
Harness boot timeout120 s120 s
Repetitions3 (battery), 5 (calibration)3

How a gap-bench task is admitted

A task counts only if it is genuinely beyond the model: with no search the agent must fail, and with the answer excerpt handed over it must succeed. Calibration is repeated for each harness and model.

flowchart LR
  T[candidate brief<br/>event after 2026-01-01] --> FL[floor: no search<br/>5 runs]
  T --> CE[ceiling: answer excerpt<br/>in the prompt, 5 runs]
  FL --> Q{floor passes ≤ 1<br/>and ceiling passes ≥ 4?}
  CE --> Q
  Q -- yes --> AD[admitted for this<br/>harness and model]
  Q -- no --> RJ[rejected]
  AD --> BAT[battery: every arm x 3]
  BAT --> GC[gap closure =<br/>arm - floor / ceiling - floor]
gap closure  =  (arm pass rate − floor pass rate) / (ceiling pass rate − floor pass rate)

   floor            arm                 ceiling
   0% ●─────────────────●──────────────────● 95%
       └──── closed ────┘└──── remaining ────┘

How answers are graded

sequenceDiagram
  participant A as Agent answer
  participant V as Schema validator
  participant S as Source capture
  participant J1 as Judge 1 (Claude Code)
  participant J2 as Judge 2 (codex, gap bench)
  A->>V: JSON deliverable (brief, claims, citations)
  V-->>A: reject malformed answers
  A->>S: cited URLs
  S->>J1: readable source text + rubric (blinded to arm)
  S->>J2: same inputs
  J1-->>A: key-fact recall, unsupported claims, decision
  J2-->>A: independent labels, agreement recorded
  Note over A,J2: pass = right decision AND recall >= 0.7 AND unsupported <= 0.25

Arm configurations

Pinned in config/provider-arm-tools.yaml. Every tool a server advertised was available to the agent; no vendor tool was hidden.

ArmInterfaceVersion / endpointTools exposed to the agent
Bravestdio MCP server (npx)@brave/brave-search-mcp-server@2.1.48: brave_image_search, brave_llm_context, brave_local_search, brave_news_search, brave_place_search, brave_summarizer, brave_video_search, brave_web_search
Exastdio MCP server (npx)exa-mcp-server@3.4.12: web_fetch_exa, web_search_exa
Firecrawlstdio MCP server (npx)firecrawl-mcp@3.26.029: firecrawl_agent, firecrawl_agent_status, firecrawl_check_crawl_status, firecrawl_crawl, firecrawl_credit_usage, firecrawl_developer_search, firecrawl_extract, firecrawl_feedback, firecrawl_find_tools, firecrawl_interact, firecrawl_interact_stop, firecrawl_map, firecrawl_monitor_check, firecrawl_monitor_checks, firecrawl_monitor_create, firecrawl_monitor_delete, firecrawl_monitor_get, firecrawl_monitor_list, firecrawl_monitor_run, firecrawl_monitor_update, firecrawl_parse, firecrawl_research_inspect_paper, firecrawl_research_read_paper, firecrawl_research_related_papers, firecrawl_research_search_github, firecrawl_research_search_papers, firecrawl_scrape, firecrawl_search, firecrawl_search_feedback
Parallelhosted MCP via mcp-remotemcp-remote@0.14.3 → https://search.parallel.ai/mcp-oauth2: web_fetch, web_search
Perplexitystdio MCP server (npx)@perplexity-ai/mcp-server@1.3.04: perplexity_ask, perplexity_reason, perplexity_research, perplexity_search
Tavilystdio MCP server (npx)tavily-mcp@0.2.225: tavily_crawl, tavily_extract, tavily_map, tavily_research, tavily_search
native (Claude Code)harness built-inClaude Code 2.1.282WebSearch, WebFetch (refused before 2026-10-06), Read
native (codex)harness built-incodex --searchbuilt-in web search and page views

Runs

Each run is the recorded report, transcribed into summary.json and checked by scripts/check_reports.py.

Web search bakeoff (2026-09-29)

Full recorded report

WSB live battery — initial results (2026-09-29)

Correction (2026-10-06): the native arm's page fetches were mostly refused. The arm exposed WebSearch and WebFetch but pre-approved only WebSearch, and headless Claude Code refuses tools that are not pre-approved. WebFetch was refused in 40 of 54 native cells (110 calls). Provider arms were unaffected. The native rows therefore understate Claude Code's own web tools. The configuration is fixed and pinned by a test; the recorded numbers below are unchanged. Details: agent search behavior.

Pack: web-search-bakeoff (WSB) on the Search Evaluation Workbench (SEW). Harness / model: claude-code, Claude Opus 5.5 (claude-opus-5-5[1m]), authenticated execution on the original host. Status: historical initial grading; headline and Results tables are superseded by the 2026-09-30 regrade, where Firecrawl leads at 90%. The numbers below are graded with SEWBENCH-01 applied to the stored deliverables. The same deliverables graded without those fixes are summarised in Failure dive and were materially wrong.

Headline

Historical figures under the initial grader: rankings and cost lower bounds here apply only to that grading, not the later operator-ruled catalog.

On the 14 competitive production tasks (42 cells per arm), every search arm except Brave scores above answering without search:

  • Firecrawl and Perplexity lead at 83%. Exa is at 81%, Tavily 79%, native web search and Parallel 76%, no-search 71% and Brave 69%.
  • The 95% intervals overlap. At n=42 per arm the ordering is directional, not conclusive: no arm's difference from no-search or from native has an interval that excludes zero.
  • Search matters most on tasks that need live data. On the cloud egress pricing table only Firecrawl and Perplexity went 3/3 (Brave 2/3, the rest 0/3). The AWS pricing pages need JavaScript rendering, and most arms reported the intra-AZ cell as unavailable, honestly.
  • Among fully priced arms, native search costs least per success ($0.167), followed by Tavily ($0.206). Perplexity (at least $0.148; 37/42 cells priced) and Firecrawl (at least $0.159; 39/42 priced) are lower bounds, excluded from the cost ranking; Perplexity's ask spend is unmetered. Brave and no-search are also partially priced (at least $0.28 each); no-search had a few long runaway answers.
  • Brave is last, failing unanswerable-nonexistent-postgres-guc and upstream-diagnosis-node-openssl3-md4 in all three repetitions.

What ran

RunArmsCellsWindow (UTC)
first arm groupno-search, native, brave, tavily2162026-09-29 03:29 – 13:26
second arm groupexa, parallel-web, firecrawl, perplexity2162026-09-29 13:39 – 22:49
  • Tasks: all 18 production tasks (production task catalog), × 3 repetitions.
    • 14 are competitive.
    • 4 are expected_fail by design: no retrieval surface can answer them, and the honest deliverable is a partial answer or a decline.
  • Arms:
    • Provider arms expose one vendor MCP server each, and nothing else.
    • The native arm uses the harness's own web search.
    • The no-search arm has no network tools.
  • Pinned MCP servers:
  • Grading:
    • Deterministic validators where a task has one.
    • Otherwise a blinded claude-code judge scoring the task's rubric.
    • bakeoff grade --regrade re-scores the stored deliverables. Nothing was re-run.
  • Cost: model tokens are priced at the published Opus 5.5 rates. Vendor spend is metered per MCP call where a tariff exists. It is unknown for Exa, Parallel and Perplexity's ask tool, which report no spend over MCP. In the tables, partial is the priced cells' spend divided by all of the arm's successes. Unpriced cells can only add spend, so each partial figure is a lower bound on the full-arm $/success, and these arms are excluded from the cost ranking. bakeoff report divides by the successes among priced cells instead, which gives $0.172 for Perplexity, $0.163 for Firecrawl and $0.29 for Brave on the competitive set, and no figure for Exa or Parallel on the expected-fail set, where every success is in an unpriced cell.

Results

All Results tables below use the historical initial grading. The ungraded column counts delivered answers that could not be graded. Cells that exhausted their budget before delivering are counted under budget-exhausted and as failures, rather than under ungraded.

Competitive tasks (14 tasks × 3 repetitions)
armpass95% CI$/successpricedtokens/cellprovider callsbudget-exhaustedungraded
firecrawl35/42 (83%)69–92%partial $0.159 (39/42 priced)39/4255k21920
perplexity35/42 (83%)69–92%partial $0.148 (37/42 priced)37/4242k14500
exa34/42 (81%)67–90%unpriced0/4254k10500
tavily33/42 (79%)64–88%$0.20642/4252k11900
native32/42 (76%)61–87%$0.16742/4228k25600
parallel-web32/42 (76%)61–87%unpriced0/4254k9010
no-search30/42 (71%)56–83%partial $0.285 (41/42 priced)41/4215k000
brave29/42 (69%)54–81%partial $0.280 (41/42 priced)41/4265k15400
Competitive deltas against controls

Each arm against each control over the same 14 tasks × 3 repetitions. Intervals are 95% Newcombe (hybrid Wilson score) intervals for a difference of two independent proportions. They ignore the task pairing, so they are conservative. Token ratios compare mean tokens per cell. Provider arms and their controls come from two runs about ten hours apart (see Caveats).

armvs no-search: pass Δ (95% CI)vs no-search: tokensvs native: pass Δ (95% CI)vs native: tokens
firecrawl+12 pp (−6 to +29)3.49×+7 pp (−10 to +24)1.95×
perplexity+12 pp (−6 to +29)2.63×+7 pp (−10 to +24)1.47×
exa+10 pp (−9 to +27)3.39×+5 pp (−13 to +22)1.90×
tavily+7 pp (−11 to +25)3.29×+2 pp (−15 to +20)1.84×
native+5 pp (−14 to +23)1.79×——
parallel-web+5 pp (−14 to +23)3.44×+0 pp (−18 to +18)1.92×
no-search——−5 pp (−23 to +14)0.56×
brave−2 pp (−21 to +17)4.13×−7 pp (−25 to +12)2.31×
Expected-fail tasks (4 tasks × 3 repetitions)

Passing here means an honest partial answer or decline, which the validator can confirm.

armpass95% CI$/successpricedtokens/cellprovider callsbudget-exhaustedungraded
exa5/12 (42%)19–68%partial $0.046 (6/12 priced)6/1267k3000
no-search3/12 (25%)9–53%partial $2.827 (9/12 priced)9/1243k000
native3/12 (25%)9–53%$0.27312/1216k2800
brave3/12 (25%)9–53%partial $0.763 (11/12 priced)11/1266k4510
firecrawl3/12 (25%)9–53%partial $0.672 (10/12 priced)10/1275k10420
perplexity3/12 (25%)9–53%partial $0.197 (9/12 priced)9/1236k3000
parallel-web2/12 (17%)5–45%partial $0.116 (6/12 priced)6/1260k3410
tavily1/12 (8%)1–35%partial $1.923 (11/12 priced)11/1264k4910
Per task
taskoutcomeno-searchnativebravetavilyexaparallel-webfirecrawlperplexity
change-detection-kubernetes-dockershim-removalcompetitive2/32/31/33/32/33/33/33/3
change-detection-openapi-30-to-31competitive1/31/30/30/32/32/30/31/3
competitive-table-cloud-egress-pricingcompetitive0/30/32/30/30/30/33/33/3
competitive-table-copyleft-obligationscompetitive3/33/33/33/33/33/33/33/3
list-build-kubernetes-122-api-removalscompetitive3/33/33/33/32/32/33/33/3
list-build-python313-pep594-removalscompetitive3/33/33/33/33/33/33/33/3
list-build-sca-tool-shortlistcompetitive0/32/32/31/33/33/31/33/3
multi-hop-cve-to-fixed-releasecompetitive3/31/33/33/33/32/32/31/3
multi-hop-python-feature-pepscompetitive3/33/33/33/33/33/33/33/3
unanswerable-nonexistent-postgres-guccompetitive2/33/30/33/32/32/33/31/3
unanswerable-private-company-audited-arrcompetitive3/33/33/33/33/31/33/33/3
unanswerable-python4-ga-datecompetitive3/33/33/33/33/33/33/33/3
upstream-diagnosis-node-openssl3-md4competitive1/32/30/32/32/32/32/32/3
upstream-diagnosis-urllib3-libressl-import-errorcompetitive3/33/33/33/33/33/33/33/3
change-detection-unversioned-runtime-driftexpected_fail0/30/30/30/30/30/30/30/3
competitive-table-unpublished-enterprise-pricingexpected_fail3/33/33/31/32/31/32/33/3
multi-hop-npm-semver-dependentsexpected_fail0/30/30/30/33/31/31/30/3
upstream-diagnosis-unindexed-ordering-regressionexpected_fail0/30/30/30/30/30/30/30/3

Findings

1. Search earns its cost on live or long-tail data, not on stable knowledge.

  • Tasks whose answers are stable and well known were 3/3 in every arm, including no-search: copyleft obligations, PEP 594 removals, Python feature PEPs and the urllib3 diagnosis.
  • The separation comes from the tasks that need current data:
    • the egress pricing table;
    • the SCA tool shortlist (no-search 0/3; Exa, Parallel and Perplexity 3/3);
    • the dockershim change-detection task.

2. Rendering matters for vendor pricing pages. Firecrawl (which renders JavaScript) and Perplexity were the only arms to read AWS's intra-AZ transfer pricing. The others honestly reported it unavailable, and the validator fails an unavailable cell. 3. Declining is not the default for search arms. unanswerable-nonexistent-postgres-guc separates the arms: Brave 0/3 and Perplexity 1/3 against 3/3 for native, Tavily and Firecrawl. 4. Exa handles impossible tasks best (5/12 honest partials on the expected-fail set). It is the only arm to go 3/3 on multi-hop-npm-semver-dependents, where the honest move is to state the method and its limits rather than invent a ranking. 5. Provider arms use more reported tokens per cell. The MCP provider arms report 42k–65k tokens per cell against 15k for no-search. Native search reports 28k. These totals do not separate input from output tokens, which have different prices, and some vendor spend is unknown; the totals alone cannot establish what dominates cost per success.

Failure dive

Two tasks were initially at 0/24 across all arms; a third mostly failed, and the judged SCA task was mostly ungraded. All four exposed bench bugs. Each was confirmed by reading the deliverables, and each is fixed in the bench fix:

TaskAs first gradedGraded with SEWBENCH-01Cause
unanswerable-python4-ga-date0/2424/24The decline rejected any date, including "as of 2026-09-28" and other releases' dates.
unanswerable-private-company-audited-arr3/23 graded22/23 gradedThe prompt asks for how circulating revenue figures were derived; any dollar figure failed. The 24th cell (Parallel) hit its budget before delivering, is not graded, and counts as a failure in the tables, so the per-task row reads 22/24.
upstream-diagnosis-urllib3-libressl-import-error0/2424/24A URL-only check on a field every arm filled in prose; the URLs were in evidence_urls.
list-build-sca-tool-shortlist (judged)mostly ungradedgradedThe judge sometimes fences its JSON; the default profile name also counted as an arm-identity leak.

With the fixes, the competitive pass rate rose between 14 and 29 points per arm (no-search +14, Exa +29). It rose most for search arms, which had been penalised for citing and explaining evidence. The first reading of this battery, that no-search was about as good as any provider, was an artifact of these bugs.

Genuine failures that remain:

  • Egress pricing: unavailable AWS intra-AZ cells, as above.
  • change-detection-openapi-30-to-31: most misses omit example→examples from the breaking changes. OAS 3.1 deprecates example rather than removing it. The tables above still require it; the operator has since ruled it is not a break (see the update below).
  • Budget exhaustion: 8 cells. Six are Firecrawl or Parallel on long list-building and enumeration tasks; Brave and Tavily have one each.
  • By design: the two remaining expected_fail tasks (unversioned-runtime-drift, unindexed-ordering-regression) are 0/3 everywhere.

Codex (partial)

The codex harness (gpt-6-sol) ran one competitive task on 2026-09-28 before its weekly quota ran out:

  • Egress pricing: native 2/3, no-search 1/3. Brave and Tavily went 0/3 on the token budget, reading 393k–448k tokens per cell.
  • Expected-fail tasks: codex's search arms burned the full provider-call budget on multi-hop-npm-semver-dependents instead of declining.

A full codex battery waits for the quota reset (2026-10-04 12:52Z).

Caveats

  • Sample size: n=42 per arm on the competitive set. Adjacent arms are within each other's intervals.
  • Timing: the two runs were about ten hours apart. Same harness, model, tasks and budgets.
  • Pricing: vendor spend is unpriced for Exa and Parallel, and for Perplexity's ask tool. Their $/success counts model tokens plus metered calls only.
  • Grading: the judged tasks use a single claude judge, so agreement is not measured.

Update (2026-09-30): operator decisions

  • OpenAPI: example→examples is a deprecation, not a break, so the task no longer requires it. The task also stops requiring a construct labelled as the JSON Schema dialect: the dialect has its own field, and complete answers list its consequences instead. Both changes are in the bench fix. On replay, the task goes from 7/24 to 22/24. The two remaining misses cite hosts outside the allowlist.
  • Codex native search is priced at Brave's per-request rate as a proxy for its backing vendor. Claude native search keeps Anthropic's published rate.
  • Tokens and outcome rate are reported together. Tokens per success come from the bench fix.

Competitive pass rates regraded with the corrected catalog at the recorded grader revision:

armpasstokens/success
firecrawl38/42 (90%)61.5k
perplexity37/42 (88%)47.7k
tavily36/42 (86%)61.3k
exa35/42 (83%)65.0k
native33/42 (79%)36.3k
parallel-web33/42 (79%)69.8k
no-search32/42 (76%)20.9k
brave31/42 (74%)89.2k

The historical tokens/cell column averages measured total-billable tokens over attempted cells with measured usage; unknown or estimated usage is excluded, never treated as zero. The update's tokens/success sums measured tokens over graded cells and divides by measured successes, as defined by the bench fix; it excludes ungraded and unmeasured cells and is not a full-arm lower bound. The per-arm measured coverage counts were not retained here. Multiplying a measured-cell mean by all 42 attempts need not reproduce it. The previously quoted fresh/cache/output means are withdrawn: their cell coverage and bucket accounting were not recorded here and they do not reconcile for Firecrawl, Parallel or no-search. They cannot support a token-mix comparison from this document.

The output-reduction claim is also withdrawn. The quoted no-search mean included long runaway answers; without a median or an outlier-excluded comparison, this battery does not establish a typical reduction or a causal effect of search on output length. Among the six MCP providers, a higher pass rate goes with fewer tokens per success: Perplexity is best on tokens, and Brave is worst on both. Spearman ρ ≈ −0.83 (pass rate versus tokens/success, n=6) is suggestive, not established, and is not significant at α=0.05. No arm's pass-rate difference from no-search has an interval that excludes zero. Firecrawl comes closest at +14 points (−2 to +30).

Methodology

WSB methodology

The experiment used Claude Code with claude-opus-5-5[1m], all 18 production tasks, eight tool arms and three repetitions. The two arm groups ran about ten hours apart. Arms differ in exposed search tools: no-search, native, or exactly one provider MCP server. Budgets and contamination checks apply to each cell.

The 14 competitive tasks and four expected-fail tasks have separate denominators. Expected-fail success means an honest partial answer or decline. Deterministic validators grade structured tasks; a blinded Claude judge grades rubric tasks. Single-judge agreement was not measured. Pass-rate intervals are Wilson 95%; control deltas use independent-proportion Newcombe intervals, ignoring pairing. Adjacent rankings are not statistically established.

The original bench corrections fixed overly strict date/dollar decline checks, a prose-versus-URL field mismatch and judge JSON parsing/blinding. The later regrade relaxed incorrect OpenAPI breaking-change requirements. Both grade the same stored answers; neither reran agents. REPORT.md retains initial and updated tables with their original precision.

Historical cost lower bounds sum priced spend and divide by all successes. The report command instead uses priced-cell successes, so its output can differ. Unknown vendor spend is not zero; partially priced arms are excluded from cost ranking. Updated tokens per success use measured tokens from graded cells divided by measured successes; coverage counts were not retained. Do not multiply the historical tokens/cell means by attempted-cell counts to derive that update. The output-reduction and fresh/cache/output comparisons were withdrawn.

See recorded results, summary and reproduction. Calibration was not part of this experiment: calibration record.

Files: reports/2026-09-29-web-search-bakeoff/

Search gap bench (2026-10-03)

Full recorded report

GAP battery results, 2026-10-03: does search change the outcome of the job?

Correction (2026-10-06): the Claude Code native arm could not fetch pages. The arm exposed WebSearch and WebFetch but pre-approved only WebSearch, and headless Claude Code refuses tools that are not pre-approved. Every WebFetch call in the native arm was refused (50 calls across all 21 native cells). Each provider arm could use its vendor's fetch tool, so the Claude Code native row measures search without page fetching and is not comparable to the provider rows. The codex native arm was unaffected. The configuration is fixed and pinned by a test; the recorded numbers below are unchanged. Details: agent search behavior.

The first full Search Gap Bench (GAP) battery ran on 2026-10-03, operator-approved: brief tasks on two harnesses, nine arms each, three repetitions.

  • GAP grades the job, not the retrieval.
  • A task counts only after calibration shows its knowledge gap: the no-search arm (floor) passes at most 1 of 5, and the arm given the answer excerpt (ceiling) passes at least 4 of 5.
  • Each search arm is scored by gap closure: how much of the floor-to-ceiling distance it covers.
  • Tokens are reported beside every outcome.

Headline. On these tasks, search decides the outcome.

  • The floor passed 0% on both harnesses, and every search arm passed 62–95%.
  • Which provider is best depends on the harness: Parallel led on claude-code and Brave on codex.
  • Pass rate did not predict cost: providers with similar pass rates differed almost twofold in tokens per success.

These are the recorded results after the Codex reruns. Claude-code provider tool availability remains unverified in the published evidence: no transcript audit for those cells is documented, and they lack availability records. Its provider rankings are provisional pending that audit. In the first codex pass, 31 of 108 provider cells ran without their search tool. They were re-run on 2026-10-04 under the GAPMCP-01 fix, every re-run cell loaded its tool, and the codex tables below include them (see "Execution changes").

Setup

Harnessesclaude-code on claude-opus-5-5, codex on gpt-6.1-sol
Armsfloor (no search), ceiling (answer excerpt in the prompt), native, Brave, Tavily, Exa, Parallel, Firecrawl, Perplexity
Tasks10 brief tasks from GAP-09: 6 decision briefs and 4 research briefs. The code corpus (GAP-07) merged during the run and gets its own battery.
Repetitions5 per reference arm in calibration; 3 per arm in the battery
GradingDecision correctness, weighted key-fact recall (≥ 0.7) and unsupported-claim rate (≤ 0.25), judged by two blinded judges (claude-code primary, codex secondary) against captured primary sources
Spendabout 50M tokens in total: battery agents 21.1M; battery judges 15.3M; calibration agents 4.2M; calibration judges 9.2M

Calibration

Final admissions under GAPRULE-01. Each cell shows floor passes and ceiling passes out of 5.

TaskFamilyclaude-codecodex
billing-reportingresearchFloor: 0, Ceiling: 5 — admittedFloor: 0, Ceiling: 5 — admitted
classroom-termdecisionFloor: 0, Ceiling: 5 — admittedFloor: 0, Ceiling: 5 — admitted
models-productiondecisionFloor: 0, Ceiling: 5 — admittedFloor: 0, Ceiling: 5 — admitted
npm-token-scoperesearchFloor: 0, Ceiling: 5 — admittedFloor: 0, Ceiling: 5 — admitted
org-migrationdecisionFloor: 0, Ceiling: 5 — admittedFloor: 0, Ceiling: 4 — admitted
python-metadataresearchFloor: 0, Ceiling: 5 — admittedFloor: 0, Ceiling: 5 — admitted
tarfile-uploaddecisionFloor: 0, Ceiling: 5 — admittedFloor: 0, Ceiling: 0 — rejected
cli-trustdecisionFloor: 0, Ceiling: 2 — rejectedFloor: 0, Ceiling: 2 — rejected
spark-editingdecisionFloor: 0, Ceiling: 1 — rejectedFloor: 0, Ceiling: 0 — rejected
python-security-marchresearchFloor: 0, Ceiling: 0 — rejectedFloor: 0, Ceiling: 1 — rejected

Every floor scored 0, so every task has a real knowledge gap. All rejections are ceiling failures, which point at task or rubric defects:

  • spark-editing: the heaviest key fact is not what the prompt asks about.
  • python-security-march: the excerpt omitted the date that its top fact requires. The excerpt fix is pending for the next battery; GAPRULE-01 changed the unsupported-claim limit, not the excerpt.
  • cli-trust: with three claims, a single unsupported claim fails the brief.
  • tarfile-upload on codex: the codex briefs miss key facts.

Gap closure

claude-code (7 tasks, 21 cells per arm)
ArmGap closure (95% CI)Pass rate (Wilson 95% CI)Tokens per successvs floor (exact McNemar)
Parallel1.00 (0.86–1.17)95% (77–99%)64kp < 0.0001
Firecrawl0.95 (0.83–1.00)90% (71–97%)116kp < 0.0001
Brave0.85 (0.62–1.00)81% (60–92%)93kp < 0.0001
Perplexity0.80 (0.50–1.00)76% (55–89%)74kp < 0.0001
Exa0.75 (0.47–0.95)71% (50–86%)89kp < 0.0001
Tavily0.75 (0.43–1.00)71% (50–86%)83kp < 0.0001
native0.65 (0.37–0.91)62% (41–79%)66kp = 0.0002
ceiling (reference)1.0095%4k
floor (reference)0.000%—

By family:

  • Research briefs: every provider arm passed 100% (native 89%).
  • Decision briefs: they separate the providers. Parallel 92% and Firecrawl 83% lead; Brave scored 67%, Perplexity 58%, Exa and Tavily 50% each, and native 42%.
codex (6 tasks, 18 cells per arm)
ArmGap closure (95% CI)Pass rate (Wilson 95% CI)Tokens per successvs floor (exact McNemar)
Brave1.06 (0.88–1.29)94% (74–99%)111kp < 0.0001
Firecrawl1.00 (1.00–1.00)89% (67–97%)166kp < 0.0001
Tavily1.00 (0.71–1.29)89% (67–97%)129kp < 0.0001
Exa0.88 (0.47–1.29)78% (55–91%)146kp = 0.0001
native0.88 (0.47–1.20)78% (55–91%)117kp = 0.0001
Parallel0.81 (0.41–1.14)72% (49–88%)157kp = 0.0002
Perplexity0.81 (0.47–1.12)72% (49–88%)117kp = 0.0002
ceiling (reference)1.0089% (67–97%)12k
floor (reference)0.000% (0–18%)—

No search arm beat native significantly; the largest gain was Brave at +16.7 points (p = 0.25).

By family:

  • Research briefs: every arm passed 100%, except Exa at 89%.
  • Decision briefs: they separate the providers. Brave scored 89%; Firecrawl and Tavily 78% each; Exa 67%; native 56%; Parallel and Perplexity 44% each.

The first pass's estimates that excluded the 31 tool-less cells were close to the re-run results:

ArmEstimateAfter re-run
Brave93%94%
Firecrawl85%89%
Tavily85%89%
Exa76%78%
Parallel67%72%
Perplexity62%72%

What this shows

1. On these tasks, search changes the outcome. Without search, both harnesses failed every brief. With a good provider they reached the answer-excerpt arm's level. The difference against the floor is significant for every search arm on both harnesses (exact McNemar, p ≤ 0.0002). 2. The best provider depends on the harness. On claude-code, Parallel (95%) and Firecrawl (90%) led. On codex, Brave led (94%), ahead of Firecrawl and Tavily (89% each).

  • Parallel tied for last on codex (72%), even with its tool loaded in every cell. On codex decision briefs it passed 44%, against 92% on claude-code.
  • So a provider's ranking on one harness does not carry over to another.

3. Pass rate does not predict cost between providers. Tokens per success counts every agent token in the arm, failed cells included, divided by passing cells.

  • The hypothesis that better search shows up as both lower token spend and higher success holds against no search, which spent tokens and passed nothing. It also holds for Parallel on claude-code, which led on both measures (64k tokens per success).
  • It does not hold as a ranking. Among the seven search arms, the rank correlation between pass rate and tokens per success was +0.13 on claude-code and −0.20 on codex; negative means higher-passing arms cost less. With seven arms, neither value is distinguishable from zero.
  • Per-cell spend varied widely: 41k (native) to 105k (Firecrawl) on claude-code, and 84k (Perplexity) to 148k (Firecrawl) on codex. Firecrawl passed 90% at 116k tokens per success, against Parallel's 95% at 64k.
  • Choosing a provider means weighing pass rate and cost separately.

4. Decision briefs separate the providers; research briefs mostly don't. Research briefs passed at or near 100% on both harnesses. Decision briefs hinge on one recent fact, and they spread the search arms from 42% to 92% on claude-code and from 44% to 89% on codex. 5. Native search was not the best choice on either harness. It came last among search arms on claude-code (62%). On codex it was mid-pack (78%, level with Exa), behind Brave, Firecrawl and Tavily, though no arm beat it significantly there.

Changes made during the run

The first live runs surfaced several bench defects. Two changed how cells execute, and two changed how cells are graded. The timeline was recorded from the battery change log and cell start events.

Execution changes

Every cell in the tables above ran under both of these fixes:

  • GAPBUDGET-01: brief budgets became runaway caps (1M tokens, 60 calls). The old 30k cap was below the median codex search cell.
    • It was part of the pinned export before any calibration or battery cell ran.
    • No cell ran under the old caps.
  • GAPARM-01: brief search arms resolve the GAP catalog. Before it, no brief search arm could run, so no search-arm cell predates it.
    • It was applied at 18:30:04Z on 2026-10-03.
    • The earliest search-arm cell started at 18:30:29Z (codex round 1). claude-code started at 19:12Z.

The tool-isolation remediation, which forbids the GAP catalog to every GAP cell, was applied at 18:50:40Z and the codex battery was restarted to load it.

  • 18 codex round-1 cells started before it: classroom-term repetitions 1 and 2, all arms.
  • No catalog access in them. Their transcripts contain only search calls and agent messages, with no file reads or commands, so the catalog was not read.
  • They were kept, not re-run.
  • Every other cell ran after the remediation.

The codex provider re-run (2026-10-04). In the first pass, 31 of codex's 108 provider cells made no provider call:

ArmCells without a provider call
Parallel12
Firecrawl5
Tavily5
Perplexity5
Brave3
Exa1
  • Cause: GAPMCP-01 reproduced codex's 30-second default MCP startup timeout. A 35-second server fails under it and starts with a 120-second allowance.
  • How the cells were selected: the 31 cells were reclassified from an audit of their transcripts. All of them had zero provider calls and a list_mcp_resources probe; 29 of the 31 also said in their answer that no retrieval was available. The private audit is not distributed.
  • How they were re-run: with --rerun-unavailable under GAPMCP-01 (a 120-second startup allowance, required=true and pre-warmed pinned packages), using the same matrix, calibration and catalog.
  • Every re-run cell loaded its tool. Each one recorded a completed initialize and tool listing.
  • A second fix was needed. The first re-run attempt misclassified working cells, because the MCP meter could not load its core (SEV2 METERAVAIL, below). Making the metering library available in each provider server's environment fixed it, and those cells were re-run again.

These 31 cells ran under the startup allowance; the other 77 provider cells ran without it but had their tool. No agent-side setting changed.

Grading changes

These changed only how stored answers were judged. Every affected cell was regraded from the agents' stored answers, with no agent re-runs:

  • GAPJUDGE-01: judges read readable text instead of raw HTML, and each source once. This cut judge tokens per cell from about 250k to about 15k. It also covered:
    • safe redirects;
    • gzip decoding (python.org served gzip unconditionally, which garbled every python.org citation);
    • over-budget sources withheld instead of the whole cell going ungraded;
    • the blinding leak check reading decoded text, and generic arm labels no longer counted as leaks.
  • GAPRULE-01, operator-approved: the unsupported-claim limit was raised from 0.10 to 0.25. Ceiling agents see a short excerpt, and their hedges about it ("the notice does not specify X") were contradicted by the full pages the judges read. All stored grades were recomputed from their saved judge labels; 49 outcomes flipped, all from fail to pass.

Open issues

  • SEV2 GAPMCP (fix GAPMCP-01, the bench fix): resolved. The 31 affected codex cells were re-run and the codex table updated.
  • SEV2 METERAVAIL (fix METERAVAIL-01): provider-call metering was silently off for this battery on both harnesses. The meter runs inside the harness's restricted MCP child environment, and from the exported workbench tree it could not load its core.
    • Vendor spend per cell was therefore not metered.
    • Only the 31 re-run cells have availability records; the report's "availability coverage" column counts those. Claude-code provider availability has not been verified in the published evidence.
  • SEV2 ARGVKEY (fix ARGVKEY-01): the Parallel API key was passed in mcp-remote's process arguments. The battery configuration now passes it through a 0600 header file.
  • The claude-code floor arm sometimes writes pretend searches as text when it has no tools. It is a legitimate failure, but slow and expensive (up to about 130k output tokens in 13 minutes).
  • Next battery:
    • the code corpus (GAP-07: 16 code tasks and 4 controls);
    • the python-security-march excerpt fix;
    • rubric fixes for spark-editing and cli-trust.
Methodology

GAP methodology

GAP grades completed jobs against captured primary sources. The initial corpus contained ten briefs: six decision and four research. Calibration is per harness and model: five repetitions of floor (no search) and ceiling (answer excerpt). Admission requires floor at most one pass and ceiling at least four passes. Seven Claude tasks and six Codex tasks were admitted. Code tasks were not run in this battery. The committed calibration file is an admission summary, not a runtime calibration file with captured oracle sources.

Each admitted task ran three times on nine arms. Brief grading requires decision correctness, weighted key-fact recall at least 0.7 and unsupported-claim rate at most 0.25. Two blinded judges used Claude Code as primary and Codex as secondary. Gap closure is (arm pass rate − floor pass rate) / (ceiling pass rate − floor pass rate); it can exceed one. Wilson 95% intervals accompany pass rates. Gap-closure intervals use a seeded paired task bootstrap with 10,000 resamples. Exact McNemar tests pair task and repetition, against floor and native. The findings publish floor comparisons and the largest Codex native comparison, not every native test.

Tokens per success sum all agent tokens in an arm, including failures, divided by passes. Judge and calibration spend is separate. A zero-success floor has undefined tokens per success. Vendor spend was unmetered; tokens are not dollar cost.

GAPJUDGE-01 changed readable source extraction, redirects, gzip, source budget handling and blinding. GAPRULE-01 raised the unsupported-claim limit from 0.10 to 0.25; recomputing saved judge labels flipped 49 failures to passes. These grading changes did not rerun agents. Execution changes and the later rerun of 31 Codex cells with unavailable provider tools are preserved separately in REPORT.md.

See recorded results, summary, admission counts and reproduction.

Files: reports/2026-10-03-search-gap-bench/

How agents used search (2026-10-05)

Full recorded report

How agents used search: query style, fetch-from-memory and primary-source retrieval (2026-10-05)

This report analyses the stored transcripts of the two published batteries, the web search bakeoff and the search gap bench. No agent was rerun. It asks how agents used the search tools they were given, rather than which provider scored best. Every number below comes from scripts/analyze_search_behavior.py, whose classification rules are listed in methodology. The analysis covers 974 search queries from the GAP battery (both harnesses) and 655 from the bakeoff (Claude Code).

Headline

  • Agents write keyword queries whatever the engine asks for. Of 974 GAP search queries, 7.3% read as natural language. In the bakeoff the figure is 6.0%. Exa's tool asks for a description of the ideal page rather than keywords; its queries were natural language in 3% (Claude Code) and 2% (codex) of GAP calls. Because the automatic rule can miss descriptive phrasing, all 111 Exa queries from both batteries were also read by hand (appendix): none is phrased as a description of a page. The closest are a headline-like phrase used three times ("Python 4.0 release date announced by Python Steering Council") and a page title.
  • The two harnesses express constraints differently. On the six provider arms, codex put site: into 31% to 67% of its queries, even where a structured domain filter existed. Claude Code wrote site: only on Firecrawl (6%) and otherwise used structured domain filters where a tool offered one (Perplexity search_domain_filter, Tavily include_domains, WebSearch allowed_domains).
  • Agents often fetch remembered URLs instead of searching. In the bakeoff, Claude Code fetched pages without any search in 36 of 55 Exa cells, 35 of 55 Firecrawl cells, 34 of 54 Tavily cells and 30 of 55 Parallel cells. Those provider arms therefore measured page fetching as much as search. On post-cutoff GAP tasks the same habit failed: Claude Code's Exa cells that only fetched passed 1/3, against 14/18 for its Exa cells that searched.
  • Retrieving the primary source mattered most on Claude Code. Pooled over provider and native arms, Claude Code passed 94/111 cells (85%) when the task's primary source appeared in search results and 20/33 (61%) when it did not. Codex surfaced the primary source in most searched cells, so the comparison is thin there (95/116 against 8/10). Two arms passed without the primary page: Parallel (6/7) and Firecrawl (9/9) on Claude Code, which indicates their results carried the facts from other pages.

These are observations about agent behavior with each vendor's MCP server at its pinned version and default settings. They are not measurements of the vendors' APIs.

What each search tool asks for

Descriptions are quoted from the pinned MCP server packages and, for Parallel's hosted server, from its published documentation.

ArmSearch interface the agent sawWhat the tool tells the agent about queries
Exaweb_search_exa(query, numResults)"describe the ideal page, not keywords"; query is a "semantically rich description of the ideal page, not just keywords"
Parallelweb_search(objective, search_queries, …)a natural-language objective plus "concise, related keyword queries of 3–6 words each"
Perplexityperplexity_search(query, …) plus ask, research, reason"Supports recency filters, domain restrictions, and a lower-latency fast search mode"
Tavilytavily_search(query, search_depth, …)query is described only as "Search query"
Firecrawlfirecrawl_search(query, …)"Query operators, domain filters, categories … are described on their parameters"
Bravebrave_web_search(query, …)lists when to use the tool; no guidance on query phrasing
native (Claude Code)WebSearch(query, allowed_domains, …)harness built-in
native (codex)built-in web searchharness built-in; no parameters exposed to the transcript

Query style (GAP battery)

Parallel's search_queries are keyword queries by design; its natural-language part is the separate objective field, which is counted as a structured parameter and not as a query. Perplexity's user messages count as queries alongside its query inputs; system and assistant messages are excluded. Codex native search reports its query lists in the transcript; page views are counted as fetches.

HarnessArmQueriesNatural languageContains a yearQuoted phrasesite:Median wordsStructured parameters passed (calls)
claude-codenative1017%43%39%0%11allowed_domains ×11
claude-codeBrave7217%44%18%0%8count ×23, extra_snippets ×21, maximum_number_of_tokens ×18, freshness ×1, maximum_number_of_urls ×1
claude-codeTavily3813%45%18%0%9max_results ×38, include_domains ×19, search_depth ×17, include_raw_content ×1, start_date ×1
claude-codeExa383%58%0%0%10numResults ×38
claude-codeParallel945%24%2%0%6model_name ×27, objective ×27, session_id ×27
claude-codeFirecrawl340%59%9%6%7limit ×34, sources ×34, includeDomains ×3
claude-codePerplexity7020%36%10%0%9max_results ×60, search_domain_filter ×42, max_tokens_per_page ×5, messages ×5, search_context_size ×3, search_recency_filter ×2, search_type ×1
codexnative480%50%33%10%8none
codexBrave1142%52%26%46%8maximum_number_of_tokens ×30, count ×15, maximum_number_of_urls ×13, extra_snippets ×7, url ×3
codexTavily1306%50%59%67%8max_results ×123, search_depth ×26, include_raw_content ×25, end_date ×3, include_domains ×3
codexExa522%60%35%58%10numResults ×10
codexParallel626%39%27%48%7objective ×28, session_id ×10, max_results ×3, max_chars_per_result ×1
codexFirecrawl422%74%26%31%8limit ×33, sources ×10
codexPerplexity7914%71%33%42%10max_results ×20, messages ×18, max_tokens_per_page ×16

Query style (web search bakeoff, Claude Code)

ArmCellsQueriesNatural languageContains a yearQuoted phrasesite:Median wordsStructured parameters passed (calls)
native541372%34%9%4%7allowed_domains ×36
Brave541994%7%11%6%6count ×85, maximum_number_of_tokens ×80, freshness ×17, goggles ×16, extra_snippets ×11, maximum_number_of_urls ×9, maximum_number_of_tokens_per_url ×3
Tavily54250%32%12%0%10max_results ×18, include_domains ×10, search_depth ×4, time_range ×2, start_date ×1
Exa55210%19%5%0%10numResults ×21
Parallel55770%20%0%1%5model_name ×22, objective ×22, session_id ×22
Firecrawl55200%25%15%5%7limit ×20, sources ×20
Perplexity5517616%13%10%0%8max_results ×154, search_domain_filter ×115, max_tokens_per_page ×21, search_type ×21, messages ×13, search_context_size ×13, search_recency_filter ×5

Searching versus fetching from memory (web search bakeoff, Claude Code)

A cell is "fetched without searching" when it made at least one fetch call and no search call. Brave's and Perplexity's servers expose no page-fetch tool, so their agents had to search or answer from memory. Every arm has a handful of cells with no provider call: answers given from memory and cells that ended before any tool call.

ArmCellsFetch tool in the armFetched without searchingNo provider callCalls refused by the harness
native54WebFetch (mostly refused; see below)126110
Brave54no070
Tavily54yes (tavily_extract)3460
Exa55yes (web_fetch_exa)3660
Parallel55yes (web_fetch)3060
Firecrawl55yes (firecrawl_scrape)3560
Perplexity55no070

On the bakeoff's stable-knowledge tasks, fetching a remembered URL is efficient: those tasks passed in every arm, including no-search. It means the bakeoff's provider arms with a fetch tool mostly measured that fetch path. The GAP battery, whose answers post-date the models, shows the cost of the same habit on new information.

Harness refusals: the Claude Code native arm

Claude Code's native arm exposed WebSearch and WebFetch through --tools but pre-approved only WebSearch through --allowedTools. Headless Claude Code refuses a tool that is exposed but not pre-approved ("Claude requested permissions to use WebFetch, but you haven't granted it yet"). In the GAP battery every native WebFetch call was refused (50 calls in 21 of 21 cells); in the bakeoff, 110 calls in 40 of 54 cells. No provider arm and no codex arm had a refused call. The Claude Code native results in both batteries therefore understate that harness's web tools. The arm configuration is fixed and covered by a regression test. The fetch counts in this report include refused attempts.

Primary-source retrieval (GAP battery)

Each GAP brief has one primary-source URL (the catalog oracle). A cell surfaced it when that URL appeared among the URLs of returned search results, excluding search inputs. Unknown coverage is excluded from both pass-rate groups. Recomputing against returned evidence leaves the historical source counts unchanged; no searched cell has unknown coverage. Facts can also come from secondary pages, so surfacing is a proxy for retrieval, not a requirement for passing.

HarnessArmCells that searchedUnknown coveragePrimary source surfacedPass when surfacedPass when not surfaced
claude-codenative21018/2113/180/3
claude-codeBrave21018/2115/182/3
claude-codeTavily21016/2113/162/5
claude-codeExa18014/1813/141/4
claude-codeParallel21014/2114/146/7
claude-codeFirecrawl21012/2110/129/9
claude-codePerplexity21019/2116/190/2
codexnative18017/1813/171/1
codexBrave18016/1815/162/2
codexTavily18016/1814/162/2
codexExa18017/1814/170/1
codexParallel18017/1812/171/1
codexFirecrawl18015/1814/152/3
codexPerplexity18018/1813/18—

Limits

  • Cell counts per arm are small (18 to 21 in GAP, 54 to 55 in the bakeoff). Per-arm percentages move by about five points per cell.
  • The natural-language rule is a heuristic, stated in methodology; readers can rerun the script with their own rule.
  • Each vendor was used through its MCP server at the pinned version and default settings. Other interfaces of the same vendor (direct API, other tools, other modes) were not tested.
  • The two harnesses ran different task subsets in GAP (7 tasks for Claude Code, 6 for codex), so cross-harness differences mix agent behavior with task mix.

Appendix: every Exa query

Every query sent to web_search_exa in both batteries, in transcript order, so readers can judge the phrasing directly. Other arms' queries are summarised above but not reproduced.

BatteryHarnessQuery
GAPclaude-codeCopilot Billing Preview app deprecated usage report
GAPclaude-codeGitHub Copilot premium request budgets per user cap and usage report CSV export docs
GAPclaude-codeGitHub docs AI usage page billing settings group filter export AI credits Copilot limitations
GAPclaude-codeGitHub Copilot Billing Preview app deprecated usage-based billing budgets
GAPclaude-codeGitHub docs Copilot premium requests usage report export CSV billing 2026
GAPclaude-codeGitHub docs AI usage page billing settings group filter export AI credits usage report limitations
GAPclaude-codeGitHub Copilot Billing Preview app deprecated usage report
GAPclaude-codeGitHub docs set budget for individual Copilot premium requests user-level budget
GAPclaude-codeGitHub docs AI usage page billing settings group filter export AI credits Copilot usage report limitations
GAPclaude-codegithub.blog changelog 2026-08-04 Retiring the Copilot Billing Preview app
GAPclaude-codeGitHub Classroom deprecation sunset announcement 2026
GAPclaude-codeClassroom 50 CS50 foundation GitHub Classroom alternative import autograding LTI
GAPclaude-codeCodio GitHub Classroom migration partner
GAPclaude-codeGitHub Classroom sunset deprecation announcement 2026
GAPclaude-codeClassroom 50 CS50 open source GitHub Classroom replacement import migration
GAPclaude-codeGitHub Classroom deprecation sunset announcement 2026
GAPclaude-codeGitHub Models deprecation retirement changelog 2026 BYOK inference
GAPclaude-codeGitHub Models deprecation retirement changelog 2026 BYOK paid usage
GAPclaude-codeAzure AI Inference beta SDK deprecated retirement August 26 2026 migrate to OpenAI v1 API
GAPclaude-codeGitHub Models deprecation retirement changelog 2026 BYOK
GAPclaude-codeGitHub secret scanning npm token revocation credential types 2026 changelog
GAPclaude-codeGitHub changelog npm granular access tokens publish and manage package access 2026
GAPclaude-codeGitHub secret scanning npm token revocation credential types in scope 2026
GAPclaude-codeGitHub changelog npm granular access tokens bypass 2FA publish manage package access 2026
GAPclaude-codeGitHub changelog npm granular access tokens bypass 2FA publish package access management 2026
GAPclaude-codeGitHub secret scanning npm token revocation credential types 2026 changelog
GAPclaude-codeGitHub changelog npm staged publishing generally available npm stage publish 2026
GAPclaude-codePython 3.10 security release tarfile extraction filter bypass CVE-2025-4517
GAPclaude-codeCPython tarfile data filter vulnerability 2026 CVE security-announce
GAPclaude-codePython 3.10 security release tarfile extraction filter bypass fix
GAPclaude-codeCVE-2026-4360 tarfile extract filter hardlinks Python fixed version
GAPclaude-codeGitHub docs converting a user into an organization warning cannot sign in personal account
GAPclaude-codeGitHub Enterprise Server docs promoting or demoting site administrator converting user to organization
GAPclaude-codepython.org release metadata API authentication vulnerability security advisory 2026
GAPclaude-codepython.org release metadata vulnerability unauthenticated API modify release files 2026
GAPclaude-codepython.org release API authentication bypass DEVCORE Splitline CVE advisory exploitation
GAPclaude-codepython.org release metadata vulnerability unauthenticated modify release files advisory 2026
GAPclaude-codePython Software Foundation security incident python.org downloads release API authentication
GAPcodexGitHub Classroom sunset July 2026 September 2026 partner repositories
GAPcodexGitHub Classroom retirement September 1 2026 repositories partner
GAPcodexsite:github.com/orgs/community/discussions/205975 "September 4"
GAPcodexGitHub Classroom sunset September 2026 repositories partner retirement
GAPcodexsite:github.com/orgs/community/discussions GitHub Classroom "September 4" "2026"
GAPcodexGitHub Models inference BYOK retirement July August 2026 paid existing usage
GAPcodexGitHub Models retiring July August 2026 paid inference BYOK July 10
GAPcodexGitHub Models retirement July August 2026 paid inference BYOK existing usage
GAPcodexsite:docs.github.com converting user into organization January 2026 transfer repositories personal account enterprise server
GAPcodexsite:github.blog/changelog 2026 January converting user organization disabled January 2026
GAPcodexsite:docs.github.com converting a user into an organization January 2026 transfer repositories
GAPcodexsite:github.blog/changelog 2026 January converting personal accounts organizations January 20
GAPcodexsite:docs.github.com converting user into organization January 2026 transfer repositories personal account GHES
GAPcodexsite:github.blog/changelog 2026 January 20 convert account organization transfer personal account
GAPcodexGitHub Copilot Billing Preview app retired individual budgets usage export billing reports
GAPcodexsite:docs.github.com Copilot AI usage page reports limitations billing API user level budgets export usage credits
GAPcodexsite:docs.github.com "AI usage" "report" "does not" billing usage reports
GAPcodexsite:docs.github.com "Setting up budgets" "user-level" "Budgets and alerts"
GAPcodex"Copilot Billing Preview" budgets export usage
GAPcodexsite:docs.github.com AI usage user-level budgets billing reports Copilot AI credits
GAPcodexsite:docs.github.com "AI usage" "does not" reports billing
GAPcodexsite:docs.github.com "Viewing your usage" "AI" "report" billing
GAPcodexsite:docs.github.com "Automating usage reporting" "AI"
GAPcodexsite:docs.github.com Copilot usage metrics billing data coverage IDE chat github CLI
GAPcodexsite:docs.github.com "Setting" "user-level budgets" "Budgets"
GAPcodexCopilot Billing Preview app individual budgets export raw usage finance billing dashboard
GAPcodexsite.github.blog/changelog 2026-08-04 retiring copilot billing preview
GAPcodexsite:docs.github.com AI usage page billing reports user-level budgets coverage limitations Copilot
GAPcodexsite:docs.github.com "AI usage" "does not" reports billing
GAPcodexsite:docs.github.com "AI usage" "coverage"
GAPcodexsite:docs.github.com "Viewing your usage of metered products" "AI usage"
GAPcodexsite:github.blog/changelog "user-level budgets" 2026
GAPcodexsite:docs.github.com copilot usage metrics "billing" "coverage"
GAPcodexsite:github.blog/changelog billing CSV usage reports API 2026
GAPcodexnpm token security granular tokens package access management July 2026 August 2026 GitHub credentials
GAPcodexsite:github.blog/changelog/2026 staged publishing npm May 22 trusted publishing tokens approve
GAPcodexsite:github.blog/changelog/ "2026-05-22" "npm"
GAPcodexnpm token security granular tokens package access github August 2026 publishing tokens
GAPcodexsite:github.blog/changelog npm December 9 2025 classic tokens revoked staged publishing July 2026
GAPcodexnpm token security August 2026 package access granular tokens GitHub credentials
GAPcodexsite:github.blog/changelog npm staged publishing 2026 July 2FA
GAPcodexpython.org release metadata vulnerability unsigned downloads API incident 2026 2025
GAPcodexsite.python.org Sigstore verify Python releases identity PEP 761
GAPcodexsite.blog.python.org/2026/06/mitigated-api-bypass "Timeline" "30" "June"
GAPcodex"python/pythondotorg" "3014" "merged" "June"
GAPcodexpython.org release metadata vulnerability authentication 2026 2025 security incident
GAPcodexsite:blog.python.org/2026/06/mitigated-api-bypass "Remediations" "Timeline"
GAPcodexsite:github.com/python/pythondotorg/issues/3010 OR site:github.com/python/pythondotorg/issues/3011
GAPcodexpython.org release metadata security incident authentication bypass before June 30 2026 Trail of Bits
GAPcodexpython.org release metadata vulnerability authentication 2026 2025 API incident
GAPcodexsite.python.org security incident release metadata 2026 2025 downloads authentication
GAPcodexPython Windows code signing certificates revoked August 2025 update investigation June 2026
bakeoffclaude-codePython 4.0 release date announced by Python Steering Council
bakeoffclaude-codewebpack md4 OpenSSL 3 Node 17 hash wasm md4 fix output.hashFunction xxhash64 sokra comment
bakeoffclaude-codegrype offline air-gapped vulnerability database import GRYPE_DB_AUTO_UPDATE false
bakeoffclaude-codecargo audit --no-fetch --db local advisory database path offline
bakeoffclaude-codeDependency-Track air-gapped offline vulnerability mirror configuration NVD OSV feeds URL
bakeoffclaude-codePython 4.0 release date announced by Python Steering Council
bakeoffclaude-codePython 4.0 release date announced by Python Steering Council
bakeoffclaude-codeNode.js documentation --openssl-legacy-provider option "Enable OpenSSL 3.0 legacy provider"
bakeoffclaude-codeaws.amazon.com data transfer out to internet US East (N. Virginia) first 10 TB per GB price same availability zone free
bakeoffclaude-codegrype air-gapped offline database import GRYPE_DB_AUTO_UPDATE false documentation
bakeoffclaude-codeDependency-Track air-gapped offline NVD mirror configuration documentation
bakeoffclaude-codeSafety CLI 3 requires account login API key vulnerability database commercial license
bakeoffclaude-codeAWS EC2 on-demand pricing data transfer out to internet US East (N. Virginia) first 10 TB per month per GB; data transfer same availability zone free
bakeoffclaude-codeStripe annual letter 2025 total payment volume private company revenue
bakeoffclaude-codeStripe revenue estimate UK Companies House filing Stripe Payments Europe accounts net revenue
bakeoffclaude-codeStripe annual letter 2025 total payment volume private company revenue
bakeoffclaude-codeStripe revenue estimate net revenue report private company does not disclose financials
bakeoffclaude-codeStripe Irish entity accounts filed Companies Registration Office revenue $5.1bn pre-tax profit 2024
bakeoffclaude-codeStripe annual letter 2025 total payment volume private company revenue estimate
bakeoffclaude-codeGrype offline air-gapped vulnerability database GRYPE_DB_AUTO_UPDATE false grype db import
bakeoffclaude-codeDependency-Track air-gapped offline mirror NVD OSV internal mirror documentation
Methodology

Agent search behavior: methodology

The analysis reads the stored transcripts of the web search bakeoff (Claude Code, eight arms, 18 tasks, three repetitions) and the search gap bench (Claude Code and codex, brief tasks, three repetitions per admitted task). It runs no agent and changes no grade. Pass and fail come from each cell's recorded evaluation.

scripts/analyze_search_behavior.py extracts every completed provider tool call from a transcript:

  • codex MCP tool calls, and codex built-in web_search items;
  • Claude Code tool_use and tool_result pairs, including vendor MCP tools, WebSearch and WebFetch.

Classification rules

  • Search call: a call to a search tool (web_search_exa, brave_web_search, tavily_search, Parallel web_search, firecrawl_search, the Perplexity search tools, Claude Code WebSearch, and codex built-in search actions).
  • Query: each entry of a multi-query call (Parallel search_queries, codex built-in query lists) counts as one query. Each user message in Perplexity's conversational inputs also counts as one query; system and assistant messages are excluded. Parallel's objective is counted as a structured parameter when search_queries is present, not as another query.
  • Fetch call: a page-retrieval tool call (web_fetch_exa, tavily_extract, firecrawl_scrape, Parallel web_fetch, Claude Code WebFetch, and codex built-in open_page, find_in_page and other page views).
  • Natural-language query: ends with "?", starts with an interrogative or auxiliary verb, or contains at least two function words from a fixed list (the, a, an, is, are, was, were, does, do, did, how, what, which, why, when, where, who, that, this, of, to, for, with, will, should, can, about, on, from, by, as, be, has, have).
  • Contains a year: a token from 2000 to 2099. Quoted phrase: contains a double quote. site:: contains site:.
  • Fetched without searching: the cell made at least one fetch call and no search call.
  • Primary source surfaced (GAP only): the task's catalog oracle source_url, normalised (scheme, query, fragment and trailing slash removed), appears among the URLs in any returned search result of the cell, never in the search input. It is measured only over cells that searched. A cell is unknown when any search lacks exposed results and no observed result contains the primary source. Unknown cells are excluded from both surfaced and not-surfaced pass-rate denominators. Recomputing against returned evidence leaves the historical source counts unchanged; no searched cell has unknown coverage.

The rules are deliberately simple so they can be audited. The natural-language rule undercounts descriptive keyword-free phrases and overcounts keyword strings that happen to contain function words. Readers can substitute their own rule in the script and rerun it.

Scope

Pooled figures cover the provider and native arms. Reference arms (no-search, floor and ceiling) make no search calls and are excluded from query statistics. GAP cell counts differ by harness because calibration admitted seven brief tasks for Claude Code and six for codex.

See the report, summary data and reproduction. Calibration does not apply: calibration record.

Files: reports/2026-10-05-agent-search-behavior/

Reproduce, correct, contribute

Install Searchlight, run sew doctor, then follow each run's reproduce.md. Raw transcripts and provider responses are not distributed, so exact replay of the published aggregates is not possible; a new run measures the same method. If you maintain a tested service and a configuration here misrepresents it, open an issue with the configuration you recommend: corrections are rerun and published beside the original, never silently replaced.

pipx install 'git+https://github.com/laceyenterprises/searchlight'
sew doctor
python3 scripts/check_reports.py
python3 scripts/build_site.py --check