Methodology
How Searchlight tests search
A controlled experiment: the same agent, model and instructions in every test run, with only the search provider changing. This page describes the apparatus in full.
Terms used in the reports
From task to report
The runner gives the agent a fresh scratch workspace, starts it with exactly the tools its search setup allows, records every tool call and the final answer, and grades the answer without telling the grader which setup produced it.
flowchart LR
subgraph catalog[Task catalogs]
T1[production tasks<br/>web search bakeoff]
T2[gap briefs<br/>post-cutoff events]
end
subgraph cell[One test run = task x search setup x repetition]
W[fresh scratch<br/>workspace]
H[agent<br/>Claude Code or Codex]
M[metering proxy]
V[(search provider MCP server<br/>pinned version)]
end
subgraph grade[Grading]
D[deterministic<br/>validators]
J[blinded judges<br/>vs captured sources]
end
T1 & T2 --> W --> H
H <-->|tool calls| M <--> V
H --> B[(evidence bundle<br/>transcript, final answer,<br/>setup audit, metrics)]
B --> D & J --> R[reports/*/summary.json] --> S[this site]
What each setup can touch
Search setups differ only in what the agent can use to search. Claude Code is started with
--strict-mcp-config, so the only MCP server it can reach is the provider's, and --tools limits its
built-ins. Codex runs read-only with its shell, browser, app and plugin features disabled. After each run an audit compares
every observed tool call with the setup's contract; a call outside it marks the run contaminated.
flowchart TB A[search setup contract] --> CC[Claude Code spawn<br/>--strict-mcp-config<br/>--tools: the setup's built-ins<br/>--allowedTools: the same set] A --> CX[codex spawn<br/>exec --sandbox read-only<br/>--disable shell_tool, browser_use,<br/>apps, plugins, ...] CC --> P1[search provider:<br/>one provider MCP server<br/>+ ToolSearch + Read] CC --> N1[built-in search:<br/>WebSearch + WebFetch + Read] CC --> F1[no-search / floor / ceiling:<br/>no tools] CX --> P2[search provider:<br/>one provider MCP server] CX --> N2[built-in search:<br/>codex --search] CX --> F2[no-search / floor / ceiling:<br/>no tools] P1 & N1 & F1 & P2 & N2 & F2 --> AU[setup audit:<br/>every observed tool call<br/>checked against the contract]
Claude Code invocation
claude --print --output-format stream-json --verbose --no-session-persistence \ --model claude-opus-5-5 --mcp-config <run>/claude-mcp.json --strict-mcp-config \ --tools ToolSearch,Read --allowedTools 'mcp__<provider>__*' # one search provider --tools WebSearch,WebFetch,Read --allowedTools WebSearch,WebFetch # built-in search (fixed) --tools '' --allowedTools '' # no search, floor, ceiling
Codex invocation
codex [--search] exec --json --ephemeral --skip-git-repo-check --sandbox read-only \ --model gpt-6.1-sol --disable shell_tool --disable unified_exec \ --disable browser_use --disable browser_use_external --disable computer_use \ --disable in_app_browser --disable apps --disable plugins --disable remote_plugin ... # --search only for built-in search; a provider setup adds one MCP server via the run config
Limits per test run
| Limit | Search gap bench | Web search bakeoff |
|---|---|---|
| Wall clock | 1,200 s | 300 to 1,200 s (set per task) |
| Total tokens | 1,000,000 | 120,000 to 350,000 (set per task) |
| Search provider calls | 60 | 15 to 50 (set per task) |
| Agent start-up timeout | 120 s | 120 s |
| Repetitions | 3 (battery), 5 (calibration) | 3 |
Which tasks count
A search gap task counts only if it is genuinely beyond the model: with no search the agent must fail, and with the answer excerpt handed over it must succeed. Calibration is repeated for each agent and model.
flowchart LR
T[candidate brief<br/>event after 2026-01-01] --> FL[floor: no search<br/>5 runs]
T --> CE[ceiling: answer excerpt<br/>in the prompt, 5 runs]
FL --> Q{floor passes ≤ 1<br/>and ceiling passes ≥ 4?}
CE --> Q
Q -- yes --> AD[admitted for this<br/>agent and model]
Q -- no --> RJ[rejected]
AD --> BAT[battery: every search setup x 3]
BAT --> GC[gap closure =<br/>setup - floor / ceiling - floor]
gap closure = (setup pass rate − floor pass rate) / (ceiling pass rate − floor pass rate)
floor search setup ceiling
0% ●─────────────────●──────────────────● 95%
└──── closed ────┘└──── remaining ────┘How answers are graded
sequenceDiagram participant A as Agent answer participant V as Schema validator participant S as Source capture participant J1 as Judge 1 (Claude Code) participant J2 as Judge 2 (Codex, gap bench) A->>V: JSON deliverable (brief, claims, citations) V-->>A: reject malformed answers A->>S: cited URLs S->>J1: readable source text + rubric (blinded to the search setup) S->>J2: same inputs J1-->>A: key-fact recall, unsupported claims, decision J2-->>A: independent labels, agreement recorded Note over A,J2: pass = right decision AND recall >= 0.7 AND unsupported <= 0.25
Search gap briefs pass when the decision is right, at least 70% of key facts are covered and at most 25% of claims are unsupported. Code fixes pass when hidden tests pass in an isolated workspace. Bakeoff tasks use automatic checks where a task has one, and otherwise a single blinded grader scores the task rubric.
Search provider configurations
Pinned in config/provider-arm-tools.yaml. Every tool a provider's server advertised was available to the agent; no provider tool was hidden.
| Search provider | Connection | Version / endpoint | Tools exposed to the agent |
|---|---|---|---|
| Brave | stdio MCP server (npx) | @brave/brave-search-mcp-server@2.1.4 | 8: brave_image_search, brave_llm_context, brave_local_search, brave_news_search, brave_place_search, brave_summarizer, brave_video_search, brave_web_search |
| Exa | stdio MCP server (npx) | exa-mcp-server@3.4.1 | 2: web_fetch_exa, web_search_exa |
| Firecrawl | stdio MCP server (npx) | firecrawl-mcp@3.26.0 | 29: firecrawl_agent, firecrawl_agent_status, firecrawl_check_crawl_status, firecrawl_crawl, firecrawl_credit_usage, firecrawl_developer_search, firecrawl_extract, firecrawl_feedback, firecrawl_find_tools, firecrawl_interact, firecrawl_interact_stop, firecrawl_map, firecrawl_monitor_check, firecrawl_monitor_checks, firecrawl_monitor_create, firecrawl_monitor_delete, firecrawl_monitor_get, firecrawl_monitor_list, firecrawl_monitor_run, firecrawl_monitor_update, firecrawl_parse, firecrawl_research_inspect_paper, firecrawl_research_read_paper, firecrawl_research_related_papers, firecrawl_research_search_github, firecrawl_research_search_papers, firecrawl_scrape, firecrawl_search, firecrawl_search_feedback |
| Parallel | hosted MCP via mcp-remote | mcp-remote@0.14.3 → https://search.parallel.ai/mcp-oauth | 2: web_fetch, web_search |
| Perplexity | stdio MCP server (npx) | @perplexity-ai/mcp-server@1.3.0 | 4: perplexity_ask, perplexity_reason, perplexity_research, perplexity_search |
| Tavily | stdio MCP server (npx) | tavily-mcp@0.2.22 | 5: tavily_crawl, tavily_extract, tavily_map, tavily_research, tavily_search |
| Built-in (Claude Code) | agent built-in | Claude Code 2.1.282 | WebSearch, WebFetch (refused before 2026-10-06), Read |
| Built-in (Codex) | agent built-in | codex --search | built-in web search and page views |
The search API head-to-head
One study tests the search APIs themselves rather than an agent using them. It sends the same pre-registered research questions to Exa, Tavily, Parallel and Firecrawl, each with the parameters its own documentation recommends, on the same day and with the same number of results. The results are pooled and shuffled, and blind model judges verify every counted claim on its source page. A provider is credited only when its own result page states the claim, and a claim counts as new only if the knowledge base the questions came from did not already have it. Unlike the benchmarks above, each question ran once, so the study reports counts without intervals.
flowchart LR
Q[pre-registered question sets<br/>from one knowledge base] --> E[Exa] & T[Tavily] & P[Parallel] & F[Firecrawl]
E & T & P & F --> PO[(results pooled,<br/>deduplicated, shuffled)]
PO --> J[blind model judges<br/>verify each claim on its page]
J --> K{already in the<br/>knowledge base?}
K --> A[attribution rejoined:<br/>credit only where the provider's<br/>own page states the claim]
A --> R[reports/*/summary.json]
Its sets, endpoints, amendments and costs are in the study's methodology, and its full pre-registered record is in reports/2026-09-26-search-api-head-to-head/.
Reproduce, correct, contribute
Install Searchlight, run sew doctor, then follow each study's reproduce.md. Raw transcripts and
provider responses are not distributed, so exact replay of the published numbers is not possible; a new run measures the
same method. The one exception is the search API head-to-head's Monitors probe, whose raw responses are published. If you run a tested service and a configuration here misrepresents it, open an issue with the configuration you
recommend: corrections are rerun and published beside the original, never silently replaced.
pipx install 'git+https://github.com/laceyenterprises/searchlight'
sew doctor
python3 scripts/check_reports.py
python3 scripts/build_site.py --check