SEARCHLIGHTGitHub

Methodology

How Searchlight tests search

A controlled experiment: the same agent, model and instructions in every test run, with only the search provider changing. This page describes the apparatus in full.

Terms used in the reports

Search setup (report term: arm)One way of giving the agent search: one provider's MCP server, the agent's built-in search, or no search at all.
Test run (report term: cell)One task, one search setup, one repetition, graded once.
Agent (report term: harness)The AI coding agent program: Claude Code or Codex.
Floor and ceilingThe same task with no search (floor) and with the answer excerpt handed over (ceiling).
Gap closureHow much of the distance from floor to ceiling a search setup recovers.
95% rangeThe Wilson interval around a pass rate. Overlapping ranges mean the data cannot separate two setups.

From task to report

The runner gives the agent a fresh scratch workspace, starts it with exactly the tools its search setup allows, records every tool call and the final answer, and grades the answer without telling the grader which setup produced it.

flowchart LR
  subgraph catalog[Task catalogs]
    T1[production tasks<br/>web search bakeoff]
    T2[gap briefs<br/>post-cutoff events]
  end
  subgraph cell[One test run = task x search setup x repetition]
    W[fresh scratch<br/>workspace]
    H[agent<br/>Claude Code or Codex]
    M[metering proxy]
    V[(search provider MCP server<br/>pinned version)]
  end
  subgraph grade[Grading]
    D[deterministic<br/>validators]
    J[blinded judges<br/>vs captured sources]
  end
  T1 & T2 --> W --> H
  H <-->|tool calls| M <--> V
  H --> B[(evidence bundle<br/>transcript, final answer,<br/>setup audit, metrics)]
  B --> D & J --> R[reports/*/summary.json] --> S[this site]

What each setup can touch

Search setups differ only in what the agent can use to search. Claude Code is started with --strict-mcp-config, so the only MCP server it can reach is the provider's, and --tools limits its built-ins. Codex runs read-only with its shell, browser, app and plugin features disabled. After each run an audit compares every observed tool call with the setup's contract; a call outside it marks the run contaminated.

flowchart TB
  A[search setup contract] --> CC[Claude Code spawn<br/>--strict-mcp-config<br/>--tools: the setup's built-ins<br/>--allowedTools: the same set]
  A --> CX[codex spawn<br/>exec --sandbox read-only<br/>--disable shell_tool, browser_use,<br/>apps, plugins, ...]
  CC --> P1[search provider:<br/>one provider MCP server<br/>+ ToolSearch + Read]
  CC --> N1[built-in search:<br/>WebSearch + WebFetch + Read]
  CC --> F1[no-search / floor / ceiling:<br/>no tools]
  CX --> P2[search provider:<br/>one provider MCP server]
  CX --> N2[built-in search:<br/>codex --search]
  CX --> F2[no-search / floor / ceiling:<br/>no tools]
  P1 & N1 & F1 & P2 & N2 & F2 --> AU[setup audit:<br/>every observed tool call<br/>checked against the contract]

Claude Code invocation

claude --print --output-format stream-json --verbose --no-session-persistence \
  --model claude-opus-5-5 --mcp-config <run>/claude-mcp.json --strict-mcp-config \
  --tools ToolSearch,Read --allowedTools 'mcp__<provider>__*'     # one search provider
  --tools WebSearch,WebFetch,Read --allowedTools WebSearch,WebFetch   # built-in search (fixed)
  --tools '' --allowedTools ''                                       # no search, floor, ceiling

Codex invocation

codex [--search] exec --json --ephemeral --skip-git-repo-check --sandbox read-only \
  --model gpt-6.1-sol --disable shell_tool --disable unified_exec \
  --disable browser_use --disable browser_use_external --disable computer_use \
  --disable in_app_browser --disable apps --disable plugins --disable remote_plugin ...
# --search only for built-in search; a provider setup adds one MCP server via the run config

Limits per test run

LimitSearch gap benchWeb search bakeoff
Wall clock1,200 s300 to 1,200 s (set per task)
Total tokens1,000,000120,000 to 350,000 (set per task)
Search provider calls6015 to 50 (set per task)
Agent start-up timeout120 s120 s
Repetitions3 (battery), 5 (calibration)3

Which tasks count

A search gap task counts only if it is genuinely beyond the model: with no search the agent must fail, and with the answer excerpt handed over it must succeed. Calibration is repeated for each agent and model.

flowchart LR
  T[candidate brief<br/>event after 2026-01-01] --> FL[floor: no search<br/>5 runs]
  T --> CE[ceiling: answer excerpt<br/>in the prompt, 5 runs]
  FL --> Q{floor passes ≤ 1<br/>and ceiling passes ≥ 4?}
  CE --> Q
  Q -- yes --> AD[admitted for this<br/>agent and model]
  Q -- no --> RJ[rejected]
  AD --> BAT[battery: every search setup x 3]
  BAT --> GC[gap closure =<br/>setup - floor / ceiling - floor]
gap closure  =  (setup pass rate − floor pass rate) / (ceiling pass rate − floor pass rate)

   floor          search setup           ceiling
   0% ●─────────────────●──────────────────● 95%
       └──── closed ────┘└──── remaining ────┘

How answers are graded

sequenceDiagram
  participant A as Agent answer
  participant V as Schema validator
  participant S as Source capture
  participant J1 as Judge 1 (Claude Code)
  participant J2 as Judge 2 (Codex, gap bench)
  A->>V: JSON deliverable (brief, claims, citations)
  V-->>A: reject malformed answers
  A->>S: cited URLs
  S->>J1: readable source text + rubric (blinded to the search setup)
  S->>J2: same inputs
  J1-->>A: key-fact recall, unsupported claims, decision
  J2-->>A: independent labels, agreement recorded
  Note over A,J2: pass = right decision AND recall >= 0.7 AND unsupported <= 0.25

Search gap briefs pass when the decision is right, at least 70% of key facts are covered and at most 25% of claims are unsupported. Code fixes pass when hidden tests pass in an isolated workspace. Bakeoff tasks use automatic checks where a task has one, and otherwise a single blinded grader scores the task rubric.

Search provider configurations

Pinned in config/provider-arm-tools.yaml. Every tool a provider's server advertised was available to the agent; no provider tool was hidden.

Search providerConnectionVersion / endpointTools exposed to the agent
Bravestdio MCP server (npx)@brave/brave-search-mcp-server@2.1.48: brave_image_search, brave_llm_context, brave_local_search, brave_news_search, brave_place_search, brave_summarizer, brave_video_search, brave_web_search
Exastdio MCP server (npx)exa-mcp-server@3.4.12: web_fetch_exa, web_search_exa
Firecrawlstdio MCP server (npx)firecrawl-mcp@3.26.029: firecrawl_agent, firecrawl_agent_status, firecrawl_check_crawl_status, firecrawl_crawl, firecrawl_credit_usage, firecrawl_developer_search, firecrawl_extract, firecrawl_feedback, firecrawl_find_tools, firecrawl_interact, firecrawl_interact_stop, firecrawl_map, firecrawl_monitor_check, firecrawl_monitor_checks, firecrawl_monitor_create, firecrawl_monitor_delete, firecrawl_monitor_get, firecrawl_monitor_list, firecrawl_monitor_run, firecrawl_monitor_update, firecrawl_parse, firecrawl_research_inspect_paper, firecrawl_research_read_paper, firecrawl_research_related_papers, firecrawl_research_search_github, firecrawl_research_search_papers, firecrawl_scrape, firecrawl_search, firecrawl_search_feedback
Parallelhosted MCP via mcp-remotemcp-remote@0.14.3 → https://search.parallel.ai/mcp-oauth2: web_fetch, web_search
Perplexitystdio MCP server (npx)@perplexity-ai/mcp-server@1.3.04: perplexity_ask, perplexity_reason, perplexity_research, perplexity_search
Tavilystdio MCP server (npx)tavily-mcp@0.2.225: tavily_crawl, tavily_extract, tavily_map, tavily_research, tavily_search
Built-in (Claude Code)agent built-inClaude Code 2.1.282WebSearch, WebFetch (refused before 2026-10-06), Read
Built-in (Codex)agent built-incodex --searchbuilt-in web search and page views

The search API head-to-head

One study tests the search APIs themselves rather than an agent using them. It sends the same pre-registered research questions to Exa, Tavily, Parallel and Firecrawl, each with the parameters its own documentation recommends, on the same day and with the same number of results. The results are pooled and shuffled, and blind model judges verify every counted claim on its source page. A provider is credited only when its own result page states the claim, and a claim counts as new only if the knowledge base the questions came from did not already have it. Unlike the benchmarks above, each question ran once, so the study reports counts without intervals.

flowchart LR
  Q[pre-registered question sets<br/>from one knowledge base] --> E[Exa] & T[Tavily] & P[Parallel] & F[Firecrawl]
  E & T & P & F --> PO[(results pooled,<br/>deduplicated, shuffled)]
  PO --> J[blind model judges<br/>verify each claim on its page]
  J --> K{already in the<br/>knowledge base?}
  K --> A[attribution rejoined:<br/>credit only where the provider's<br/>own page states the claim]
  A --> R[reports/*/summary.json]

Its sets, endpoints, amendments and costs are in the study's methodology, and its full pre-registered record is in reports/2026-09-26-search-api-head-to-head/.

Reproduce, correct, contribute

Install Searchlight, run sew doctor, then follow each study's reproduce.md. Raw transcripts and provider responses are not distributed, so exact replay of the published numbers is not possible; a new run measures the same method. The one exception is the search API head-to-head's Monitors probe, whose raw responses are published. If you run a tested service and a configuration here misrepresents it, open an issue with the configuration you recommend: corrections are rerun and published beside the original, never silently replaced.

pipx install 'git+https://github.com/laceyenterprises/searchlight'
sew doctor
python3 scripts/check_reports.py
python3 scripts/build_site.py --check