hbmartin/agent-web-search

Client-first TypeScript search aggregation library for multiple web search providers.

TypeScript

2

63 commits

updated Oct 5, 2026

See the code

README

agent-web-search

Client-first TypeScript search aggregation library for multiple web search providers, built for AI agents.

Query several web search APIs through a single, normalized interface. One call fans out to every configured engine in parallel, and you get back a consistent result shape per engine — including per-engine errors, warnings, rate-limit info, and optional raw payloads — without one slow or failing provider blocking the rest. Then merge everything into one deduplicated, rank-fused list, format it for an LLM prompt, or expose the whole thing as an agent tool or MCP server.

  • One shape for every provider — SearchResult, Answer, and metadata are normalized across engines.
  • Per-engine isolation — each engine returns its own ok: true | false result; a failure in one never rejects the others.
  • Cross-engine aggregation — aggregate() dedupes by canonical URL and fuses rankings with reciprocal rank fusion.
  • LLM-ready output — formatForLLM() renders results as a compact markdown or XML block for prompts.
  • Agent tools built in — one-line tool definitions for the Anthropic API, OpenAI function calling, and the Vercel AI SDK, plus an MCP stdio server mode.
  • Execution strategies — fan out to all engines, race them, fall back in priority order, or hedge with staggered starts; with an overall deadline.
  • Cost & rate-limit aware — per-engine concurrency/pacing throttles, proactive backoff on exhausted provider rate limits, and a client-wide cost budget.
  • Streaming — consume results and answer deltas as they arrive via async iterables.
  • Capability-aware — unsupported query params are surfaced as warnings (or hard errors) instead of being silently dropped.
  • Bring your own engines — register custom adapters alongside the built-ins.
  • Typed and validated — runtime-validated with Zod; ESM + CJS builds with full type declarations.
  • Zero runtime deps beyond Zod (peer), browser-safe core, and a CLI for quick searches from the terminal.

Supported engines

EngineidCredentials env var (CLI)
BravebraveBRAVE_API_KEY
CeramicceramicCERAMIC_API_KEY
DuckDuckGo Instant Answersduckduckgo— (keyless)
ExaexaEXA_API_KEY
FirecrawlfirecrawlFIRECRAWL_API_KEY
GDELTgdelt— (keyless)
Hacker News (Algolia)hackernews— (keyless)
Jina SearchjinaJINA_API_KEY
KagikagiKAGI_API_KEY
LinkuplinkupLINKUP_API_KEY
ParallelparallelPARALLEL_API_KEY
SearXNG (self-hosted)searxngSEARXNG_BASE_URL (+ optional SEARXNG_API_KEY)
SerpAPIserpapiSERPAPI_API_KEY
Serper.devserperSERPER_API_KEY
Perplexity SonarsonarPERPLEXITY_API_KEY
TavilytavilyTAVILY_API_KEY
You.comyouYOU_API_KEY

Notes: duckduckgo hits the free Instant Answer API — encyclopedic abstracts and related topics, not full web results. searxng requires a self-hosted instance with the JSON output format enabled (search.formats: [html, json] in settings.yml). gdelt returns global news metadata only (title, URL, date, image) with a ~15 minute refresh — no snippets or page text; its publishedDate is GDELT's seendate, i.e. when GDELT first saw the article rather than the publisher's own date. GDELT's domain: filter performs suffix matching, so domain:example.com can also match a longer domain ending in example.com; this adapter intentionally preserves that fuzzy behavior rather than using the exact domainis: operator. hackernews queries the public Algolia index and defaults to tags=story; override tags for comments, ask_hn, show_hn, front_page, or author_<name>. linkup defaults to the cheaper searchResults output type; set defaults: { outputType: "sourcedAnswer" } for a cited answer, or defaults: { depth: "deep" } for its slower, broader crawl.

Installation

npm install agent-web-search zod
# or
pnpm add agent-web-search zod

zod (v4) is a peer dependency. Requires Node.js >= 22 for the CLI and Node builds; the library core also runs in browsers and edge runtimes (see Browser & edge usage).

Quick start

import { search } from "agent-web-search";

const response = await search(
  { query: "best espresso machines 2026", count: 5 },
  {
    brave: { apiKey: process.env.BRAVE_API_KEY! },
    exa: { apiKey: process.env.EXA_API_KEY! },
  },
);

for (const [engine, result] of Object.entries(response)) {
  if (result.ok) {
    console.log(`${engine}: ${result.results.length} results`);
    for (const item of result.results) {
      console.log(`  ${item.title} — ${item.url}`);
    }
  } else {
    console.error(`${engine} failed: ${result.error.kind} — ${result.error.message}`);
  }
}

search() returns a SearchResponse — a record keyed by engine id, where each value is either an EngineSuccess (ok: true) or an EngineFailure (ok: false).

Usage

Reusable client

Create a client once and reuse it for many searches. Configuration is validated up front.

import { createSearchClient } from "agent-web-search";

const client = createSearchClient({
  brave: { apiKey: process.env.BRAVE_API_KEY! },
  sonar: { apiKey: process.env.PERPLEXITY_API_KEY! },
});

const response = await client.search({ query: "what is a vector database" });

search() and searchStream() are one-shot convenience wrappers that build a client per call. Prefer createSearchClient when you issue more than one search — the cost budget and rate-limit state also live on the client.

Engine defaults are overridable: normalized query fields take precedence when supplied, and overrides take precedence over both. Domain filters must contain nonblank entries. Firecrawl's native includeDomains and excludeDomains defaults or overrides must be arrays; You also accepts comma-separated native GET filters. Invalid native defaults reject client creation, and invalid selected-engine overrides reject a search before any provider request. Native count values must be positive integers (numeric strings are rejected). Counts above a documented provider limit are clamped with a warning.

Aggregation: one deduplicated, rank-fused list

aggregate() merges a multi-engine response into a single result list. URLs are canonicalized for deduplication (protocol, www., fragments, trailing slashes, and tracking params like utm_*/gclid/fbclid are ignored) and ordered by reciprocal rank fusion: each engine contributes weight / (k + rank), so results that several engines agree on rise to the top.

import { aggregate } from "agent-web-search";

const merged = aggregate(response, {
  k: 60,                       // RRF smoothing constant (default 60)
  weights: { exa: 2 },         // trust some engines more
  maxResults: 10,
});

for (const result of merged.results) {
  // result.engines — which engines returned it
  // result.engineRank — its 1-based rank per engine
  // result.fusedScore — the RRF score used for ordering
  console.log(result.fusedScore.toFixed(4), result.engines, result.url);
}

merged.answers;   // Record<engine, Answer> from answer engines (sonar, tavily, …)
merged.succeeded; // engines that returned ok
merged.failed;    // Record<engine, SearchEngineError>

LLM-ready formatting

formatForLLM() turns a response (raw or pre-aggregated) into a compact, citation-friendly block to drop into a prompt.

import { formatForLLM } from "agent-web-search";

const block = formatForLLM(response, {
  format: "markdown",   // or "xml"
  maxResults: 8,
  maxSnippetChars: 400,
});

Markdown output has an ## Answers section (when engines produced answers) and a numbered ## Search results list with title, date, URL, snippet, and source engines. XML output emits <search_results> with <answer> and <result> elements, fully escaped.

Agent tool definitions

Ready-made web-search tools for the common LLM SDK wire formats — validation via Zod, execution via your configured client, output via formatForLLM.

import {
  aiSdkWebSearchTool,
  anthropicWebSearchTool,
  createSearchClient,
  openaiWebSearchTool,
} from "agent-web-search";
// or: import { ... } from "agent-web-search/tools";

const client = createSearchClient({ brave: { apiKey: "..." } });

// Anthropic API
const tool = anthropicWebSearchTool(client);
// tools: [{ name: tool.name, description: tool.description, input_schema: tool.input_schema }]
// on tool_use: const text = await tool.execute(toolUse.input);

// OpenAI function calling
const { definition, execute } = openaiWebSearchTool(client);

// Vercel AI SDK
// tools: { web_search: aiSdkWebSearchTool(client) }

All variants accept { name, description, format } options. The generic createWebSearchTool(client) exposes the Zod schema, the JSON schema, and execute for anything else.

MCP server mode

Run the library as a zero-dependency MCP stdio server exposing a web_search tool:

TAVILY_API_KEY=... agent-web-search mcp
# restrict engines:
BRAVE_API_KEY=... agent-web-search mcp --engine brave

For Claude Code: claude mcp add web-search -e TAVILY_API_KEY=... -- npx agent-web-search mcp.

Programmatic (Node-only) usage via the agent-web-search/mcp subpath:

import { runMcpServer } from "agent-web-search/mcp";
await runMcpServer(client, { serverVersion: "1.0.0" });

Execution strategies

By default every configured engine is queried in parallel ("all"). Three more strategies are available per client or per request:

// First success wins; everything else is aborted.
await client.search({ query }, { strategy: "race" });

// Try engines sequentially in priority order, stop at the first success.
await client.search({ query }, { strategy: "fallback", order: ["brave", "exa"] });

// Start engines staggered by hedgeDelayMs; first success aborts the rest.
await client.search({ query }, { strategy: "hedged", order: ["brave", "exa"], hedgeDelayMs: 300 });

// Overall deadline across all engines and retries (any strategy).
await client.search({ query }, { deadlineMs: 5000 });

With "fallback" and "hedged", engines that were never started are omitted from the response; with "race", aborted engines settle as failures and are included. searchStream always fans out to all engines but honors deadlineMs and order.

Throttling, rate limits, and cost budget

const client = createSearchClient(
  {
    brave: {
      apiKey: "...",
      throttle: { maxConcurrent: 2, minIntervalMs: 100 }, // client-side pacing
      costPerRequestUsd: 0.005,                            // your cost estimate
    },
  },
  {
    respectRateLimits: true,        // fail fast while a provider reports remaining: 0
    budget: { maxCostUsd: 1 },      // hard ceiling across all searches on this client
  },
);
  • throttle.maxConcurrent caps in-flight requests per engine; minIntervalMs spaces request starts.
  • With respectRateLimits: true, an engine whose last response reported an exhausted rate limit fails fast with a rate_limit error until the provider-reported reset time, instead of burning a request.
  • The budget accrues provider-reported costs (usage.costUsd, e.g. Exa) or your costPerRequestUsd estimate; once reached, engines fail fast with a quota error.

Streaming

searchStream yields events as each engine produces them. Engines that support native streaming (e.g. Sonar) emit answer_delta events; non-streaming engines emit their terminal events when they complete.

import { searchStream } from "agent-web-search";

const stream = searchStream(
  { query: "summarize the latest in fusion energy" },
  { sonar: { apiKey: process.env.PERPLEXITY_API_KEY! } },
);

for await (const event of stream) {
  switch (event.type) {
    case "answer_delta":
      process.stdout.write(event.text);
      break;
    case "answer_done":
      console.log("\ncitations:", event.answer.citations.length);
      break;
    case "results":
      console.log(`[${event.engine}] ${event.results.length} results`);
      break;
    case "metadata":
    case "error":
    case "done":
      break;
  }
}

Telemetry hooks

Observe requests, responses, retries, errors, and settlements without affecting results. Hooks can be set per client, per request, or per engine.

const client = createSearchClient(engines, {
  hooks: {
    onRequest: ({ engine, url, attempt }) => log(`→ ${engine} ${url} (try ${attempt})`),
    onResponse: ({ engine, status, latencyMs }) => log(`← ${engine} ${status} ${latencyMs}ms`),
    onRetry: ({ engine, delayMs }) => log(`↻ ${engine} retrying in ${delayMs}ms`),
    onError: ({ engine, error }) => log(`✗ ${engine} ${error.kind}`),
    onSettled: ({ engine, result }) => log(`✓ ${engine} ok=${result.ok}`),
  },
});

Cancellation

Pass an AbortSignal to cancel in-flight requests (and the stream).

const controller = new AbortController();
const promise = client.search({ query: "..." }, { signal: controller.signal });
controller.abort();

Custom engines

Implement an EngineAdapter and register it via options.adapters. defineEngine is an identity helper that preserves config types.

import { createSearchClient, defineEngine, KeyedEngineConfigSchema } from "agent-web-search";

const myAdapter = defineEngine({
  id: "my-engine",
  capabilities: { /* ... */ },
  configSchema: KeyedEngineConfigSchema,
  buildRequest(query, config, warnings) {
    return { method: "GET", url: "https://api.example.com/search", query: { q: query.query } };
  },
  parseResponse(res, ctx) {
    /* return an EngineResult */
  },
});

const client = createSearchClient(
  { "my-engine": { apiKey: "..." } },
  { adapters: [myAdapter] },
);

You can also import the built-in adapters individually:

import { braveAdapter } from "agent-web-search/adapters/brave";
import { exaAdapter, tavilyAdapter } from "agent-web-search/adapters";

See CONTRIBUTING.md for the full checklist to add a built-in adapter, and the examples directory for runnable samples.

Capability matrix

Engineanswercontentstreamingmulti-querycountdateRangefreshnessincludeDomainsexcludeDomainscountrylanguagesafeSearchverticals
brave————✓✓✓emulatedemulated✓✓✓web
ceramic————————————web
duckduckgo✓———————————web
exa—✓——✓✓✓nativenative✓——web, news
firecrawl—✓——✓✓✓nativenative✓——web, news, images
gdelt————✓✓✓emulatedemulated———news
hackernews————✓✓✓—————web
jina—✓——✓——emulatedemulated✓✓—web
kagi————✓———————web, news
linkup—✓——✓✓✓nativenative———web
parallel———✓✓✓✓nativenative✓——web
searxng✓———✓—✓emulatedemulated—✓✓web, news, images, video
serpapi————✓✓✓emulatedemulated✓✓✓web, news
serper————✓✓✓emulatedemulated✓✓—web, news
sonar✓—✓——✓✓nativeemulated—✓—web
tavily✓✓——✓✓✓nativenative———web, news
you—✓——✓✓✓nativenative✓✓✓web, news

native = handled by the provider's API; emulated = approximated by the adapter (e.g. via query operators). Serper and DuckDuckGo also surface opportunistic answers (Google answer box / instant answer) when the provider returns one, even where answer is not a guaranteed capability; so does Linkup when configured with outputType: "sourcedAnswer". GDELT emulates domain filters with its suffix-matching domain: operator rather than site: or exact domainis:, and exposes language/country only through GDELT-specific codes — reach those with overrides.

CLI

The package ships an agent-web-search binary. Set the relevant API key env vars, then:

agent-web-search --query "best espresso machines" --engine brave --engine exa --count 5

# merged + LLM-ready markdown
agent-web-search -q "fusion energy news" --format markdown

# aggregated JSON, racing engines with a 5s deadline
agent-web-search -q "vector databases" --aggregate --strategy race --deadline-ms 5000

# MCP stdio server
agent-web-search mcp

By default it queries every engine that has a matching API key set in the environment (the keyless duckduckgo, gdelt, and hackernews engines join only when named with --engine; searxng joins when SEARXNG_BASE_URL is set).

Options
  -q, --query <text>            Search query. Positional text is also accepted.
  -e, --engine <id>             Engine id. Repeat or comma-separate.
      --count <number>          Desired result count per engine.
      --freshness <range>       day, week, month, or year.
      --country <code>          ISO country code.
      --language <code>         ISO language code.
      --safe-search <mode>      off, moderate, or strict.
      --include-domain <host>   Domain allowlist. Repeat or comma-separate.
      --exclude-domain <host>   Domain blocklist. Repeat or comma-separate.
      --content <fields>        true or comma list: markdown,html,text,summary.
      --raw                     Include top-level raw provider payloads.
      --stream                  Emit stream events as NDJSON.
      --format <name>           json (default), ndjson, markdown, or xml.
      --ndjson                  Shorthand for --format ndjson.
      --aggregate               Merge engines into one deduplicated,
                                rank-fused list (json/ndjson output).
      --strategy <name>         all (default), race, fallback, or hedged.
      --deadline-ms <number>    Overall deadline across engines and retries.
  -h, --help                    Show help.
  -v, --version                 Show version.

Result shapes

type SearchResponse = Record<string, EngineResult>;
type EngineResult = EngineSuccess | EngineFailure;

A successful engine result (ok: true) contains results: SearchResult[], an optional answer, and metadata. A normalized SearchResult includes url, title, snippet/snippets, publishedDate, author, score, source, optional content/highlights/image/favicon, and the provider's raw payload.

A failed engine result (ok: false) contains an error whose kind is one of: auth, rate_limit, quota, bad_request, unsupported, timeout, network, upstream, or parse, along with metadata.

Both Zod schemas (SearchResponseSchema, SearchResultSchema, AnswerSchema, …) and TypeScript types are exported from the package root. API documentation is generated with TypeDoc and published from the docs workflow.

Browser & edge usage

Everything importable from the package root (client, adapters, aggregation, formatting, tools) is browser-safe: no Node builtins, fetch-based transport, works in browsers, Cloudflare Workers, Deno, and Bun. CI enforces this with an esbuild browser-platform bundle check (pnpm check:browser). The CLI and the agent-web-search/mcp subpath are Node-only.

Caveat: most search providers do not send CORS headers and your API keys should not ship to untrusted clients — in real browser apps, proxy provider calls through your backend (baseUrl is configurable per engine, and you can pass a custom fetch). The keyless GDELT and Hacker News endpoints are exceptions: both are CORS-enabled and can be called directly from browsers.

Development

pnpm install
pnpm build          # emit ESM + CJS builds to dist/
pnpm test           # run the vitest suite with coverage
pnpm typecheck      # tsc --noEmit
pnpm lint           # biome check
pnpm format         # biome check --write
pnpm check:browser  # browser-platform bundle smoke test
pnpm docs           # generate TypeDoc API docs

Requires Node.js >= 22. This repo uses pnpm (pinned via packageManager), Vitest for tests (including per-adapter contract fixtures in test/fixtures/, property-based tests with fast-check, and an env-gated live suite), and Biome for linting and formatting. Versions and changelogs are prepared manually with Changesets. A scheduled live workflow runs the real-API integration tests weekly to catch provider drift.

To release, run pnpm changeset version, commit and merge the version and changelog changes to main, then publish a GitHub Release tagged v<package.json version>. The release workflow publishes the tagged commit to npm exclusively through trusted publishing (GitHub OIDC), with provenance. Stable releases use latest; prereleases use next. See the release checklist for npm setup and the full release process.

License

MIT © Harold Martin

hbmartin/agent-web-search

Client-first TypeScript search aggregation library for multiple web search providers.

TypeScript

2

63 commits

updated Oct 5, 2026

See the code

README

agent-web-search

Client-first TypeScript search aggregation library for multiple web search providers, built for AI agents.

Query several web search APIs through a single, normalized interface. One call fans out to every configured engine in parallel, and you get back a consistent result shape per engine — including per-engine errors, warnings, rate-limit info, and optional raw payloads — without one slow or failing provider blocking the rest. Then merge everything into one deduplicated, rank-fused list, format it for an LLM prompt, or expose the whole thing as an agent tool or MCP server.

  • One shape for every provider — SearchResult, Answer, and metadata are normalized across engines.
  • Per-engine isolation — each engine returns its own ok: true | false result; a failure in one never rejects the others.
  • Cross-engine aggregation — aggregate() dedupes by canonical URL and fuses rankings with reciprocal rank fusion.
  • LLM-ready output — formatForLLM() renders results as a compact markdown or XML block for prompts.
  • Agent tools built in — one-line tool definitions for the Anthropic API, OpenAI function calling, and the Vercel AI SDK, plus an MCP stdio server mode.
  • Execution strategies — fan out to all engines, race them, fall back in priority order, or hedge with staggered starts; with an overall deadline.
  • Cost & rate-limit aware — per-engine concurrency/pacing throttles, proactive backoff on exhausted provider rate limits, and a client-wide cost budget.
  • Streaming — consume results and answer deltas as they arrive via async iterables.
  • Capability-aware — unsupported query params are surfaced as warnings (or hard errors) instead of being silently dropped.
  • Bring your own engines — register custom adapters alongside the built-ins.
  • Typed and validated — runtime-validated with Zod; ESM + CJS builds with full type declarations.
  • Zero runtime deps beyond Zod (peer), browser-safe core, and a CLI for quick searches from the terminal.

Supported engines

EngineidCredentials env var (CLI)
BravebraveBRAVE_API_KEY
CeramicceramicCERAMIC_API_KEY
DuckDuckGo Instant Answersduckduckgo— (keyless)
ExaexaEXA_API_KEY
FirecrawlfirecrawlFIRECRAWL_API_KEY
GDELTgdelt— (keyless)
Hacker News (Algolia)hackernews— (keyless)
Jina SearchjinaJINA_API_KEY
KagikagiKAGI_API_KEY
LinkuplinkupLINKUP_API_KEY
ParallelparallelPARALLEL_API_KEY
SearXNG (self-hosted)searxngSEARXNG_BASE_URL (+ optional SEARXNG_API_KEY)
SerpAPIserpapiSERPAPI_API_KEY
Serper.devserperSERPER_API_KEY
Perplexity SonarsonarPERPLEXITY_API_KEY
TavilytavilyTAVILY_API_KEY
You.comyouYOU_API_KEY

Notes: duckduckgo hits the free Instant Answer API — encyclopedic abstracts and related topics, not full web results. searxng requires a self-hosted instance with the JSON output format enabled (search.formats: [html, json] in settings.yml). gdelt returns global news metadata only (title, URL, date, image) with a ~15 minute refresh — no snippets or page text; its publishedDate is GDELT's seendate, i.e. when GDELT first saw the article rather than the publisher's own date. GDELT's domain: filter performs suffix matching, so domain:example.com can also match a longer domain ending in example.com; this adapter intentionally preserves that fuzzy behavior rather than using the exact domainis: operator. hackernews queries the public Algolia index and defaults to tags=story; override tags for comments, ask_hn, show_hn, front_page, or author_<name>. linkup defaults to the cheaper searchResults output type; set defaults: { outputType: "sourcedAnswer" } for a cited answer, or defaults: { depth: "deep" } for its slower, broader crawl.

Installation

npm install agent-web-search zod
# or
pnpm add agent-web-search zod

zod (v4) is a peer dependency. Requires Node.js >= 22 for the CLI and Node builds; the library core also runs in browsers and edge runtimes (see Browser & edge usage).

Quick start

import { search } from "agent-web-search";

const response = await search(
  { query: "best espresso machines 2026", count: 5 },
  {
    brave: { apiKey: process.env.BRAVE_API_KEY! },
    exa: { apiKey: process.env.EXA_API_KEY! },
  },
);

for (const [engine, result] of Object.entries(response)) {
  if (result.ok) {
    console.log(`${engine}: ${result.results.length} results`);
    for (const item of result.results) {
      console.log(`  ${item.title} — ${item.url}`);
    }
  } else {
    console.error(`${engine} failed: ${result.error.kind} — ${result.error.message}`);
  }
}

search() returns a SearchResponse — a record keyed by engine id, where each value is either an EngineSuccess (ok: true) or an EngineFailure (ok: false).

Usage

Reusable client

Create a client once and reuse it for many searches. Configuration is validated up front.

import { createSearchClient } from "agent-web-search";

const client = createSearchClient({
  brave: { apiKey: process.env.BRAVE_API_KEY! },
  sonar: { apiKey: process.env.PERPLEXITY_API_KEY! },
});

const response = await client.search({ query: "what is a vector database" });

search() and searchStream() are one-shot convenience wrappers that build a client per call. Prefer createSearchClient when you issue more than one search — the cost budget and rate-limit state also live on the client.

Engine defaults are overridable: normalized query fields take precedence when supplied, and overrides take precedence over both. Domain filters must contain nonblank entries. Firecrawl's native includeDomains and excludeDomains defaults or overrides must be arrays; You also accepts comma-separated native GET filters. Invalid native defaults reject client creation, and invalid selected-engine overrides reject a search before any provider request. Native count values must be positive integers (numeric strings are rejected). Counts above a documented provider limit are clamped with a warning.

Aggregation: one deduplicated, rank-fused list

aggregate() merges a multi-engine response into a single result list. URLs are canonicalized for deduplication (protocol, www., fragments, trailing slashes, and tracking params like utm_*/gclid/fbclid are ignored) and ordered by reciprocal rank fusion: each engine contributes weight / (k + rank), so results that several engines agree on rise to the top.

import { aggregate } from "agent-web-search";

const merged = aggregate(response, {
  k: 60,                       // RRF smoothing constant (default 60)
  weights: { exa: 2 },         // trust some engines more
  maxResults: 10,
});

for (const result of merged.results) {
  // result.engines — which engines returned it
  // result.engineRank — its 1-based rank per engine
  // result.fusedScore — the RRF score used for ordering
  console.log(result.fusedScore.toFixed(4), result.engines, result.url);
}

merged.answers;   // Record<engine, Answer> from answer engines (sonar, tavily, …)
merged.succeeded; // engines that returned ok
merged.failed;    // Record<engine, SearchEngineError>

LLM-ready formatting

formatForLLM() turns a response (raw or pre-aggregated) into a compact, citation-friendly block to drop into a prompt.

import { formatForLLM } from "agent-web-search";

const block = formatForLLM(response, {
  format: "markdown",   // or "xml"
  maxResults: 8,
  maxSnippetChars: 400,
});

Markdown output has an ## Answers section (when engines produced answers) and a numbered ## Search results list with title, date, URL, snippet, and source engines. XML output emits <search_results> with <answer> and <result> elements, fully escaped.

Agent tool definitions

Ready-made web-search tools for the common LLM SDK wire formats — validation via Zod, execution via your configured client, output via formatForLLM.

import {
  aiSdkWebSearchTool,
  anthropicWebSearchTool,
  createSearchClient,
  openaiWebSearchTool,
} from "agent-web-search";
// or: import { ... } from "agent-web-search/tools";

const client = createSearchClient({ brave: { apiKey: "..." } });

// Anthropic API
const tool = anthropicWebSearchTool(client);
// tools: [{ name: tool.name, description: tool.description, input_schema: tool.input_schema }]
// on tool_use: const text = await tool.execute(toolUse.input);

// OpenAI function calling
const { definition, execute } = openaiWebSearchTool(client);

// Vercel AI SDK
// tools: { web_search: aiSdkWebSearchTool(client) }

All variants accept { name, description, format } options. The generic createWebSearchTool(client) exposes the Zod schema, the JSON schema, and execute for anything else.

MCP server mode

Run the library as a zero-dependency MCP stdio server exposing a web_search tool:

TAVILY_API_KEY=... agent-web-search mcp
# restrict engines:
BRAVE_API_KEY=... agent-web-search mcp --engine brave

For Claude Code: claude mcp add web-search -e TAVILY_API_KEY=... -- npx agent-web-search mcp.

Programmatic (Node-only) usage via the agent-web-search/mcp subpath:

import { runMcpServer } from "agent-web-search/mcp";
await runMcpServer(client, { serverVersion: "1.0.0" });

Execution strategies

By default every configured engine is queried in parallel ("all"). Three more strategies are available per client or per request:

// First success wins; everything else is aborted.
await client.search({ query }, { strategy: "race" });

// Try engines sequentially in priority order, stop at the first success.
await client.search({ query }, { strategy: "fallback", order: ["brave", "exa"] });

// Start engines staggered by hedgeDelayMs; first success aborts the rest.
await client.search({ query }, { strategy: "hedged", order: ["brave", "exa"], hedgeDelayMs: 300 });

// Overall deadline across all engines and retries (any strategy).
await client.search({ query }, { deadlineMs: 5000 });

With "fallback" and "hedged", engines that were never started are omitted from the response; with "race", aborted engines settle as failures and are included. searchStream always fans out to all engines but honors deadlineMs and order.

Throttling, rate limits, and cost budget

const client = createSearchClient(
  {
    brave: {
      apiKey: "...",
      throttle: { maxConcurrent: 2, minIntervalMs: 100 }, // client-side pacing
      costPerRequestUsd: 0.005,                            // your cost estimate
    },
  },
  {
    respectRateLimits: true,        // fail fast while a provider reports remaining: 0
    budget: { maxCostUsd: 1 },      // hard ceiling across all searches on this client
  },
);
  • throttle.maxConcurrent caps in-flight requests per engine; minIntervalMs spaces request starts.
  • With respectRateLimits: true, an engine whose last response reported an exhausted rate limit fails fast with a rate_limit error until the provider-reported reset time, instead of burning a request.
  • The budget accrues provider-reported costs (usage.costUsd, e.g. Exa) or your costPerRequestUsd estimate; once reached, engines fail fast with a quota error.

Streaming

searchStream yields events as each engine produces them. Engines that support native streaming (e.g. Sonar) emit answer_delta events; non-streaming engines emit their terminal events when they complete.

import { searchStream } from "agent-web-search";

const stream = searchStream(
  { query: "summarize the latest in fusion energy" },
  { sonar: { apiKey: process.env.PERPLEXITY_API_KEY! } },
);

for await (const event of stream) {
  switch (event.type) {
    case "answer_delta":
      process.stdout.write(event.text);
      break;
    case "answer_done":
      console.log("\ncitations:", event.answer.citations.length);
      break;
    case "results":
      console.log(`[${event.engine}] ${event.results.length} results`);
      break;
    case "metadata":
    case "error":
    case "done":
      break;
  }
}

Telemetry hooks

Observe requests, responses, retries, errors, and settlements without affecting results. Hooks can be set per client, per request, or per engine.

const client = createSearchClient(engines, {
  hooks: {
    onRequest: ({ engine, url, attempt }) => log(`→ ${engine} ${url} (try ${attempt})`),
    onResponse: ({ engine, status, latencyMs }) => log(`← ${engine} ${status} ${latencyMs}ms`),
    onRetry: ({ engine, delayMs }) => log(`↻ ${engine} retrying in ${delayMs}ms`),
    onError: ({ engine, error }) => log(`✗ ${engine} ${error.kind}`),
    onSettled: ({ engine, result }) => log(`✓ ${engine} ok=${result.ok}`),
  },
});

Cancellation

Pass an AbortSignal to cancel in-flight requests (and the stream).

const controller = new AbortController();
const promise = client.search({ query: "..." }, { signal: controller.signal });
controller.abort();

Custom engines

Implement an EngineAdapter and register it via options.adapters. defineEngine is an identity helper that preserves config types.

import { createSearchClient, defineEngine, KeyedEngineConfigSchema } from "agent-web-search";

const myAdapter = defineEngine({
  id: "my-engine",
  capabilities: { /* ... */ },
  configSchema: KeyedEngineConfigSchema,
  buildRequest(query, config, warnings) {
    return { method: "GET", url: "https://api.example.com/search", query: { q: query.query } };
  },
  parseResponse(res, ctx) {
    /* return an EngineResult */
  },
});

const client = createSearchClient(
  { "my-engine": { apiKey: "..." } },
  { adapters: [myAdapter] },
);

You can also import the built-in adapters individually:

import { braveAdapter } from "agent-web-search/adapters/brave";
import { exaAdapter, tavilyAdapter } from "agent-web-search/adapters";

See CONTRIBUTING.md for the full checklist to add a built-in adapter, and the examples directory for runnable samples.

Capability matrix

Engineanswercontentstreamingmulti-querycountdateRangefreshnessincludeDomainsexcludeDomainscountrylanguagesafeSearchverticals
brave————✓✓✓emulatedemulated✓✓✓web
ceramic————————————web
duckduckgo✓———————————web
exa—✓——✓✓✓nativenative✓——web, news
firecrawl—✓——✓✓✓nativenative✓——web, news, images
gdelt————✓✓✓emulatedemulated———news
hackernews————✓✓✓—————web
jina—✓——✓——emulatedemulated✓✓—web
kagi————✓———————web, news
linkup—✓——✓✓✓nativenative———web
parallel———✓✓✓✓nativenative✓——web
searxng✓———✓—✓emulatedemulated—✓✓web, news, images, video
serpapi————✓✓✓emulatedemulated✓✓✓web, news
serper————✓✓✓emulatedemulated✓✓—web, news
sonar✓—✓——✓✓nativeemulated—✓—web
tavily✓✓——✓✓✓nativenative———web, news
you—✓——✓✓✓nativenative✓✓✓web, news

native = handled by the provider's API; emulated = approximated by the adapter (e.g. via query operators). Serper and DuckDuckGo also surface opportunistic answers (Google answer box / instant answer) when the provider returns one, even where answer is not a guaranteed capability; so does Linkup when configured with outputType: "sourcedAnswer". GDELT emulates domain filters with its suffix-matching domain: operator rather than site: or exact domainis:, and exposes language/country only through GDELT-specific codes — reach those with overrides.

CLI

The package ships an agent-web-search binary. Set the relevant API key env vars, then:

agent-web-search --query "best espresso machines" --engine brave --engine exa --count 5

# merged + LLM-ready markdown
agent-web-search -q "fusion energy news" --format markdown

# aggregated JSON, racing engines with a 5s deadline
agent-web-search -q "vector databases" --aggregate --strategy race --deadline-ms 5000

# MCP stdio server
agent-web-search mcp

By default it queries every engine that has a matching API key set in the environment (the keyless duckduckgo, gdelt, and hackernews engines join only when named with --engine; searxng joins when SEARXNG_BASE_URL is set).

Options
  -q, --query <text>            Search query. Positional text is also accepted.
  -e, --engine <id>             Engine id. Repeat or comma-separate.
      --count <number>          Desired result count per engine.
      --freshness <range>       day, week, month, or year.
      --country <code>          ISO country code.
      --language <code>         ISO language code.
      --safe-search <mode>      off, moderate, or strict.
      --include-domain <host>   Domain allowlist. Repeat or comma-separate.
      --exclude-domain <host>   Domain blocklist. Repeat or comma-separate.
      --content <fields>        true or comma list: markdown,html,text,summary.
      --raw                     Include top-level raw provider payloads.
      --stream                  Emit stream events as NDJSON.
      --format <name>           json (default), ndjson, markdown, or xml.
      --ndjson                  Shorthand for --format ndjson.
      --aggregate               Merge engines into one deduplicated,
                                rank-fused list (json/ndjson output).
      --strategy <name>         all (default), race, fallback, or hedged.
      --deadline-ms <number>    Overall deadline across engines and retries.
  -h, --help                    Show help.
  -v, --version                 Show version.

Result shapes

type SearchResponse = Record<string, EngineResult>;
type EngineResult = EngineSuccess | EngineFailure;

A successful engine result (ok: true) contains results: SearchResult[], an optional answer, and metadata. A normalized SearchResult includes url, title, snippet/snippets, publishedDate, author, score, source, optional content/highlights/image/favicon, and the provider's raw payload.

A failed engine result (ok: false) contains an error whose kind is one of: auth, rate_limit, quota, bad_request, unsupported, timeout, network, upstream, or parse, along with metadata.

Both Zod schemas (SearchResponseSchema, SearchResultSchema, AnswerSchema, …) and TypeScript types are exported from the package root. API documentation is generated with TypeDoc and published from the docs workflow.

Browser & edge usage

Everything importable from the package root (client, adapters, aggregation, formatting, tools) is browser-safe: no Node builtins, fetch-based transport, works in browsers, Cloudflare Workers, Deno, and Bun. CI enforces this with an esbuild browser-platform bundle check (pnpm check:browser). The CLI and the agent-web-search/mcp subpath are Node-only.

Caveat: most search providers do not send CORS headers and your API keys should not ship to untrusted clients — in real browser apps, proxy provider calls through your backend (baseUrl is configurable per engine, and you can pass a custom fetch). The keyless GDELT and Hacker News endpoints are exceptions: both are CORS-enabled and can be called directly from browsers.

Development

pnpm install
pnpm build          # emit ESM + CJS builds to dist/
pnpm test           # run the vitest suite with coverage
pnpm typecheck      # tsc --noEmit
pnpm lint           # biome check
pnpm format         # biome check --write
pnpm check:browser  # browser-platform bundle smoke test
pnpm docs           # generate TypeDoc API docs

Requires Node.js >= 22. This repo uses pnpm (pinned via packageManager), Vitest for tests (including per-adapter contract fixtures in test/fixtures/, property-based tests with fast-check, and an env-gated live suite), and Biome for linting and formatting. Versions and changelogs are prepared manually with Changesets. A scheduled live workflow runs the real-API integration tests weekly to catch provider drift.

To release, run pnpm changeset version, commit and merge the version and changelog changes to main, then publish a GitHub Release tagged v<package.json version>. The release workflow publishes the tagged commit to npm exclusively through trusted publishing (GitHub OIDC), with provenance. Stable releases use latest; prereleases use next. See the release checklist for npm setup and the full release process.

License

MIT © Harold Martin