Client-first TypeScript search aggregation library for multiple web search providers.
See the codeClient-first TypeScript search aggregation library for multiple web search providers, built for AI agents.
Query several web search APIs through a single, normalized interface. One call fans out to every configured engine in parallel, and you get back a consistent result shape per engine — including per-engine errors, warnings, rate-limit info, and optional raw payloads — without one slow or failing provider blocking the rest. Then merge everything into one deduplicated, rank-fused list, format it for an LLM prompt, or expose the whole thing as an agent tool or MCP server.
SearchResult, Answer, and metadata are normalized across engines.ok: true | false result; a failure in one never rejects the others.aggregate() dedupes by canonical URL and fuses rankings with reciprocal rank fusion.formatForLLM() renders results as a compact markdown or XML block for prompts.| Engine | id | Credentials env var (CLI) |
|---|---|---|
| Brave | brave | BRAVE_API_KEY |
| Ceramic | ceramic | CERAMIC_API_KEY |
| DuckDuckGo Instant Answers | duckduckgo | — (keyless) |
| Exa | exa | EXA_API_KEY |
| Firecrawl | firecrawl | FIRECRAWL_API_KEY |
| GDELT | gdelt | — (keyless) |
| Hacker News (Algolia) | hackernews | — (keyless) |
| Jina Search | jina | JINA_API_KEY |
| Kagi | kagi | KAGI_API_KEY |
| Linkup | linkup | LINKUP_API_KEY |
| Parallel | parallel | PARALLEL_API_KEY |
| SearXNG (self-hosted) | searxng | SEARXNG_BASE_URL (+ optional SEARXNG_API_KEY) |
| SerpAPI | serpapi | SERPAPI_API_KEY |
| Serper.dev | serper | SERPER_API_KEY |
| Perplexity Sonar | sonar | PERPLEXITY_API_KEY |
| Tavily | tavily | TAVILY_API_KEY |
| You.com | you | YOU_API_KEY |
Notes: duckduckgo hits the free Instant Answer API — encyclopedic abstracts and related topics, not full web results. searxng requires a self-hosted instance with the JSON output format enabled (search.formats: [html, json] in settings.yml). gdelt returns global news metadata only (title, URL, date, image) with a ~15 minute refresh — no snippets or page text; its publishedDate is GDELT's seendate, i.e. when GDELT first saw the article rather than the publisher's own date. GDELT's domain: filter performs suffix matching, so domain:example.com can also match a longer domain ending in example.com; this adapter intentionally preserves that fuzzy behavior rather than using the exact domainis: operator. hackernews queries the public Algolia index and defaults to tags=story; override tags for comments, ask_hn, show_hn, front_page, or author_<name>. linkup defaults to the cheaper searchResults output type; set defaults: { outputType: "sourcedAnswer" } for a cited answer, or defaults: { depth: "deep" } for its slower, broader crawl.
npm install agent-web-search zod
# or
pnpm add agent-web-search zod
zod (v4) is a peer dependency. Requires Node.js >= 22 for the CLI and Node builds; the library core also runs in browsers and edge runtimes (see Browser & edge usage).
import { search } from "agent-web-search";
const response = await search(
{ query: "best espresso machines 2026", count: 5 },
{
brave: { apiKey: process.env.BRAVE_API_KEY! },
exa: { apiKey: process.env.EXA_API_KEY! },
},
);
for (const [engine, result] of Object.entries(response)) {
if (result.ok) {
console.log(`${engine}: ${result.results.length} results`);
for (const item of result.results) {
console.log(` ${item.title} — ${item.url}`);
}
} else {
console.error(`${engine} failed: ${result.error.kind} — ${result.error.message}`);
}
}
search() returns a SearchResponse — a record keyed by engine id, where each value is either an EngineSuccess (ok: true) or an EngineFailure (ok: false).
Create a client once and reuse it for many searches. Configuration is validated up front.
import { createSearchClient } from "agent-web-search";
const client = createSearchClient({
brave: { apiKey: process.env.BRAVE_API_KEY! },
sonar: { apiKey: process.env.PERPLEXITY_API_KEY! },
});
const response = await client.search({ query: "what is a vector database" });
search() and searchStream() are one-shot convenience wrappers that build a client per call. Prefer createSearchClient when you issue more than one search — the cost budget and rate-limit state also live on the client.
Engine defaults are overridable: normalized query fields take precedence when supplied, and overrides take precedence over both. Domain filters must contain nonblank entries. Firecrawl's native includeDomains and excludeDomains defaults or overrides must be arrays; You also accepts comma-separated native GET filters. Invalid native defaults reject client creation, and invalid selected-engine overrides reject a search before any provider request. Native count values must be positive integers (numeric strings are rejected). Counts above a documented provider limit are clamped with a warning.
aggregate() merges a multi-engine response into a single result list. URLs are canonicalized for deduplication (protocol, www., fragments, trailing slashes, and tracking params like utm_*/gclid/fbclid are ignored) and ordered by reciprocal rank fusion: each engine contributes weight / (k + rank), so results that several engines agree on rise to the top.
import { aggregate } from "agent-web-search";
const merged = aggregate(response, {
k: 60, // RRF smoothing constant (default 60)
weights: { exa: 2 }, // trust some engines more
maxResults: 10,
});
for (const result of merged.results) {
// result.engines — which engines returned it
// result.engineRank — its 1-based rank per engine
// result.fusedScore — the RRF score used for ordering
console.log(result.fusedScore.toFixed(4), result.engines, result.url);
}
merged.answers; // Record<engine, Answer> from answer engines (sonar, tavily, …)
merged.succeeded; // engines that returned ok
merged.failed; // Record<engine, SearchEngineError>
formatForLLM() turns a response (raw or pre-aggregated) into a compact, citation-friendly block to drop into a prompt.
import { formatForLLM } from "agent-web-search";
const block = formatForLLM(response, {
format: "markdown", // or "xml"
maxResults: 8,
maxSnippetChars: 400,
});
Markdown output has an ## Answers section (when engines produced answers) and a numbered ## Search results list with title, date, URL, snippet, and source engines. XML output emits <search_results> with <answer> and <result> elements, fully escaped.
Ready-made web-search tools for the common LLM SDK wire formats — validation via Zod, execution via your configured client, output via formatForLLM.
import {
aiSdkWebSearchTool,
anthropicWebSearchTool,
createSearchClient,
openaiWebSearchTool,
} from "agent-web-search";
// or: import { ... } from "agent-web-search/tools";
const client = createSearchClient({ brave: { apiKey: "..." } });
// Anthropic API
const tool = anthropicWebSearchTool(client);
// tools: [{ name: tool.name, description: tool.description, input_schema: tool.input_schema }]
// on tool_use: const text = await tool.execute(toolUse.input);
// OpenAI function calling
const { definition, execute } = openaiWebSearchTool(client);
// Vercel AI SDK
// tools: { web_search: aiSdkWebSearchTool(client) }
All variants accept { name, description, format } options. The generic createWebSearchTool(client) exposes the Zod schema, the JSON schema, and execute for anything else.
Run the library as a zero-dependency MCP stdio server exposing a web_search tool:
TAVILY_API_KEY=... agent-web-search mcp
# restrict engines:
BRAVE_API_KEY=... agent-web-search mcp --engine brave
For Claude Code: claude mcp add web-search -e TAVILY_API_KEY=... -- npx agent-web-search mcp.
Programmatic (Node-only) usage via the agent-web-search/mcp subpath:
import { runMcpServer } from "agent-web-search/mcp";
await runMcpServer(client, { serverVersion: "1.0.0" });
By default every configured engine is queried in parallel ("all"). Three more strategies are available per client or per request:
// First success wins; everything else is aborted.
await client.search({ query }, { strategy: "race" });
// Try engines sequentially in priority order, stop at the first success.
await client.search({ query }, { strategy: "fallback", order: ["brave", "exa"] });
// Start engines staggered by hedgeDelayMs; first success aborts the rest.
await client.search({ query }, { strategy: "hedged", order: ["brave", "exa"], hedgeDelayMs: 300 });
// Overall deadline across all engines and retries (any strategy).
await client.search({ query }, { deadlineMs: 5000 });
With "fallback" and "hedged", engines that were never started are omitted from the response; with "race", aborted engines settle as failures and are included. searchStream always fans out to all engines but honors deadlineMs and order.
const client = createSearchClient(
{
brave: {
apiKey: "...",
throttle: { maxConcurrent: 2, minIntervalMs: 100 }, // client-side pacing
costPerRequestUsd: 0.005, // your cost estimate
},
},
{
respectRateLimits: true, // fail fast while a provider reports remaining: 0
budget: { maxCostUsd: 1 }, // hard ceiling across all searches on this client
},
);
throttle.maxConcurrent caps in-flight requests per engine; minIntervalMs spaces request starts.respectRateLimits: true, an engine whose last response reported an exhausted rate limit fails fast with a rate_limit error until the provider-reported reset time, instead of burning a request.usage.costUsd, e.g. Exa) or your costPerRequestUsd estimate; once reached, engines fail fast with a quota error.searchStream yields events as each engine produces them. Engines that support native streaming (e.g. Sonar) emit answer_delta events; non-streaming engines emit their terminal events when they complete.
import { searchStream } from "agent-web-search";
const stream = searchStream(
{ query: "summarize the latest in fusion energy" },
{ sonar: { apiKey: process.env.PERPLEXITY_API_KEY! } },
);
for await (const event of stream) {
switch (event.type) {
case "answer_delta":
process.stdout.write(event.text);
break;
case "answer_done":
console.log("\ncitations:", event.answer.citations.length);
break;
case "results":
console.log(`[${event.engine}] ${event.results.length} results`);
break;
case "metadata":
case "error":
case "done":
break;
}
}
Observe requests, responses, retries, errors, and settlements without affecting results. Hooks can be set per client, per request, or per engine.
const client = createSearchClient(engines, {
hooks: {
onRequest: ({ engine, url, attempt }) => log(`→ ${engine} ${url} (try ${attempt})`),
onResponse: ({ engine, status, latencyMs }) => log(`← ${engine} ${status} ${latencyMs}ms`),
onRetry: ({ engine, delayMs }) => log(`↻ ${engine} retrying in ${delayMs}ms`),
onError: ({ engine, error }) => log(`✗ ${engine} ${error.kind}`),
onSettled: ({ engine, result }) => log(`✓ ${engine} ok=${result.ok}`),
},
});
Pass an AbortSignal to cancel in-flight requests (and the stream).
const controller = new AbortController();
const promise = client.search({ query: "..." }, { signal: controller.signal });
controller.abort();
Implement an EngineAdapter and register it via options.adapters. defineEngine is an identity helper that preserves config types.
import { createSearchClient, defineEngine, KeyedEngineConfigSchema } from "agent-web-search";
const myAdapter = defineEngine({
id: "my-engine",
capabilities: { /* ... */ },
configSchema: KeyedEngineConfigSchema,
buildRequest(query, config, warnings) {
return { method: "GET", url: "https://api.example.com/search", query: { q: query.query } };
},
parseResponse(res, ctx) {
/* return an EngineResult */
},
});
const client = createSearchClient(
{ "my-engine": { apiKey: "..." } },
{ adapters: [myAdapter] },
);
You can also import the built-in adapters individually:
import { braveAdapter } from "agent-web-search/adapters/brave";
import { exaAdapter, tavilyAdapter } from "agent-web-search/adapters";
See CONTRIBUTING.md for the full checklist to add a built-in adapter, and the examples directory for runnable samples.
| Engine | answer | content | streaming | multi-query | count | dateRange | freshness | includeDomains | excludeDomains | country | language | safeSearch | verticals |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| brave | — | — | — | — | ✓ | ✓ | ✓ | emulated | emulated | ✓ | ✓ | ✓ | web |
| ceramic | — | — | — | — | — | — | — | — | — | — | — | — | web |
| duckduckgo | ✓ | — | — | — | — | — | — | — | — | — | — | — | web |
| exa | — | ✓ | — | — | ✓ | ✓ | ✓ | native | native | ✓ | — | — | web, news |
| firecrawl | — | ✓ | — | — | ✓ | ✓ | ✓ | native | native | ✓ | — | — | web, news, images |
| gdelt | — | — | — | — | ✓ | ✓ | ✓ | emulated | emulated | — | — | — | news |
| hackernews | — | — | — | — | ✓ | ✓ | ✓ | — | — | — | — | — | web |
| jina | — | ✓ | — | — | ✓ | — | — | emulated | emulated | ✓ | ✓ | — | web |
| kagi | — | — | — | — | ✓ | — | — | — | — | — | — | — | web, news |
| linkup | — | ✓ | — | — | ✓ | ✓ | ✓ | native | native | — | — | — | web |
| parallel | — | — | — | ✓ | ✓ | ✓ | ✓ | native | native | ✓ | — | — | web |
| searxng | ✓ | — | — | — | ✓ | — | ✓ | emulated | emulated | — | ✓ | ✓ | web, news, images, video |
| serpapi | — | — | — | — | ✓ | ✓ | ✓ | emulated | emulated | ✓ | ✓ | ✓ | web, news |
| serper | — | — | — | — | ✓ | ✓ | ✓ | emulated | emulated | ✓ | ✓ | — | web, news |
| sonar | ✓ | — | ✓ | — | — | ✓ | ✓ | native | emulated | — | ✓ | — | web |
| tavily | ✓ | ✓ | — | — | ✓ | ✓ | ✓ | native | native | — | — | — | web, news |
| you | — | ✓ | — | — | ✓ | ✓ | ✓ | native | native | ✓ | ✓ | ✓ | web, news |
native = handled by the provider's API; emulated = approximated by the adapter (e.g. via query operators). Serper and DuckDuckGo also surface opportunistic answers (Google answer box / instant answer) when the provider returns one, even where answer is not a guaranteed capability; so does Linkup when configured with outputType: "sourcedAnswer". GDELT emulates domain filters with its suffix-matching domain: operator rather than site: or exact domainis:, and exposes language/country only through GDELT-specific codes — reach those with overrides.
The package ships an agent-web-search binary. Set the relevant API key env vars, then:
agent-web-search --query "best espresso machines" --engine brave --engine exa --count 5
# merged + LLM-ready markdown
agent-web-search -q "fusion energy news" --format markdown
# aggregated JSON, racing engines with a 5s deadline
agent-web-search -q "vector databases" --aggregate --strategy race --deadline-ms 5000
# MCP stdio server
agent-web-search mcp
By default it queries every engine that has a matching API key set in the environment (the keyless duckduckgo, gdelt, and hackernews engines join only when named with --engine; searxng joins when SEARXNG_BASE_URL is set).
Options
-q, --query <text> Search query. Positional text is also accepted.
-e, --engine <id> Engine id. Repeat or comma-separate.
--count <number> Desired result count per engine.
--freshness <range> day, week, month, or year.
--country <code> ISO country code.
--language <code> ISO language code.
--safe-search <mode> off, moderate, or strict.
--include-domain <host> Domain allowlist. Repeat or comma-separate.
--exclude-domain <host> Domain blocklist. Repeat or comma-separate.
--content <fields> true or comma list: markdown,html,text,summary.
--raw Include top-level raw provider payloads.
--stream Emit stream events as NDJSON.
--format <name> json (default), ndjson, markdown, or xml.
--ndjson Shorthand for --format ndjson.
--aggregate Merge engines into one deduplicated,
rank-fused list (json/ndjson output).
--strategy <name> all (default), race, fallback, or hedged.
--deadline-ms <number> Overall deadline across engines and retries.
-h, --help Show help.
-v, --version Show version.
type SearchResponse = Record<string, EngineResult>;
type EngineResult = EngineSuccess | EngineFailure;
A successful engine result (ok: true) contains results: SearchResult[], an optional answer, and metadata. A normalized SearchResult includes url, title, snippet/snippets, publishedDate, author, score, source, optional content/highlights/image/favicon, and the provider's raw payload.
A failed engine result (ok: false) contains an error whose kind is one of: auth, rate_limit, quota, bad_request, unsupported, timeout, network, upstream, or parse, along with metadata.
Both Zod schemas (SearchResponseSchema, SearchResultSchema, AnswerSchema, …) and TypeScript types are exported from the package root. API documentation is generated with TypeDoc and published from the docs workflow.
Everything importable from the package root (client, adapters, aggregation, formatting, tools) is browser-safe: no Node builtins, fetch-based transport, works in browsers, Cloudflare Workers, Deno, and Bun. CI enforces this with an esbuild browser-platform bundle check (pnpm check:browser). The CLI and the agent-web-search/mcp subpath are Node-only.
Caveat: most search providers do not send CORS headers and your API keys should not ship to untrusted clients — in real browser apps, proxy provider calls through your backend (baseUrl is configurable per engine, and you can pass a custom fetch). The keyless GDELT and Hacker News endpoints are exceptions: both are CORS-enabled and can be called directly from browsers.
pnpm install
pnpm build # emit ESM + CJS builds to dist/
pnpm test # run the vitest suite with coverage
pnpm typecheck # tsc --noEmit
pnpm lint # biome check
pnpm format # biome check --write
pnpm check:browser # browser-platform bundle smoke test
pnpm docs # generate TypeDoc API docs
Requires Node.js >= 22. This repo uses pnpm (pinned via packageManager), Vitest for tests (including per-adapter contract fixtures in test/fixtures/, property-based tests with fast-check, and an env-gated live suite), and Biome for linting and formatting. Versions and changelogs are prepared manually with Changesets. A scheduled live workflow runs the real-API integration tests weekly to catch provider drift.
To release, run pnpm changeset version, commit and merge the version and changelog changes to main, then publish a GitHub Release tagged v<package.json version>. The release workflow publishes the tagged commit to npm exclusively through trusted publishing (GitHub OIDC), with provenance. Stable releases use latest; prereleases use next. See the release checklist for npm setup and the full release process.
MIT © Harold Martin
Client-first TypeScript search aggregation library for multiple web search providers.
See the codeClient-first TypeScript search aggregation library for multiple web search providers, built for AI agents.
Query several web search APIs through a single, normalized interface. One call fans out to every configured engine in parallel, and you get back a consistent result shape per engine — including per-engine errors, warnings, rate-limit info, and optional raw payloads — without one slow or failing provider blocking the rest. Then merge everything into one deduplicated, rank-fused list, format it for an LLM prompt, or expose the whole thing as an agent tool or MCP server.
SearchResult, Answer, and metadata are normalized across engines.ok: true | false result; a failure in one never rejects the others.aggregate() dedupes by canonical URL and fuses rankings with reciprocal rank fusion.formatForLLM() renders results as a compact markdown or XML block for prompts.| Engine | id | Credentials env var (CLI) |
|---|---|---|
| Brave | brave | BRAVE_API_KEY |
| Ceramic | ceramic | CERAMIC_API_KEY |
| DuckDuckGo Instant Answers | duckduckgo | — (keyless) |
| Exa | exa | EXA_API_KEY |
| Firecrawl | firecrawl | FIRECRAWL_API_KEY |
| GDELT | gdelt | — (keyless) |
| Hacker News (Algolia) | hackernews | — (keyless) |
| Jina Search | jina | JINA_API_KEY |
| Kagi | kagi | KAGI_API_KEY |
| Linkup | linkup | LINKUP_API_KEY |
| Parallel | parallel | PARALLEL_API_KEY |
| SearXNG (self-hosted) | searxng | SEARXNG_BASE_URL (+ optional SEARXNG_API_KEY) |
| SerpAPI | serpapi | SERPAPI_API_KEY |
| Serper.dev | serper | SERPER_API_KEY |
| Perplexity Sonar | sonar | PERPLEXITY_API_KEY |
| Tavily | tavily | TAVILY_API_KEY |
| You.com | you | YOU_API_KEY |
Notes: duckduckgo hits the free Instant Answer API — encyclopedic abstracts and related topics, not full web results. searxng requires a self-hosted instance with the JSON output format enabled (search.formats: [html, json] in settings.yml). gdelt returns global news metadata only (title, URL, date, image) with a ~15 minute refresh — no snippets or page text; its publishedDate is GDELT's seendate, i.e. when GDELT first saw the article rather than the publisher's own date. GDELT's domain: filter performs suffix matching, so domain:example.com can also match a longer domain ending in example.com; this adapter intentionally preserves that fuzzy behavior rather than using the exact domainis: operator. hackernews queries the public Algolia index and defaults to tags=story; override tags for comments, ask_hn, show_hn, front_page, or author_<name>. linkup defaults to the cheaper searchResults output type; set defaults: { outputType: "sourcedAnswer" } for a cited answer, or defaults: { depth: "deep" } for its slower, broader crawl.
npm install agent-web-search zod
# or
pnpm add agent-web-search zod
zod (v4) is a peer dependency. Requires Node.js >= 22 for the CLI and Node builds; the library core also runs in browsers and edge runtimes (see Browser & edge usage).
import { search } from "agent-web-search";
const response = await search(
{ query: "best espresso machines 2026", count: 5 },
{
brave: { apiKey: process.env.BRAVE_API_KEY! },
exa: { apiKey: process.env.EXA_API_KEY! },
},
);
for (const [engine, result] of Object.entries(response)) {
if (result.ok) {
console.log(`${engine}: ${result.results.length} results`);
for (const item of result.results) {
console.log(` ${item.title} — ${item.url}`);
}
} else {
console.error(`${engine} failed: ${result.error.kind} — ${result.error.message}`);
}
}
search() returns a SearchResponse — a record keyed by engine id, where each value is either an EngineSuccess (ok: true) or an EngineFailure (ok: false).
Create a client once and reuse it for many searches. Configuration is validated up front.
import { createSearchClient } from "agent-web-search";
const client = createSearchClient({
brave: { apiKey: process.env.BRAVE_API_KEY! },
sonar: { apiKey: process.env.PERPLEXITY_API_KEY! },
});
const response = await client.search({ query: "what is a vector database" });
search() and searchStream() are one-shot convenience wrappers that build a client per call. Prefer createSearchClient when you issue more than one search — the cost budget and rate-limit state also live on the client.
Engine defaults are overridable: normalized query fields take precedence when supplied, and overrides take precedence over both. Domain filters must contain nonblank entries. Firecrawl's native includeDomains and excludeDomains defaults or overrides must be arrays; You also accepts comma-separated native GET filters. Invalid native defaults reject client creation, and invalid selected-engine overrides reject a search before any provider request. Native count values must be positive integers (numeric strings are rejected). Counts above a documented provider limit are clamped with a warning.
aggregate() merges a multi-engine response into a single result list. URLs are canonicalized for deduplication (protocol, www., fragments, trailing slashes, and tracking params like utm_*/gclid/fbclid are ignored) and ordered by reciprocal rank fusion: each engine contributes weight / (k + rank), so results that several engines agree on rise to the top.
import { aggregate } from "agent-web-search";
const merged = aggregate(response, {
k: 60, // RRF smoothing constant (default 60)
weights: { exa: 2 }, // trust some engines more
maxResults: 10,
});
for (const result of merged.results) {
// result.engines — which engines returned it
// result.engineRank — its 1-based rank per engine
// result.fusedScore — the RRF score used for ordering
console.log(result.fusedScore.toFixed(4), result.engines, result.url);
}
merged.answers; // Record<engine, Answer> from answer engines (sonar, tavily, …)
merged.succeeded; // engines that returned ok
merged.failed; // Record<engine, SearchEngineError>
formatForLLM() turns a response (raw or pre-aggregated) into a compact, citation-friendly block to drop into a prompt.
import { formatForLLM } from "agent-web-search";
const block = formatForLLM(response, {
format: "markdown", // or "xml"
maxResults: 8,
maxSnippetChars: 400,
});
Markdown output has an ## Answers section (when engines produced answers) and a numbered ## Search results list with title, date, URL, snippet, and source engines. XML output emits <search_results> with <answer> and <result> elements, fully escaped.
Ready-made web-search tools for the common LLM SDK wire formats — validation via Zod, execution via your configured client, output via formatForLLM.
import {
aiSdkWebSearchTool,
anthropicWebSearchTool,
createSearchClient,
openaiWebSearchTool,
} from "agent-web-search";
// or: import { ... } from "agent-web-search/tools";
const client = createSearchClient({ brave: { apiKey: "..." } });
// Anthropic API
const tool = anthropicWebSearchTool(client);
// tools: [{ name: tool.name, description: tool.description, input_schema: tool.input_schema }]
// on tool_use: const text = await tool.execute(toolUse.input);
// OpenAI function calling
const { definition, execute } = openaiWebSearchTool(client);
// Vercel AI SDK
// tools: { web_search: aiSdkWebSearchTool(client) }
All variants accept { name, description, format } options. The generic createWebSearchTool(client) exposes the Zod schema, the JSON schema, and execute for anything else.
Run the library as a zero-dependency MCP stdio server exposing a web_search tool:
TAVILY_API_KEY=... agent-web-search mcp
# restrict engines:
BRAVE_API_KEY=... agent-web-search mcp --engine brave
For Claude Code: claude mcp add web-search -e TAVILY_API_KEY=... -- npx agent-web-search mcp.
Programmatic (Node-only) usage via the agent-web-search/mcp subpath:
import { runMcpServer } from "agent-web-search/mcp";
await runMcpServer(client, { serverVersion: "1.0.0" });
By default every configured engine is queried in parallel ("all"). Three more strategies are available per client or per request:
// First success wins; everything else is aborted.
await client.search({ query }, { strategy: "race" });
// Try engines sequentially in priority order, stop at the first success.
await client.search({ query }, { strategy: "fallback", order: ["brave", "exa"] });
// Start engines staggered by hedgeDelayMs; first success aborts the rest.
await client.search({ query }, { strategy: "hedged", order: ["brave", "exa"], hedgeDelayMs: 300 });
// Overall deadline across all engines and retries (any strategy).
await client.search({ query }, { deadlineMs: 5000 });
With "fallback" and "hedged", engines that were never started are omitted from the response; with "race", aborted engines settle as failures and are included. searchStream always fans out to all engines but honors deadlineMs and order.
const client = createSearchClient(
{
brave: {
apiKey: "...",
throttle: { maxConcurrent: 2, minIntervalMs: 100 }, // client-side pacing
costPerRequestUsd: 0.005, // your cost estimate
},
},
{
respectRateLimits: true, // fail fast while a provider reports remaining: 0
budget: { maxCostUsd: 1 }, // hard ceiling across all searches on this client
},
);
throttle.maxConcurrent caps in-flight requests per engine; minIntervalMs spaces request starts.respectRateLimits: true, an engine whose last response reported an exhausted rate limit fails fast with a rate_limit error until the provider-reported reset time, instead of burning a request.usage.costUsd, e.g. Exa) or your costPerRequestUsd estimate; once reached, engines fail fast with a quota error.searchStream yields events as each engine produces them. Engines that support native streaming (e.g. Sonar) emit answer_delta events; non-streaming engines emit their terminal events when they complete.
import { searchStream } from "agent-web-search";
const stream = searchStream(
{ query: "summarize the latest in fusion energy" },
{ sonar: { apiKey: process.env.PERPLEXITY_API_KEY! } },
);
for await (const event of stream) {
switch (event.type) {
case "answer_delta":
process.stdout.write(event.text);
break;
case "answer_done":
console.log("\ncitations:", event.answer.citations.length);
break;
case "results":
console.log(`[${event.engine}] ${event.results.length} results`);
break;
case "metadata":
case "error":
case "done":
break;
}
}
Observe requests, responses, retries, errors, and settlements without affecting results. Hooks can be set per client, per request, or per engine.
const client = createSearchClient(engines, {
hooks: {
onRequest: ({ engine, url, attempt }) => log(`→ ${engine} ${url} (try ${attempt})`),
onResponse: ({ engine, status, latencyMs }) => log(`← ${engine} ${status} ${latencyMs}ms`),
onRetry: ({ engine, delayMs }) => log(`↻ ${engine} retrying in ${delayMs}ms`),
onError: ({ engine, error }) => log(`✗ ${engine} ${error.kind}`),
onSettled: ({ engine, result }) => log(`✓ ${engine} ok=${result.ok}`),
},
});
Pass an AbortSignal to cancel in-flight requests (and the stream).
const controller = new AbortController();
const promise = client.search({ query: "..." }, { signal: controller.signal });
controller.abort();
Implement an EngineAdapter and register it via options.adapters. defineEngine is an identity helper that preserves config types.
import { createSearchClient, defineEngine, KeyedEngineConfigSchema } from "agent-web-search";
const myAdapter = defineEngine({
id: "my-engine",
capabilities: { /* ... */ },
configSchema: KeyedEngineConfigSchema,
buildRequest(query, config, warnings) {
return { method: "GET", url: "https://api.example.com/search", query: { q: query.query } };
},
parseResponse(res, ctx) {
/* return an EngineResult */
},
});
const client = createSearchClient(
{ "my-engine": { apiKey: "..." } },
{ adapters: [myAdapter] },
);
You can also import the built-in adapters individually:
import { braveAdapter } from "agent-web-search/adapters/brave";
import { exaAdapter, tavilyAdapter } from "agent-web-search/adapters";
See CONTRIBUTING.md for the full checklist to add a built-in adapter, and the examples directory for runnable samples.
| Engine | answer | content | streaming | multi-query | count | dateRange | freshness | includeDomains | excludeDomains | country | language | safeSearch | verticals |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| brave | — | — | — | — | ✓ | ✓ | ✓ | emulated | emulated | ✓ | ✓ | ✓ | web |
| ceramic | — | — | — | — | — | — | — | — | — | — | — | — | web |
| duckduckgo | ✓ | — | — | — | — | — | — | — | — | — | — | — | web |
| exa | — | ✓ | — | — | ✓ | ✓ | ✓ | native | native | ✓ | — | — | web, news |
| firecrawl | — | ✓ | — | — | ✓ | ✓ | ✓ | native | native | ✓ | — | — | web, news, images |
| gdelt | — | — | — | — | ✓ | ✓ | ✓ | emulated | emulated | — | — | — | news |
| hackernews | — | — | — | — | ✓ | ✓ | ✓ | — | — | — | — | — | web |
| jina | — | ✓ | — | — | ✓ | — | — | emulated | emulated | ✓ | ✓ | — | web |
| kagi | — | — | — | — | ✓ | — | — | — | — | — | — | — | web, news |
| linkup | — | ✓ | — | — | ✓ | ✓ | ✓ | native | native | — | — | — | web |
| parallel | — | — | — | ✓ | ✓ | ✓ | ✓ | native | native | ✓ | — | — | web |
| searxng | ✓ | — | — | — | ✓ | — | ✓ | emulated | emulated | — | ✓ | ✓ | web, news, images, video |
| serpapi | — | — | — | — | ✓ | ✓ | ✓ | emulated | emulated | ✓ | ✓ | ✓ | web, news |
| serper | — | — | — | — | ✓ | ✓ | ✓ | emulated | emulated | ✓ | ✓ | — | web, news |
| sonar | ✓ | — | ✓ | — | — | ✓ | ✓ | native | emulated | — | ✓ | — | web |
| tavily | ✓ | ✓ | — | — | ✓ | ✓ | ✓ | native | native | — | — | — | web, news |
| you | — | ✓ | — | — | ✓ | ✓ | ✓ | native | native | ✓ | ✓ | ✓ | web, news |
native = handled by the provider's API; emulated = approximated by the adapter (e.g. via query operators). Serper and DuckDuckGo also surface opportunistic answers (Google answer box / instant answer) when the provider returns one, even where answer is not a guaranteed capability; so does Linkup when configured with outputType: "sourcedAnswer". GDELT emulates domain filters with its suffix-matching domain: operator rather than site: or exact domainis:, and exposes language/country only through GDELT-specific codes — reach those with overrides.
The package ships an agent-web-search binary. Set the relevant API key env vars, then:
agent-web-search --query "best espresso machines" --engine brave --engine exa --count 5
# merged + LLM-ready markdown
agent-web-search -q "fusion energy news" --format markdown
# aggregated JSON, racing engines with a 5s deadline
agent-web-search -q "vector databases" --aggregate --strategy race --deadline-ms 5000
# MCP stdio server
agent-web-search mcp
By default it queries every engine that has a matching API key set in the environment (the keyless duckduckgo, gdelt, and hackernews engines join only when named with --engine; searxng joins when SEARXNG_BASE_URL is set).
Options
-q, --query <text> Search query. Positional text is also accepted.
-e, --engine <id> Engine id. Repeat or comma-separate.
--count <number> Desired result count per engine.
--freshness <range> day, week, month, or year.
--country <code> ISO country code.
--language <code> ISO language code.
--safe-search <mode> off, moderate, or strict.
--include-domain <host> Domain allowlist. Repeat or comma-separate.
--exclude-domain <host> Domain blocklist. Repeat or comma-separate.
--content <fields> true or comma list: markdown,html,text,summary.
--raw Include top-level raw provider payloads.
--stream Emit stream events as NDJSON.
--format <name> json (default), ndjson, markdown, or xml.
--ndjson Shorthand for --format ndjson.
--aggregate Merge engines into one deduplicated,
rank-fused list (json/ndjson output).
--strategy <name> all (default), race, fallback, or hedged.
--deadline-ms <number> Overall deadline across engines and retries.
-h, --help Show help.
-v, --version Show version.
type SearchResponse = Record<string, EngineResult>;
type EngineResult = EngineSuccess | EngineFailure;
A successful engine result (ok: true) contains results: SearchResult[], an optional answer, and metadata. A normalized SearchResult includes url, title, snippet/snippets, publishedDate, author, score, source, optional content/highlights/image/favicon, and the provider's raw payload.
A failed engine result (ok: false) contains an error whose kind is one of: auth, rate_limit, quota, bad_request, unsupported, timeout, network, upstream, or parse, along with metadata.
Both Zod schemas (SearchResponseSchema, SearchResultSchema, AnswerSchema, …) and TypeScript types are exported from the package root. API documentation is generated with TypeDoc and published from the docs workflow.
Everything importable from the package root (client, adapters, aggregation, formatting, tools) is browser-safe: no Node builtins, fetch-based transport, works in browsers, Cloudflare Workers, Deno, and Bun. CI enforces this with an esbuild browser-platform bundle check (pnpm check:browser). The CLI and the agent-web-search/mcp subpath are Node-only.
Caveat: most search providers do not send CORS headers and your API keys should not ship to untrusted clients — in real browser apps, proxy provider calls through your backend (baseUrl is configurable per engine, and you can pass a custom fetch). The keyless GDELT and Hacker News endpoints are exceptions: both are CORS-enabled and can be called directly from browsers.
pnpm install
pnpm build # emit ESM + CJS builds to dist/
pnpm test # run the vitest suite with coverage
pnpm typecheck # tsc --noEmit
pnpm lint # biome check
pnpm format # biome check --write
pnpm check:browser # browser-platform bundle smoke test
pnpm docs # generate TypeDoc API docs
Requires Node.js >= 22. This repo uses pnpm (pinned via packageManager), Vitest for tests (including per-adapter contract fixtures in test/fixtures/, property-based tests with fast-check, and an env-gated live suite), and Biome for linting and formatting. Versions and changelogs are prepared manually with Changesets. A scheduled live workflow runs the real-API integration tests weekly to catch provider drift.
To release, run pnpm changeset version, commit and merge the version and changelog changes to main, then publish a GitHub Release tagged v<package.json version>. The release workflow publishes the tagged commit to npm exclusively through trusted publishing (GitHub OIDC), with provenance. Stable releases use latest; prereleases use next. See the release checklist for npm setup and the full release process.
MIT © Harold Martin