AndreaBozzo/Ares

Rust-native, local-first, auditable structured extraction with schemas, benchmarks, persistence, and operational reliability

Rust

5

99 commits

updated Aug 9, 2026

See the code
anthropic-api
claude-code-skills
docker
json-schema
llm
ollama
openai
qwen
rust
structured-data
web-scraper

README

Ares

Ares

Web scraper with LLM-powered structured data extraction — cloud or fully local.

crates.io CI Discord


Ares fetches web pages, converts HTML to Markdown, and uses an LLM to extract structured data defined by JSON Schemas. It ships a CLI and a REST API, supports persistent job queues with retries, circuit breaking, rate-limiting, change detection, and recursive crawling.

Works with any LLM backend — OpenAI, Gemini, Anthropic (Claude), Ollama, llama.cpp, LM Studio, or a fully embedded Qwen model that runs inside the Ares process with no network connection required.

Named after the Greek god of war and courage.

Conceptual sibling of Ceres — same philosophy, different temperament. Where Ceres is the nurturing goddess of harvest, Ares charges headfirst into the web and takes what it needs.

💡 Claude Code user? Install the Ares Claude Skill to give Claude deep knowledge of Ares — architecture, traits, CLI, REST API, schemas, and extension patterns.

What's new in v0.4.0

  • 🖥️ Native local inference — run extraction entirely offline with the embedded Qwen2.5-3B model (no API key, no server, no network after first download) via the local-llm feature
  • 🦙 Ollama & llama.cpp support — point ARES_BASE_URL at any OpenAI-compatible local server (Ollama, llama.cpp, LM Studio) and extract with zero code changes
  • 🤖 Anthropic (Claude) provider — native Messages API with forced tool use for structured extraction, behind the anthropic feature flag
  • Output validation — every extraction is validated against your JSON Schema before it is saved; mismatches are surfaced as errors rather than silently stored
  • 📊 Run metadata — provider, schema version, latency, and token counts are now recorded for every extraction

Architecture

flowchart TB
  %% External Entities
  User((User / Cron))
  Admin((API Consumer))
  Web[("Target Websites")]
  LLM_API[("LLM API\n(OpenAI / Gemini)")]

  %% Entrypoints
  subgraph Interfaces["Interfaces"]
    CLI["ares-cli\n(Command Line)"]
    API["ares-api\n(REST / Axum / Swagger)"]
  end

  %% Core Business Logic
  subgraph Core["ares-core (Business Logic)"]
    Traits{{"Traits\n(Fetcher · Cleaner · Extractor\nExtractorFactory · ExtractionStore · JobQueue)"}}
    Schema["SchemaResolver\n(CRUD · name@version · registry)"]
    ScrapeSvc["ScrapeService\n(Fetcher → Cleaner → Extractor → Store)"]
    WorkerSvc["WorkerService\n(Poll queue · retry · shutdown)"]
    CB["CircuitBreaker"]
    Throttle["ThrottledFetcher\n(Per-domain rate limit)"]
    Cache["ContentCache · ExtractionCache\n(In-memory / moka)"]
    Crawl["CrawlConfig\n(Depth · Pages · Domains · Robots)"]
    NullStore["NullStore\n(No-op persistence)"]

    WorkerSvc -->|Creates per job| ScrapeSvc
    WorkerSvc -->|Guards scrape calls| CB
    ScrapeSvc -->|Optional| Cache
    WorkerSvc -->|Spawns child jobs| Crawl
  end

  %% External Adapters
  subgraph Client["ares-client (External Adapters)"]
    Reqwest["ReqwestFetcher\n(Static HTML)"]
    Browser["BrowserFetcher\n(Chromium SPA)"]
    Cleaner["HtmdCleaner\n(HTML → Markdown)"]
    LlmClient["OpenAiExtractor\n(JSON Schema extraction)"]
    Factory["OpenAiExtractorFactory\n(Creates extractors per job)"]
    LinkDisc["HtmlLinkDiscoverer\n(Anchor tag extraction)"]
    Robots["CachedRobotsChecker\n(Per-domain robots.txt)"]
  end

  %% Database
  subgraph Database["ares-db (Persistence)"]
    DB[(PostgreSQL)]
    JobRepo["ScrapeJobRepository\n(implements JobQueue)"]
    ExtRepo["ExtractionRepository\n(implements ExtractionStore)"]

    JobRepo --> DB
    ExtRepo --> DB
  end

  %% User → Interface
  User -->|Executes| CLI
  Admin -->|HTTP| API

  %% CLI wiring
  CLI -->|One-shot scrape| ScrapeSvc
  CLI -->|Start worker| WorkerSvc
  CLI -->|Resolve schemas| Schema

  %% API wiring (no WorkerService — worker is a separate process)
  API -->|One-shot scrape| ScrapeSvc
  API -->|Manage jobs| JobRepo
  API -->|Schema CRUD| Schema

  %% Trait implementations (dashed = "implements")
  Traits -.->|Implemented by| Client
  Traits -.->|Implemented by| Database
  NullStore -.->|Implements ExtractionStore| Traits

  %% External interactions
  Reqwest -->|HTTP fetch| Web
  Browser -->|Headless render| Web
  LlmClient -->|Structured extraction| LLM_API
ares-cli          CLI interface — arg parsing, wiring, output formatting, delegation
ares-api          REST API — Axum HTTP server, OpenAPI/Swagger UI, Bearer auth
ares-core         Business logic — ScrapeService, WorkerService, CircuitBreaker, CrawlConfig, ContentCache, ExtractionCache, SchemaResolver, traits
ares-client       External adapters — ReqwestFetcher, BrowserFetcher, HtmdCleaner, OpenAiExtractor, HtmlLinkDiscoverer, CachedRobotsChecker
ares-db           PostgreSQL persistence — ExtractionRepository, ScrapeJobRepository, migrations

All external dependencies are behind traits (Fetcher, Cleaner, Extractor, ExtractionStore, ExtractorFactory, JobQueue), enabling full mock-based testing. The Fetcher trait has two implementations: ReqwestFetcher for static pages and BrowserFetcher (feature-gated behind browser) for JS-rendered SPAs.

Prerequisites

  • Rust 1.88+ (edition 2024)
  • Docker (for PostgreSQL and integration tests)
  • An LLM backend — one of:
    • An API key for OpenAI, Gemini, or Anthropic
    • A local server such as Ollama or llama.cpp (no key needed)
    • Nothing — build with --features local-llm to embed Qwen2.5-3B directly in Ares
  • Chromium / Chrome (only when using --browser for JS-rendered pages)

Quick Start

# Clone and build
git clone <repo-url> && cd Ares
cargo build

# Start PostgreSQL + pgAdmin
docker compose up -d

# Configure environment
cp .env.example .env
# Edit .env with your API key. The default DATABASE_URL already matches
# the compose `db` service (ares_user / password / ares_db).

# One-shot scrape (stdout only)
cargo run -- scrape -u https://example.com -s schemas/blog/1.0.0.json

# Scrape with Ollama (no API key required)
ollama pull qwen2.5:3b && ollama serve &
ARES_BASE_URL=http://localhost:11434/v1 ARES_MODEL=qwen2.5:3b ARES_API_KEY=sk-local \
  cargo run -- scrape -u https://example.com -s blog@latest

# Scrape fully offline — embedded Qwen2.5-3B, no network after download
cargo run --features local-llm -- model pull qwen2.5-3b-instruct-q4
cargo run --features local-llm -- scrape --provider local --model qwen2.5-3b-instruct-q4 \
  -u https://example.com -s blog@latest

# Scrape a JS-rendered page with headless browser
cargo run --features browser -- scrape -u https://spa-example.com -s blog@latest --browser

# Scrape and persist to database
cargo run -- scrape -u https://example.com -s blog@latest --save

# View extraction history
cargo run -- history -u https://example.com -s blog

# Create a background job
cargo run -- job create -u https://example.com -s blog@latest

# Start a worker to process jobs
cargo run -- worker

CLI Commands

ares scrape

One-shot extraction. Fetches the URL, cleans HTML to Markdown, sends it to the LLM with the JSON Schema, and prints the extracted data to stdout.

FlagEnv VarDescription
-u, --urlTarget URL
-s, --schemaSchema path or name@version
-m, --modelARES_MODELLLM model (e.g., gpt-4o-mini, claude-haiku-4-5)
--providerARES_PROVIDERopenai (default) or anthropic (requires the anthropic feature)
-b, --base-urlARES_BASE_URLAPI base URL (defaults to the selected provider's endpoint)
-a, --api-keyARES_API_KEYAPI key
--savePersist result to database
--schema-nameOverride schema name for storage
--browserUse headless browser for JS-rendered pages (requires browser feature)
--fetch-timeoutHTTP fetch timeout in seconds (default: 30)
--llm-timeoutLLM API timeout in seconds (default: 120)
--system-promptCustom system prompt for LLM extraction
--skip-unchangedSkip saving when extracted data hasn't changed (requires --save)
--throttlePer-domain throttle delay in milliseconds (e.g., 1000 for 1s between requests)
--no-cacheDisable in-memory caching (content + extraction)
--cache-ttlARES_CACHE_TTLCache TTL in seconds (default: 3600)
--formatOutput format: json, jsonl, csv, table, jq (default: json)

ares history

Show extraction history for a URL + schema pair, with change detection.

FlagEnv VarDescription
-u, --urlTarget URL
-s, --schema-nameSchema name to filter by
-l, --limitNumber of results (default: 10)
--formatOutput format: json, jsonl, csv, table, jq (default: json)

ares job create|list|show|cancel

Manage persistent scrape jobs in the PostgreSQL queue.

ares worker

Start a background worker that polls the job queue, processes scrape jobs through the circuit breaker, handles retries with exponential backoff, and supports graceful shutdown via Ctrl+C.

FlagEnv VarDescription
--worker-idCustom worker ID (auto-generated if omitted)
--poll-intervalSeconds between job queue polls (default: 5)
-a, --api-keyARES_API_KEYAPI key
--providerARES_PROVIDERopenai (default) or anthropic (requires the anthropic feature)
--browserUse headless browser for JS-rendered pages (requires browser feature)
--fetch-timeoutHTTP fetch timeout in seconds (default: 30)
--llm-timeoutLLM API timeout in seconds (default: 120)
--system-promptCustom system prompt for LLM extraction
--skip-unchangedSkip saving when extracted data hasn't changed
--throttlePer-domain throttle delay in milliseconds
--no-cacheDisable in-memory caching
--cache-ttlARES_CACHE_TTLCache TTL in seconds (default: 3600)

ares crawl start|status|results

Recursive web crawling with link discovery and robots.txt compliance. The seed URL is fetched, links are discovered, and child jobs are created in the queue for the worker to process.

FlagDescription
-u, --urlSeed URL to start crawling from
-s, --schemaSchema path or name@version
-d, --max-depthMaximum crawl depth (default: 1)
-m, --modelLLM model
-b, --base-urlAPI base URL
--max-pagesMaximum number of pages to crawl (default: 100)
--allowed-domainsComma-separated allowed domains (defaults to seed URL domain)
--schema-nameOverride schema name
# Start a crawl (creates seed job + discovers links)
ares crawl start -u https://example.com -s blog@latest --max-depth 2 --max-pages 10

# Check progress (requires a worker running in another terminal)
ares crawl status <SESSION_ID>

# View extracted data from all crawled pages
ares crawl results <SESSION_ID>

ares schema validate

Validate a JSON Schema file against the JSON Schema specification.

ares schema validate schemas/blog/1.0.0.json

REST API

Ares ships a standalone HTTP server (ares-api) built on Axum with auto-generated OpenAPI documentation.

Running the server

# Run locally
cargo run --bin ares-api

# Or with Docker
docker build -t ares-api:latest .
docker run -p 3000:3000 --env-file .env ares-api:latest

Once running, interactive API docs are available at /swagger-ui.

Endpoints

MethodPathAuthDescription
POST/v1/scrapeBearerOne-shot scrape and extract
POST/v1/jobsBearerCreate a scrape job
GET/v1/jobsBearerList jobs (filter by status, limit)
GET/v1/jobs/{id}BearerGet job details
DELETE/v1/jobs/{id}BearerCancel a pending job
GET/v1/extractionsBearerQuery extraction history
GET/v1/schemasBearerList all schemas
GET/v1/schemas/{name}/{version}BearerGet schema definition
POST/v1/schemasBearerCreate/upload a schema version
PUT/v1/schemas/{name}/{version}BearerUpdate a schema version
DELETE/v1/schemas/{name}/{version}BearerDelete a schema version
POST/v1/jobs/{id}/retryBearerRetry a failed/cancelled job
POST/v1/crawlBearerStart a crawl session
GET/v1/crawl/{id}BearerGet crawl session status
GET/v1/crawl/{id}/resultsBearerGet crawl session results
GET/healthHealth check (database connectivity)

Authentication

Protected endpoints require a Bearer token set via ARES_ADMIN_TOKEN. Token comparison uses constant-time equality (subtle crate) to prevent timing attacks.

curl -H "Authorization: Bearer $ARES_ADMIN_TOKEN" http://localhost:3000/v1/jobs

If ARES_ADMIN_TOKEN is not set, all protected endpoints return 403 Forbidden.

Schemas

Schemas are versioned JSON Schema files stored in schemas/:

schemas/
  registry.json
  blog/1.0.0.json            # Blog posts and articles
  github_repo/1.0.0.json     # GitHub repository pages
  product/1.0.0.json         # E-commerce product pages
  news_article/1.0.0.json    # News articles
  job_listing/1.0.0.json     # Job board listings
  recipe/1.0.0.json          # Recipe pages
  event/1.0.0.json           # Event listings
  dataset/1.0.0.json         # Open data portal datasets

Reference by path (schemas/blog/1.0.0.json) or by name (blog@1.0.0, blog@latest). Validate with ares schema validate <path>.

Configuration

VariableRequiredDefaultDescription
ARES_API_KEYYes (cloud)LLM API key (not needed with --provider local)
ARES_MODELYesLLM model name
ARES_PROVIDERNoopenaiLLM provider: openai, anthropic, or local
ARES_BASE_URLNoprovider defaultAPI base URL — set this to point at any local server
DATABASE_URLFor persistencePostgreSQL connection string
DATABASE_MAX_CONNECTIONSNo5PostgreSQL connection pool size
ARES_ADMIN_TOKENNo****** for REST API auth
ARES_SERVER_PORTNo3000HTTP server listen port
ARES_SCHEMAS_DIRNoschemasPath to schemas directory
ARES_CORS_ORIGINNoAllowed CORS origins (comma-separated, or *)
ARES_RATE_LIMIT_BURSTNo30Max burst requests per IP
ARES_RATE_LIMIT_RPSNo1Request replenish rate (per second)
ARES_BODY_SIZE_LIMITNo2097152Max request body size in bytes (2 MB)
ARES_CACHE_TTLNo3600In-memory cache TTL in seconds
ARES_MODEL_DIRNoplatform cacheDirectory where native models are stored
CHROME_BINNoAuto-detectedOverride path to Chrome/Chromium binary

Local inference — Ollama or llama.cpp (no API key needed)

Point ARES_BASE_URL at any OpenAI-compatible server and you are done — no rebuild, no feature flags required. Ares sends response_format: json_schema with your schema, and every extraction is validated against it regardless of backend.

Ollama:

ollama pull qwen2.5:3b
ollama serve   # exposes http://localhost:11434/v1

export ARES_PROVIDER=openai
export ARES_BASE_URL=http://localhost:11434/v1
export ARES_MODEL=qwen2.5:3b           # must match the exact Ollama tag
export ARES_API_KEY=sk-local           # ignored by Ollama, but required by Ares
cargo run -- scrape -u https://example.com -s blog@latest

Note: Ollama's OpenAI-compatibility layer supports json_object mode but not full json_schema. Ares validates the output anyway, so malformed extractions are caught and surfaced as errors.

llama.cpp (llama-server):

llama-server -hf Qwen/Qwen2.5-3B-Instruct-GGUF:Q4_K_M \
  --port 8080 --alias qwen2.5-3b-instruct --ctx-size 8192 --temp 0

export ARES_BASE_URL=http://localhost:8080/v1
export ARES_MODEL=qwen2.5-3b-instruct
export ARES_API_KEY=sk-local
cargo run -- scrape -u https://example.com -s blog@latest

See docs/local-inference.md for more server options (LM Studio, etc.) and the bench harness that compares local vs hosted on output validity, latency, and cost.

Native local inference — embedded Qwen (fully offline)

Build Ares with the local-llm feature to embed a Qwen2.5-3B-Instruct Q4 model that runs directly inside the Ares process using Candle. No server, no API key, no network connection needed after the one-time model download (~2 GB).

# Download the model once
cargo run --features local-llm -- model pull qwen2.5-3b-instruct-q4

# Scrape with the embedded model
cargo run --features local-llm -- scrape \
  --provider local --model qwen2.5-3b-instruct-q4 \
  --url https://example.com --schema blog@latest

# Manage downloaded models
cargo run --features local-llm -- model list
cargo run --features local-llm -- model remove qwen2.5-3b-instruct-q4

The model is stored in the platform cache directory (or ARES_MODEL_DIR when set). Native generation is CPU-based and serialized per process — treat it as a one-at-a-time extractor rather than a parallel worker.

Gemini

Gemini works via its OpenAI-compatible endpoint — no special feature flag needed:

export ARES_BASE_URL="https://generativelanguage.googleapis.com/v1beta/openai"
export ARES_MODEL="gemini-2.5-flash"

Anthropic (Claude)

Anthropic's API is not OpenAI-compatible (it uses the native Messages API), so it lives behind the anthropic build feature and the anthropic provider. Build with the feature, then select the provider:

export ARES_PROVIDER="anthropic"
export ARES_API_KEY="sk-ant-..."
export ARES_MODEL="claude-haiku-4-5"   # or claude-sonnet-4-6 for complex schemas

# ARES_BASE_URL defaults to https://api.anthropic.com/v1 for the anthropic provider
cargo run --features anthropic -- scrape -u https://example.com -s blog@latest
ModelBest for
claude-haiku-4-5Fast, cheap, high-volume extraction of simple schemas
claude-sonnet-4-6Complex schemas and nuanced content (higher quality, higher cost)

Extraction uses forced tool use under the hood: the JSON Schema is passed as the tool's input_schema and Claude is required to call it, so the result is a structured object that is then validated against the schema like any other provider. When running a worker with --provider anthropic, make sure jobs target an Anthropic base URL (the per-job default is the OpenAI endpoint).

Docker

# Build the image
docker build -t ares-api:latest .

# Start the dependencies with docker compose (PostgreSQL + pgAdmin)
docker compose up -d

docker compose up -d starts PostgreSQL (port 5432) and pgAdmin (port 5050). The application server service is included but commented out in compose.yml — uncomment it to run ares-api in the same stack, or run the built image directly with docker run -p 3000:3000 --env-file .env ares-api:latest.

The Dockerfile uses a multi-stage build (Rust builder → Debian slim runtime) with Chromium pre-installed for browser-based scraping. The release binary is compiled with LTO and symbol stripping for minimal image size.

Development

# Format, lint, and test
make all

# Run unit tests only
make test-unit

# Run integration tests (requires Docker)
make test-integration

# Run database migrations
make migrate

# Start/stop PostgreSQL
make docker-up
make docker-down

CI runs on every push and PR via GitHub Actions: formatting, Clippy, unit tests, integration tests (with a Postgres service container), and a cargo-deny security audit.

License

Apache-2.0

Contributors

AndreaBozzo

89 commits

Copilot

1 commits

AndreaBozzo/Ares

Rust-native, local-first, auditable structured extraction with schemas, benchmarks, persistence, and operational reliability

Rust

5

99 commits

updated Aug 9, 2026

See the code
anthropic-api
claude-code-skills
docker
json-schema
llm
ollama
openai
qwen
rust
structured-data
web-scraper

README

Ares

Ares

Web scraper with LLM-powered structured data extraction — cloud or fully local.

crates.io CI Discord


Ares fetches web pages, converts HTML to Markdown, and uses an LLM to extract structured data defined by JSON Schemas. It ships a CLI and a REST API, supports persistent job queues with retries, circuit breaking, rate-limiting, change detection, and recursive crawling.

Works with any LLM backend — OpenAI, Gemini, Anthropic (Claude), Ollama, llama.cpp, LM Studio, or a fully embedded Qwen model that runs inside the Ares process with no network connection required.

Named after the Greek god of war and courage.

Conceptual sibling of Ceres — same philosophy, different temperament. Where Ceres is the nurturing goddess of harvest, Ares charges headfirst into the web and takes what it needs.

💡 Claude Code user? Install the Ares Claude Skill to give Claude deep knowledge of Ares — architecture, traits, CLI, REST API, schemas, and extension patterns.

What's new in v0.4.0

  • 🖥️ Native local inference — run extraction entirely offline with the embedded Qwen2.5-3B model (no API key, no server, no network after first download) via the local-llm feature
  • 🦙 Ollama & llama.cpp support — point ARES_BASE_URL at any OpenAI-compatible local server (Ollama, llama.cpp, LM Studio) and extract with zero code changes
  • 🤖 Anthropic (Claude) provider — native Messages API with forced tool use for structured extraction, behind the anthropic feature flag
  • Output validation — every extraction is validated against your JSON Schema before it is saved; mismatches are surfaced as errors rather than silently stored
  • 📊 Run metadata — provider, schema version, latency, and token counts are now recorded for every extraction

Architecture

flowchart TB
  %% External Entities
  User((User / Cron))
  Admin((API Consumer))
  Web[("Target Websites")]
  LLM_API[("LLM API\n(OpenAI / Gemini)")]

  %% Entrypoints
  subgraph Interfaces["Interfaces"]
    CLI["ares-cli\n(Command Line)"]
    API["ares-api\n(REST / Axum / Swagger)"]
  end

  %% Core Business Logic
  subgraph Core["ares-core (Business Logic)"]
    Traits{{"Traits\n(Fetcher · Cleaner · Extractor\nExtractorFactory · ExtractionStore · JobQueue)"}}
    Schema["SchemaResolver\n(CRUD · name@version · registry)"]
    ScrapeSvc["ScrapeService\n(Fetcher → Cleaner → Extractor → Store)"]
    WorkerSvc["WorkerService\n(Poll queue · retry · shutdown)"]
    CB["CircuitBreaker"]
    Throttle["ThrottledFetcher\n(Per-domain rate limit)"]
    Cache["ContentCache · ExtractionCache\n(In-memory / moka)"]
    Crawl["CrawlConfig\n(Depth · Pages · Domains · Robots)"]
    NullStore["NullStore\n(No-op persistence)"]

    WorkerSvc -->|Creates per job| ScrapeSvc
    WorkerSvc -->|Guards scrape calls| CB
    ScrapeSvc -->|Optional| Cache
    WorkerSvc -->|Spawns child jobs| Crawl
  end

  %% External Adapters
  subgraph Client["ares-client (External Adapters)"]
    Reqwest["ReqwestFetcher\n(Static HTML)"]
    Browser["BrowserFetcher\n(Chromium SPA)"]
    Cleaner["HtmdCleaner\n(HTML → Markdown)"]
    LlmClient["OpenAiExtractor\n(JSON Schema extraction)"]
    Factory["OpenAiExtractorFactory\n(Creates extractors per job)"]
    LinkDisc["HtmlLinkDiscoverer\n(Anchor tag extraction)"]
    Robots["CachedRobotsChecker\n(Per-domain robots.txt)"]
  end

  %% Database
  subgraph Database["ares-db (Persistence)"]
    DB[(PostgreSQL)]
    JobRepo["ScrapeJobRepository\n(implements JobQueue)"]
    ExtRepo["ExtractionRepository\n(implements ExtractionStore)"]

    JobRepo --> DB
    ExtRepo --> DB
  end

  %% User → Interface
  User -->|Executes| CLI
  Admin -->|HTTP| API

  %% CLI wiring
  CLI -->|One-shot scrape| ScrapeSvc
  CLI -->|Start worker| WorkerSvc
  CLI -->|Resolve schemas| Schema

  %% API wiring (no WorkerService — worker is a separate process)
  API -->|One-shot scrape| ScrapeSvc
  API -->|Manage jobs| JobRepo
  API -->|Schema CRUD| Schema

  %% Trait implementations (dashed = "implements")
  Traits -.->|Implemented by| Client
  Traits -.->|Implemented by| Database
  NullStore -.->|Implements ExtractionStore| Traits

  %% External interactions
  Reqwest -->|HTTP fetch| Web
  Browser -->|Headless render| Web
  LlmClient -->|Structured extraction| LLM_API
ares-cli          CLI interface — arg parsing, wiring, output formatting, delegation
ares-api          REST API — Axum HTTP server, OpenAPI/Swagger UI, Bearer auth
ares-core         Business logic — ScrapeService, WorkerService, CircuitBreaker, CrawlConfig, ContentCache, ExtractionCache, SchemaResolver, traits
ares-client       External adapters — ReqwestFetcher, BrowserFetcher, HtmdCleaner, OpenAiExtractor, HtmlLinkDiscoverer, CachedRobotsChecker
ares-db           PostgreSQL persistence — ExtractionRepository, ScrapeJobRepository, migrations

All external dependencies are behind traits (Fetcher, Cleaner, Extractor, ExtractionStore, ExtractorFactory, JobQueue), enabling full mock-based testing. The Fetcher trait has two implementations: ReqwestFetcher for static pages and BrowserFetcher (feature-gated behind browser) for JS-rendered SPAs.

Prerequisites

  • Rust 1.88+ (edition 2024)
  • Docker (for PostgreSQL and integration tests)
  • An LLM backend — one of:
    • An API key for OpenAI, Gemini, or Anthropic
    • A local server such as Ollama or llama.cpp (no key needed)
    • Nothing — build with --features local-llm to embed Qwen2.5-3B directly in Ares
  • Chromium / Chrome (only when using --browser for JS-rendered pages)

Quick Start

# Clone and build
git clone <repo-url> && cd Ares
cargo build

# Start PostgreSQL + pgAdmin
docker compose up -d

# Configure environment
cp .env.example .env
# Edit .env with your API key. The default DATABASE_URL already matches
# the compose `db` service (ares_user / password / ares_db).

# One-shot scrape (stdout only)
cargo run -- scrape -u https://example.com -s schemas/blog/1.0.0.json

# Scrape with Ollama (no API key required)
ollama pull qwen2.5:3b && ollama serve &
ARES_BASE_URL=http://localhost:11434/v1 ARES_MODEL=qwen2.5:3b ARES_API_KEY=sk-local \
  cargo run -- scrape -u https://example.com -s blog@latest

# Scrape fully offline — embedded Qwen2.5-3B, no network after download
cargo run --features local-llm -- model pull qwen2.5-3b-instruct-q4
cargo run --features local-llm -- scrape --provider local --model qwen2.5-3b-instruct-q4 \
  -u https://example.com -s blog@latest

# Scrape a JS-rendered page with headless browser
cargo run --features browser -- scrape -u https://spa-example.com -s blog@latest --browser

# Scrape and persist to database
cargo run -- scrape -u https://example.com -s blog@latest --save

# View extraction history
cargo run -- history -u https://example.com -s blog

# Create a background job
cargo run -- job create -u https://example.com -s blog@latest

# Start a worker to process jobs
cargo run -- worker

CLI Commands

ares scrape

One-shot extraction. Fetches the URL, cleans HTML to Markdown, sends it to the LLM with the JSON Schema, and prints the extracted data to stdout.

FlagEnv VarDescription
-u, --urlTarget URL
-s, --schemaSchema path or name@version
-m, --modelARES_MODELLLM model (e.g., gpt-4o-mini, claude-haiku-4-5)
--providerARES_PROVIDERopenai (default) or anthropic (requires the anthropic feature)
-b, --base-urlARES_BASE_URLAPI base URL (defaults to the selected provider's endpoint)
-a, --api-keyARES_API_KEYAPI key
--savePersist result to database
--schema-nameOverride schema name for storage
--browserUse headless browser for JS-rendered pages (requires browser feature)
--fetch-timeoutHTTP fetch timeout in seconds (default: 30)
--llm-timeoutLLM API timeout in seconds (default: 120)
--system-promptCustom system prompt for LLM extraction
--skip-unchangedSkip saving when extracted data hasn't changed (requires --save)
--throttlePer-domain throttle delay in milliseconds (e.g., 1000 for 1s between requests)
--no-cacheDisable in-memory caching (content + extraction)
--cache-ttlARES_CACHE_TTLCache TTL in seconds (default: 3600)
--formatOutput format: json, jsonl, csv, table, jq (default: json)

ares history

Show extraction history for a URL + schema pair, with change detection.

FlagEnv VarDescription
-u, --urlTarget URL
-s, --schema-nameSchema name to filter by
-l, --limitNumber of results (default: 10)
--formatOutput format: json, jsonl, csv, table, jq (default: json)

ares job create|list|show|cancel

Manage persistent scrape jobs in the PostgreSQL queue.

ares worker

Start a background worker that polls the job queue, processes scrape jobs through the circuit breaker, handles retries with exponential backoff, and supports graceful shutdown via Ctrl+C.

FlagEnv VarDescription
--worker-idCustom worker ID (auto-generated if omitted)
--poll-intervalSeconds between job queue polls (default: 5)
-a, --api-keyARES_API_KEYAPI key
--providerARES_PROVIDERopenai (default) or anthropic (requires the anthropic feature)
--browserUse headless browser for JS-rendered pages (requires browser feature)
--fetch-timeoutHTTP fetch timeout in seconds (default: 30)
--llm-timeoutLLM API timeout in seconds (default: 120)
--system-promptCustom system prompt for LLM extraction
--skip-unchangedSkip saving when extracted data hasn't changed
--throttlePer-domain throttle delay in milliseconds
--no-cacheDisable in-memory caching
--cache-ttlARES_CACHE_TTLCache TTL in seconds (default: 3600)

ares crawl start|status|results

Recursive web crawling with link discovery and robots.txt compliance. The seed URL is fetched, links are discovered, and child jobs are created in the queue for the worker to process.

FlagDescription
-u, --urlSeed URL to start crawling from
-s, --schemaSchema path or name@version
-d, --max-depthMaximum crawl depth (default: 1)
-m, --modelLLM model
-b, --base-urlAPI base URL
--max-pagesMaximum number of pages to crawl (default: 100)
--allowed-domainsComma-separated allowed domains (defaults to seed URL domain)
--schema-nameOverride schema name
# Start a crawl (creates seed job + discovers links)
ares crawl start -u https://example.com -s blog@latest --max-depth 2 --max-pages 10

# Check progress (requires a worker running in another terminal)
ares crawl status <SESSION_ID>

# View extracted data from all crawled pages
ares crawl results <SESSION_ID>

ares schema validate

Validate a JSON Schema file against the JSON Schema specification.

ares schema validate schemas/blog/1.0.0.json

REST API

Ares ships a standalone HTTP server (ares-api) built on Axum with auto-generated OpenAPI documentation.

Running the server

# Run locally
cargo run --bin ares-api

# Or with Docker
docker build -t ares-api:latest .
docker run -p 3000:3000 --env-file .env ares-api:latest

Once running, interactive API docs are available at /swagger-ui.

Endpoints

MethodPathAuthDescription
POST/v1/scrapeBearerOne-shot scrape and extract
POST/v1/jobsBearerCreate a scrape job
GET/v1/jobsBearerList jobs (filter by status, limit)
GET/v1/jobs/{id}BearerGet job details
DELETE/v1/jobs/{id}BearerCancel a pending job
GET/v1/extractionsBearerQuery extraction history
GET/v1/schemasBearerList all schemas
GET/v1/schemas/{name}/{version}BearerGet schema definition
POST/v1/schemasBearerCreate/upload a schema version
PUT/v1/schemas/{name}/{version}BearerUpdate a schema version
DELETE/v1/schemas/{name}/{version}BearerDelete a schema version
POST/v1/jobs/{id}/retryBearerRetry a failed/cancelled job
POST/v1/crawlBearerStart a crawl session
GET/v1/crawl/{id}BearerGet crawl session status
GET/v1/crawl/{id}/resultsBearerGet crawl session results
GET/healthHealth check (database connectivity)

Authentication

Protected endpoints require a Bearer token set via ARES_ADMIN_TOKEN. Token comparison uses constant-time equality (subtle crate) to prevent timing attacks.

curl -H "Authorization: Bearer $ARES_ADMIN_TOKEN" http://localhost:3000/v1/jobs

If ARES_ADMIN_TOKEN is not set, all protected endpoints return 403 Forbidden.

Schemas

Schemas are versioned JSON Schema files stored in schemas/:

schemas/
  registry.json
  blog/1.0.0.json            # Blog posts and articles
  github_repo/1.0.0.json     # GitHub repository pages
  product/1.0.0.json         # E-commerce product pages
  news_article/1.0.0.json    # News articles
  job_listing/1.0.0.json     # Job board listings
  recipe/1.0.0.json          # Recipe pages
  event/1.0.0.json           # Event listings
  dataset/1.0.0.json         # Open data portal datasets

Reference by path (schemas/blog/1.0.0.json) or by name (blog@1.0.0, blog@latest). Validate with ares schema validate <path>.

Configuration

VariableRequiredDefaultDescription
ARES_API_KEYYes (cloud)LLM API key (not needed with --provider local)
ARES_MODELYesLLM model name
ARES_PROVIDERNoopenaiLLM provider: openai, anthropic, or local
ARES_BASE_URLNoprovider defaultAPI base URL — set this to point at any local server
DATABASE_URLFor persistencePostgreSQL connection string
DATABASE_MAX_CONNECTIONSNo5PostgreSQL connection pool size
ARES_ADMIN_TOKENNo****** for REST API auth
ARES_SERVER_PORTNo3000HTTP server listen port
ARES_SCHEMAS_DIRNoschemasPath to schemas directory
ARES_CORS_ORIGINNoAllowed CORS origins (comma-separated, or *)
ARES_RATE_LIMIT_BURSTNo30Max burst requests per IP
ARES_RATE_LIMIT_RPSNo1Request replenish rate (per second)
ARES_BODY_SIZE_LIMITNo2097152Max request body size in bytes (2 MB)
ARES_CACHE_TTLNo3600In-memory cache TTL in seconds
ARES_MODEL_DIRNoplatform cacheDirectory where native models are stored
CHROME_BINNoAuto-detectedOverride path to Chrome/Chromium binary

Local inference — Ollama or llama.cpp (no API key needed)

Point ARES_BASE_URL at any OpenAI-compatible server and you are done — no rebuild, no feature flags required. Ares sends response_format: json_schema with your schema, and every extraction is validated against it regardless of backend.

Ollama:

ollama pull qwen2.5:3b
ollama serve   # exposes http://localhost:11434/v1

export ARES_PROVIDER=openai
export ARES_BASE_URL=http://localhost:11434/v1
export ARES_MODEL=qwen2.5:3b           # must match the exact Ollama tag
export ARES_API_KEY=sk-local           # ignored by Ollama, but required by Ares
cargo run -- scrape -u https://example.com -s blog@latest

Note: Ollama's OpenAI-compatibility layer supports json_object mode but not full json_schema. Ares validates the output anyway, so malformed extractions are caught and surfaced as errors.

llama.cpp (llama-server):

llama-server -hf Qwen/Qwen2.5-3B-Instruct-GGUF:Q4_K_M \
  --port 8080 --alias qwen2.5-3b-instruct --ctx-size 8192 --temp 0

export ARES_BASE_URL=http://localhost:8080/v1
export ARES_MODEL=qwen2.5-3b-instruct
export ARES_API_KEY=sk-local
cargo run -- scrape -u https://example.com -s blog@latest

See docs/local-inference.md for more server options (LM Studio, etc.) and the bench harness that compares local vs hosted on output validity, latency, and cost.

Native local inference — embedded Qwen (fully offline)

Build Ares with the local-llm feature to embed a Qwen2.5-3B-Instruct Q4 model that runs directly inside the Ares process using Candle. No server, no API key, no network connection needed after the one-time model download (~2 GB).

# Download the model once
cargo run --features local-llm -- model pull qwen2.5-3b-instruct-q4

# Scrape with the embedded model
cargo run --features local-llm -- scrape \
  --provider local --model qwen2.5-3b-instruct-q4 \
  --url https://example.com --schema blog@latest

# Manage downloaded models
cargo run --features local-llm -- model list
cargo run --features local-llm -- model remove qwen2.5-3b-instruct-q4

The model is stored in the platform cache directory (or ARES_MODEL_DIR when set). Native generation is CPU-based and serialized per process — treat it as a one-at-a-time extractor rather than a parallel worker.

Gemini

Gemini works via its OpenAI-compatible endpoint — no special feature flag needed:

export ARES_BASE_URL="https://generativelanguage.googleapis.com/v1beta/openai"
export ARES_MODEL="gemini-2.5-flash"

Anthropic (Claude)

Anthropic's API is not OpenAI-compatible (it uses the native Messages API), so it lives behind the anthropic build feature and the anthropic provider. Build with the feature, then select the provider:

export ARES_PROVIDER="anthropic"
export ARES_API_KEY="sk-ant-..."
export ARES_MODEL="claude-haiku-4-5"   # or claude-sonnet-4-6 for complex schemas

# ARES_BASE_URL defaults to https://api.anthropic.com/v1 for the anthropic provider
cargo run --features anthropic -- scrape -u https://example.com -s blog@latest
ModelBest for
claude-haiku-4-5Fast, cheap, high-volume extraction of simple schemas
claude-sonnet-4-6Complex schemas and nuanced content (higher quality, higher cost)

Extraction uses forced tool use under the hood: the JSON Schema is passed as the tool's input_schema and Claude is required to call it, so the result is a structured object that is then validated against the schema like any other provider. When running a worker with --provider anthropic, make sure jobs target an Anthropic base URL (the per-job default is the OpenAI endpoint).

Docker

# Build the image
docker build -t ares-api:latest .

# Start the dependencies with docker compose (PostgreSQL + pgAdmin)
docker compose up -d

docker compose up -d starts PostgreSQL (port 5432) and pgAdmin (port 5050). The application server service is included but commented out in compose.yml — uncomment it to run ares-api in the same stack, or run the built image directly with docker run -p 3000:3000 --env-file .env ares-api:latest.

The Dockerfile uses a multi-stage build (Rust builder → Debian slim runtime) with Chromium pre-installed for browser-based scraping. The release binary is compiled with LTO and symbol stripping for minimal image size.

Development

# Format, lint, and test
make all

# Run unit tests only
make test-unit

# Run integration tests (requires Docker)
make test-integration

# Run database migrations
make migrate

# Start/stop PostgreSQL
make docker-up
make docker-down

CI runs on every push and PR via GitHub Actions: formatting, Clippy, unit tests, integration tests (with a Postgres service container), and a cargo-deny security audit.

License

Apache-2.0

Contributors

AndreaBozzo

89 commits

Copilot

1 commits

Languages

Rust

96.7%

HTML

2.8%