rayoplateado/jurl

curl that reads the page for you

Rust

1

40 commits

updated Oct 5, 2026

See the code

See what people are saying

SourceMessageScoreDate

Show HN: Jurl – curl that reads the page for you

5

Oct 5, 2026

README

Ink drawing of a man in a suit and leopard-print tie, pinching his fingers as if picking something out

jurl

curl that reads the page for you.
Tell it what you want. It picks it out of the page.


$ jurl -q "how do I install it on macOS?" github.com/BurntSushi/ripgrep
…
### Installation

…

If you're a macOS Homebrew or a Linuxbrew user, then you can install ripgrep from homebrew-core:

```
$ brew install ripgrep
```

If you're a MacPorts user, then you can install ripgrep from the official ports:

```
$ sudo port install ripgrep
```
$ jurl --find "a cathedral" en.wikipedia.org/wiki/Cologne
https://thumb.wikimedia.org/wikipedia/commons/thumb/2/27/Kdom.jpg/250px-Kdom.jpg
$ jurl --links -n 3 news.ycombinator.com
https://github.com/PowderworksCode/headstart
https://www.da.vidbuchanan.co.uk/blog/hacking-time.html
https://gamehistory.org/5k-magazines/

Why jurl

  • It picks. It doesn't write. jurl doesn't use a chatbot. It uses decision models: Jev reads the text and Clef looks at the images. Neither writes a word. They only say how likely each paragraph, link or image is to be what you want, and jurl prints the winners exactly as they appear on the page. You never get a summary that drifts or a URL that doesn't exist.
  • It's fast. Most pages take under a second, end to end.
  • It's cheap. About a thousand pages per dollar of API usage.

Install

brew install rayoplateado/tap/jurl                                                   # macOS, Linux
curl -LsSf https://github.com/rayoplateado/jurl/releases/latest/download/jurl-installer.sh | sh   # no Homebrew

Windows: powershell -ExecutionPolicy Bypass -c "irm https://github.com/rayoplateado/jurl/releases/latest/download/jurl-installer.ps1 | iex" · From source: cargo install --git https://github.com/rayoplateado/jurl

Then run it. The first time, jurl asks for a TypeSafe API key, checks it and saves it. That's the whole setup. Pages that need JavaScript just work too: jurl fetches a headless browser the first time one shows up.

To use --vision and --find, run jurl init and add a Cloudflare Workers AI token.

What you can do

Get the gist

$ jurl -n 3 blog.cloudflare.com/markdown-for-agents/
# Introducing Markdown for Agents

<https://blog.cloudflare.com/markdown-for-agents/> · article (1.00)

The way content and businesses are discovered online is changing rapidly. In the past, traffic originated from traditional search engines, and SEO determined who got found first. …

## Convert HTML to markdown, automatically

Cloudflare's network now supports real-time content conversion at the source, …

Here’s how it works. To fetch the markdown version of any page from a zone with Markdown for Agents enabled, the client needs to add the **Accept** negotiation header …

You get the title, what kind of page it is, and the blocks that carry it, in reading order. Navigation, cookie banners, sign-up prompts, author bios and footers are left out.

Ask one question

$ jurl -q "What is the Jevons paradox?" -n 2 en.wikipedia.org/wiki/William_Stanley_Jevons
…
Jevons received public recognition for his work on The Coal Question (1865), in which he called attention to the gradual exhaustion of Britain's coal supplies and also put forth the view that increases in energy production efficiency leads to more, not less, consumption. …

## Practical economics

In The Coal Question, Jevons covered a breadth of concepts on energy depletion …

Out of a page with about 240 blocks, you get the two paragraphs that answer it.

Just the code

$ jurl --code -n 2 github.com/BurntSushi/ripgrep
### Installation

```
$ brew install ripgrep
```

```
$ cargo install ripgrep
```
$ jurl --links -n 3 -q "official installation instructions" github.com/BurntSushi/ripgrep
https://www.macports.org/ports.php?by=name&substr=ripgrep
https://packages.gentoo.org/packages/sys-apps/ripgrep
https://chocolatey.org/packages/ripgrep

The ranking puts content first, ahead of login, share or privacy policy. Add -q to keep only the links about something specific.

The real images

--image picks the content images by their file name, alt text and caption, with no logos, icons or tracking pixels. --vision also has Clef look at the pixels, which matters when the alt text says nothing:

Image on a Cloudflare blog postAlt text--image--vision
Stacked area chartBLOG-3162 40.640.82
DiagramBLOG-3162 30.630.82
Author avatarWill Allen0.380.21
Company logoCloudflare0.090.08

Find a photo by description

$ jurl --find "a cathedral" -n 3 en.wikipedia.org/wiki/Cologne

The page has 73 images, and Clef looks at every one of them in parallel. It took 2.3 seconds:

ImageWhy it matchedp
1Kdom.jpgCologne Cathedral0.96
2Köln_um_1890.jpgThe 1890 skyline, the cathedral towering over it0.95
3Kranhäuser_Cologne_April_2018.jpgThe Rhine at dusk, the lit cathedral on the right. No alt text, no caption: only the pixels could find it0.95

If nothing matches, jurl tells you so instead of handing you the least-bad photo:

$ jurl --find "a carnival parade" en.wikipedia.org/wiki/Cologne
jurl: no image in https://en.wikipedia.org/wiki/Cologne looks like "a carnival parade" (closest: …, p=0.02)

JavaScript apps

Some pages arrive empty because their content is built by JavaScript. jurl spots those and renders them in Lightpanda, a fast headless browser:

$ jurl -n 2 hn.algolia.com
jurl: no text without JavaScript, rendering with Lightpanda…
# HN Search powered by Algolia

<https://hn.algolia.com/> · listing (1.00)

Stephen Hawking has died(http://www.bbc.com/news/uk-43396008)

6015 points|Cogito|9 years ago|436 comments

Use -r to force it.

Recipes

Everything prints plain text or plain URLs, so jurl fits in a pipe:

# Read the top 3 Hacker News stories, one key paragraph each
jurl -l -n 3 news.ycombinator.com | xargs -n1 jurl -n 1

# Download the photo that matches
jurl -f "a bridge over a river" en.wikipedia.org/wiki/Cologne | xargs curl -sO

# Every chart in a post
jurl -f "a chart or graph" -n 10 blog.cloudflare.com/markdown-for-agents/ | xargs -n1 curl -sO

# Only the confident answers, for scripts
jurl --json -q "installation" github.com/BurntSushi/ripgrep | jq -r '.blocks[] | select(.p > 0.8) | .text'

Reference

Flag
-q, --ask "…"Keep what answers the question. Works with every mode
-c, --codeCode blocks only
-l, --linksContent links, best first
-i, --imageContent images, judged by file name, alt text and caption
--visionLike --image, plus Clef looks at the pixels
-f, --find "…"The image that best matches the description
-r, --renderRun the page's JavaScript first (automatic for empty JavaScript apps)
-n, --max NHow many results (12 blocks, 5 with --ask, 8 code blocks, 20 links, 1 with --find)
-a, --allNo limit: everything above the threshold
--threshold PMinimum probability (default 0.5)
--jsonMachine-readable output, with every probability
-t, --timingWhere the time went, on stderr
jurl initSet or replace your API keys

Keys live in ~/.config/jurl/env. Environment variables take precedence over that file: TYPESAFE_API_KEY, and for images CLOUDFLARE_ACCOUNT_ID plus CLOUDFLARE_AI_TOKEN.

Speed and cost

CommandTypical timeTypical cost
jurl <url>0.6–1 s$0.0004–0.0013
jurl -i <url>0.5–0.7 s$0.00004
jurl --vision <url>1.3–1.9 s$0.0002
jurl --find "…" <url> (73 images)2.2–2.9 s$0.002
JavaScript apps+3–6 ssame

Measured on 2026-10-04. Run any command with -t to see your own numbers.

Your keys, your data

jurl has no server and no account of its own. It talks to the model APIs directly with your keys:

  • Who you pay: usage is billed by TypeSafe (Jev, $0.042 per million input tokens) and Cloudflare (Clef-flash, $0.09 per million). Output is free on both.
  • What leaves your machine: the text of the page goes to TypeSafe. With --vision or --find, the images go to Cloudflare too. Keep that in mind for internal or private pages.
  • What jurl can't read: it sends no cookies, so pages behind a login are out of reach.

Update and uninstall

  • Update: brew upgrade jurl, or run the install script again.
  • Uninstall: brew uninstall jurl, or delete ~/.local/bin/jurl. To remove everything, also delete ~/.config/jurl (keys) and ~/Library/Caches/jurl or ~/.cache/jurl (the browser).
How it works
url ─▶ fetch (asks for markdown first) ─▶ split into blocks · links · images
          └─ empty JavaScript app? ─▶ render in Lightpanda
                          │
            one yes/no question per candidate
              ┌───────────┴───────────┐
              ▼                       ▼
       Jev reads text          Clef looks at images
   "is block 12 the point?"   "does this show a cathedral?"
          p = 0.97                  p = 0.96
              └───────────┬───────────┘
                          ▼
        rank · threshold · page order · print verbatim
  • Fetch. jurl asks for text/markdown first; sites using Cloudflare's Markdown for Agents send it already converted. Otherwise jurl's own extractor drops navigation, footers, asides and scripts, and splits the rest into headings, paragraphs, list items, code, quotes and tables.
  • Ask. Every candidate becomes one yes/no question, and they all go to Jev in a single request. Asking 100 questions takes about as long as asking one.
  • Look. Clef-flash checks the pixels of each image (downscaled to 384 px) while Jev reads the text, not after.
  • Print. jurl keeps what clears the threshold, in page order. A heading comes back only when something in its section does.
Design notes
  • Models choose, they don't write. That's why the output is safe to pipe into xargs curl.
  • Clef is better at "what is this?" than at "does this matter?" Asked whether a chart was "meaningful content", Clef said 0.12. Asked what it shows (photo, chart, logo, avatar…), it was right. --vision adds up the content classes and averages them with Jev's judgement of the page context: a portrait is content on Wikipedia and noise in an author box.
  • For --find, Clef decides alone. It sees the pixels and the alt text and caption.
  • Rendering waits for text, not silence. Single-page apps never stop talking to the network, so Lightpanda waits for the network to calm down and for 1,500 characters of visible text, up to 8 s.
  • Several tricks keep --vision fast:
    • Clef runs at the same time as Jev.
    • Each image gets its own HTTP/1 connection: sharing one HTTP/2 connection made the slowest calls about twice as slow.
    • A call that takes longer than 700 ms is sent again, and whichever copy answers first wins.
    • An image still pending at 2.5 s keeps its text-only score.
  • Lightpanda isn't bundled. It's AGPL-3.0 and about 90 MB. jurl downloads a pinned 1.0.0 from Lightpanda's official release, checks its SHA-256 and caches it. If one is already on your PATH, jurl uses that. JURL_LIGHTPANDA points to a specific binary, and JURL_NO_DOWNLOAD stops the download. Lightpanda has no Windows build, so on Windows rendering is unavailable.
Development
cargo test        # extraction: layout tables, lazy images, markdown, links, app-shell detection
cargo build --release && ./target/release/jurl -t <url>
FileWhat's in it
src/main.rsModes, chunking, hedging, output
src/extract.rsHTML and markdown → blocks, links, images
src/decide.rsJev and Clef clients
src/fetch.rs · src/lightpanda.rsFetching, rendering, the browser download
src/setup.rs · src/config.rsFirst-run key prompt, jurl init, key storage

Releases are built by cargo-dist when a v* tag is pushed.

Contributing

Issues and pull requests are welcome. Read CONTRIBUTING.md first; security problems go to SECURITY.md.

License

MIT or Apache-2.0, at your option. Lightpanda, which jurl downloads for JavaScript pages, is a separate program under AGPL-3.0.

cli
curl
llm
rust
web-scraping

rayoplateado/jurl

curl that reads the page for you

Rust

1

40 commits

updated Oct 5, 2026

See the code

See what people are saying

SourceMessageScoreDate

Show HN: Jurl – curl that reads the page for you

5

Oct 5, 2026

README

Ink drawing of a man in a suit and leopard-print tie, pinching his fingers as if picking something out

jurl

curl that reads the page for you.
Tell it what you want. It picks it out of the page.


$ jurl -q "how do I install it on macOS?" github.com/BurntSushi/ripgrep
…
### Installation

…

If you're a macOS Homebrew or a Linuxbrew user, then you can install ripgrep from homebrew-core:

```
$ brew install ripgrep
```

If you're a MacPorts user, then you can install ripgrep from the official ports:

```
$ sudo port install ripgrep
```
$ jurl --find "a cathedral" en.wikipedia.org/wiki/Cologne
https://thumb.wikimedia.org/wikipedia/commons/thumb/2/27/Kdom.jpg/250px-Kdom.jpg
$ jurl --links -n 3 news.ycombinator.com
https://github.com/PowderworksCode/headstart
https://www.da.vidbuchanan.co.uk/blog/hacking-time.html
https://gamehistory.org/5k-magazines/

Why jurl

  • It picks. It doesn't write. jurl doesn't use a chatbot. It uses decision models: Jev reads the text and Clef looks at the images. Neither writes a word. They only say how likely each paragraph, link or image is to be what you want, and jurl prints the winners exactly as they appear on the page. You never get a summary that drifts or a URL that doesn't exist.
  • It's fast. Most pages take under a second, end to end.
  • It's cheap. About a thousand pages per dollar of API usage.

Install

brew install rayoplateado/tap/jurl                                                   # macOS, Linux
curl -LsSf https://github.com/rayoplateado/jurl/releases/latest/download/jurl-installer.sh | sh   # no Homebrew

Windows: powershell -ExecutionPolicy Bypass -c "irm https://github.com/rayoplateado/jurl/releases/latest/download/jurl-installer.ps1 | iex" · From source: cargo install --git https://github.com/rayoplateado/jurl

Then run it. The first time, jurl asks for a TypeSafe API key, checks it and saves it. That's the whole setup. Pages that need JavaScript just work too: jurl fetches a headless browser the first time one shows up.

To use --vision and --find, run jurl init and add a Cloudflare Workers AI token.

What you can do

Get the gist

$ jurl -n 3 blog.cloudflare.com/markdown-for-agents/
# Introducing Markdown for Agents

<https://blog.cloudflare.com/markdown-for-agents/> · article (1.00)

The way content and businesses are discovered online is changing rapidly. In the past, traffic originated from traditional search engines, and SEO determined who got found first. …

## Convert HTML to markdown, automatically

Cloudflare's network now supports real-time content conversion at the source, …

Here’s how it works. To fetch the markdown version of any page from a zone with Markdown for Agents enabled, the client needs to add the **Accept** negotiation header …

You get the title, what kind of page it is, and the blocks that carry it, in reading order. Navigation, cookie banners, sign-up prompts, author bios and footers are left out.

Ask one question

$ jurl -q "What is the Jevons paradox?" -n 2 en.wikipedia.org/wiki/William_Stanley_Jevons
…
Jevons received public recognition for his work on The Coal Question (1865), in which he called attention to the gradual exhaustion of Britain's coal supplies and also put forth the view that increases in energy production efficiency leads to more, not less, consumption. …

## Practical economics

In The Coal Question, Jevons covered a breadth of concepts on energy depletion …

Out of a page with about 240 blocks, you get the two paragraphs that answer it.

Just the code

$ jurl --code -n 2 github.com/BurntSushi/ripgrep
### Installation

```
$ brew install ripgrep
```

```
$ cargo install ripgrep
```
$ jurl --links -n 3 -q "official installation instructions" github.com/BurntSushi/ripgrep
https://www.macports.org/ports.php?by=name&substr=ripgrep
https://packages.gentoo.org/packages/sys-apps/ripgrep
https://chocolatey.org/packages/ripgrep

The ranking puts content first, ahead of login, share or privacy policy. Add -q to keep only the links about something specific.

The real images

--image picks the content images by their file name, alt text and caption, with no logos, icons or tracking pixels. --vision also has Clef look at the pixels, which matters when the alt text says nothing:

Image on a Cloudflare blog postAlt text--image--vision
Stacked area chartBLOG-3162 40.640.82
DiagramBLOG-3162 30.630.82
Author avatarWill Allen0.380.21
Company logoCloudflare0.090.08

Find a photo by description

$ jurl --find "a cathedral" -n 3 en.wikipedia.org/wiki/Cologne

The page has 73 images, and Clef looks at every one of them in parallel. It took 2.3 seconds:

ImageWhy it matchedp
1Kdom.jpgCologne Cathedral0.96
2Köln_um_1890.jpgThe 1890 skyline, the cathedral towering over it0.95
3Kranhäuser_Cologne_April_2018.jpgThe Rhine at dusk, the lit cathedral on the right. No alt text, no caption: only the pixels could find it0.95

If nothing matches, jurl tells you so instead of handing you the least-bad photo:

$ jurl --find "a carnival parade" en.wikipedia.org/wiki/Cologne
jurl: no image in https://en.wikipedia.org/wiki/Cologne looks like "a carnival parade" (closest: …, p=0.02)

JavaScript apps

Some pages arrive empty because their content is built by JavaScript. jurl spots those and renders them in Lightpanda, a fast headless browser:

$ jurl -n 2 hn.algolia.com
jurl: no text without JavaScript, rendering with Lightpanda…
# HN Search powered by Algolia

<https://hn.algolia.com/> · listing (1.00)

Stephen Hawking has died(http://www.bbc.com/news/uk-43396008)

6015 points|Cogito|9 years ago|436 comments

Use -r to force it.

Recipes

Everything prints plain text or plain URLs, so jurl fits in a pipe:

# Read the top 3 Hacker News stories, one key paragraph each
jurl -l -n 3 news.ycombinator.com | xargs -n1 jurl -n 1

# Download the photo that matches
jurl -f "a bridge over a river" en.wikipedia.org/wiki/Cologne | xargs curl -sO

# Every chart in a post
jurl -f "a chart or graph" -n 10 blog.cloudflare.com/markdown-for-agents/ | xargs -n1 curl -sO

# Only the confident answers, for scripts
jurl --json -q "installation" github.com/BurntSushi/ripgrep | jq -r '.blocks[] | select(.p > 0.8) | .text'

Reference

Flag
-q, --ask "…"Keep what answers the question. Works with every mode
-c, --codeCode blocks only
-l, --linksContent links, best first
-i, --imageContent images, judged by file name, alt text and caption
--visionLike --image, plus Clef looks at the pixels
-f, --find "…"The image that best matches the description
-r, --renderRun the page's JavaScript first (automatic for empty JavaScript apps)
-n, --max NHow many results (12 blocks, 5 with --ask, 8 code blocks, 20 links, 1 with --find)
-a, --allNo limit: everything above the threshold
--threshold PMinimum probability (default 0.5)
--jsonMachine-readable output, with every probability
-t, --timingWhere the time went, on stderr
jurl initSet or replace your API keys

Keys live in ~/.config/jurl/env. Environment variables take precedence over that file: TYPESAFE_API_KEY, and for images CLOUDFLARE_ACCOUNT_ID plus CLOUDFLARE_AI_TOKEN.

Speed and cost

CommandTypical timeTypical cost
jurl <url>0.6–1 s$0.0004–0.0013
jurl -i <url>0.5–0.7 s$0.00004
jurl --vision <url>1.3–1.9 s$0.0002
jurl --find "…" <url> (73 images)2.2–2.9 s$0.002
JavaScript apps+3–6 ssame

Measured on 2026-10-04. Run any command with -t to see your own numbers.

Your keys, your data

jurl has no server and no account of its own. It talks to the model APIs directly with your keys:

  • Who you pay: usage is billed by TypeSafe (Jev, $0.042 per million input tokens) and Cloudflare (Clef-flash, $0.09 per million). Output is free on both.
  • What leaves your machine: the text of the page goes to TypeSafe. With --vision or --find, the images go to Cloudflare too. Keep that in mind for internal or private pages.
  • What jurl can't read: it sends no cookies, so pages behind a login are out of reach.

Update and uninstall

  • Update: brew upgrade jurl, or run the install script again.
  • Uninstall: brew uninstall jurl, or delete ~/.local/bin/jurl. To remove everything, also delete ~/.config/jurl (keys) and ~/Library/Caches/jurl or ~/.cache/jurl (the browser).
How it works
url ─▶ fetch (asks for markdown first) ─▶ split into blocks · links · images
          └─ empty JavaScript app? ─▶ render in Lightpanda
                          │
            one yes/no question per candidate
              ┌───────────┴───────────┐
              ▼                       ▼
       Jev reads text          Clef looks at images
   "is block 12 the point?"   "does this show a cathedral?"
          p = 0.97                  p = 0.96
              └───────────┬───────────┘
                          ▼
        rank · threshold · page order · print verbatim
  • Fetch. jurl asks for text/markdown first; sites using Cloudflare's Markdown for Agents send it already converted. Otherwise jurl's own extractor drops navigation, footers, asides and scripts, and splits the rest into headings, paragraphs, list items, code, quotes and tables.
  • Ask. Every candidate becomes one yes/no question, and they all go to Jev in a single request. Asking 100 questions takes about as long as asking one.
  • Look. Clef-flash checks the pixels of each image (downscaled to 384 px) while Jev reads the text, not after.
  • Print. jurl keeps what clears the threshold, in page order. A heading comes back only when something in its section does.
Design notes
  • Models choose, they don't write. That's why the output is safe to pipe into xargs curl.
  • Clef is better at "what is this?" than at "does this matter?" Asked whether a chart was "meaningful content", Clef said 0.12. Asked what it shows (photo, chart, logo, avatar…), it was right. --vision adds up the content classes and averages them with Jev's judgement of the page context: a portrait is content on Wikipedia and noise in an author box.
  • For --find, Clef decides alone. It sees the pixels and the alt text and caption.
  • Rendering waits for text, not silence. Single-page apps never stop talking to the network, so Lightpanda waits for the network to calm down and for 1,500 characters of visible text, up to 8 s.
  • Several tricks keep --vision fast:
    • Clef runs at the same time as Jev.
    • Each image gets its own HTTP/1 connection: sharing one HTTP/2 connection made the slowest calls about twice as slow.
    • A call that takes longer than 700 ms is sent again, and whichever copy answers first wins.
    • An image still pending at 2.5 s keeps its text-only score.
  • Lightpanda isn't bundled. It's AGPL-3.0 and about 90 MB. jurl downloads a pinned 1.0.0 from Lightpanda's official release, checks its SHA-256 and caches it. If one is already on your PATH, jurl uses that. JURL_LIGHTPANDA points to a specific binary, and JURL_NO_DOWNLOAD stops the download. Lightpanda has no Windows build, so on Windows rendering is unavailable.
Development
cargo test        # extraction: layout tables, lazy images, markdown, links, app-shell detection
cargo build --release && ./target/release/jurl -t <url>
FileWhat's in it
src/main.rsModes, chunking, hedging, output
src/extract.rsHTML and markdown → blocks, links, images
src/decide.rsJev and Clef clients
src/fetch.rs · src/lightpanda.rsFetching, rendering, the browser download
src/setup.rs · src/config.rsFirst-run key prompt, jurl init, key storage

Releases are built by cargo-dist when a v* tag is pushed.

Contributing

Issues and pull requests are welcome. Read CONTRIBUTING.md first; security problems go to SECURITY.md.

License

MIT or Apache-2.0, at your option. Lightpanda, which jurl downloads for JavaScript pages, is a separate program under AGPL-3.0.

cli
curl
llm
rust
web-scraping