frankzch/ai-news-brief

Open-source AI news aggregator & daily digest engine: pulls 94 sources (RSS, Hacker News, Reddit, X, YouTube transcripts, GitHub Trending), LLM-filters, scores, dedups (SimHash + pgvector) and writes EN/中文 summaries. Powers inbrief.info.

Python

9

3 commits

updated Sep 19, 2026

See the code

See what people are saying

README

AI News Brief

English | 中文

An open-source, self-hostable AI news aggregator that turns 94 sources — AI lab blogs, Hacker News, Reddit, 35 AI builders on X, 13 YouTube channels (via transcripts) and GitHub Trending — into a deduplicated daily AI digest with English and Chinese summaries. It is the engine behind inbrief.info; if you just want to read the digest, use the site or the agent skill instead of self-hosting.

https://github.com/user-attachments/assets/4c2a209e-46dd-422a-ae8f-82d92c73f68c

It continuously pulls AI-related content from RSS feeds, Hacker News, Reddit, X (Twitter), YouTube and GitHub Trending, then runs every item through an LLM pipeline — relevance filtering, bilingual (EN/ZH) summarization, tag extraction, importance scoring — deduplicates it in two stages, and stores the curated result in PostgreSQL.

scheduler (main.py)
  └─ pipeline_runner ─ fetchers (RSS / HN / Reddit / X / YouTube / GitHub Trending)
       └─ content_processor  (trafilatura + curl_cffi + DrissionPage fallbacks)
            └─ ai_engine     (LLM: filter / summarize / tag / score)
                 └─ dedup    (SimHash quick screen + pgvector semantic)
                      └─ PostgreSQL (articles, auto-created schema)

Features

  • Multi-source fetchers — RSS/Atom, Hugging Face weekly papers, Hacker News (with comment threads), Reddit (subreddit hot or keyword search), X.com keyword search, YouTube (with transcript extraction), GitHub Trending.
  • Anti-bot resilient scraping — layered strategy: curl_cffi TLS impersonation → DrissionPage real browser → httpx; Playwright with a persistent login profile for Reddit / X / YouTube.
  • LLM curation — per-category prompts (news / discussion / video / opensource) produce bilingual summaries, tags and an importance score; any OpenAI-compatible API works (DeepSeek by default).
  • Two-stage deduplication — SimHash fingerprint quick screening, then pgvector cosine similarity with dynamic thresholds (stricter for last-24h articles).
  • Engagement gates & delayed re-scan — low-traction posts are recorded and re-checked later instead of being fetched repeatedly.
  • Optional content moderation — Aliyun Green text moderation before storage (skipped when keys are absent).

Source catalog

The production instance at inbrief.info currently tracks 94 sources across four categories. The engine ships with an empty source table — use this catalog as a starting point and add the ones you want via admin_rss.py.

📰 News & blogs — 16 RSS feeds

SourceFeed
OpenAI Bloghttps://openai.com/news/rss.xml
Google DeepMind Bloghttps://deepmind.google/blog/rss.xml
Google Research Bloghttps://research.google/blog/rss/
Apple Machine Learninghttps://machinelearning.apple.com/rss.xml
Microsoft AI Bloghttps://blogs.microsoft.com/ai/feed/
Nvidia Deep Learning Bloghttps://blogs.nvidia.com/blog/category/deep-learning/feed/
Nvidia Developer Bloghttps://developer.nvidia.com/blog/feed/
Hugging Face Bloghttps://huggingface.co/blog/feed.xml
HF Daily Papers (community-voted, links to arXiv)https://huggingface.co/api/daily_papers
TechCrunch AIhttps://techcrunch.com/category/artificial-intelligence/feed/
The Vergehttps://www.theverge.com/rss/index.xml
MIT Technology Review AIhttps://www.technologyreview.com/topic/artificial-intelligence/feed/
VentureBeat AIhttps://venturebeat.com/category/ai/feed
MarkTechPosthttps://www.marktechpost.com/feed/
AI News (artificialintelligence-news.com)https://www.artificialintelligence-news.com/feed/
Machine Learning Masteryhttps://machinelearningmastery.com/feed/

💬 Discussion — Hacker News, 16 Reddit sources, 47 X sources

Hacker News — front page via https://news.ycombinator.com/rss, with full comment-thread extraction and engagement gates (min upvotes / comments).

Reddit — 11 subreddits (hot posts): r/OpenAI, r/artificial, r/MachineLearning, r/ChatGPT, r/ClaudeAI, r/GeminiAI, r/DeepSeek, r/PromptEngineering, r/ArtificialInteligence, r/openclaw, r/AIToolTesting

Reddit — 5 keyword searches: llm, codex, prompt ai, agent ai, skill ai

X.com — 35 KOL accounts (high-engagement posts from their timelines):

Sam Altman (@sama)Andrej Karpathy (@karpathy)Yann LeCun (@ylecun)Demis Hassabis (@demishassabis)
Fei-Fei Li (@drfeifei)François Chollet (@fchollet)John Carmack (@ID_AA_Carmack)Lilian Weng (@lilianweng)
Amanda Askell (@AmandaAskell)Alex Albert (@alexalbert__)Boris Cherny (@bcherny)Cat Wu (@_catwu)
Simon Willison (@simonw)swyx (@swyx)Riley Goodside (@goodside)Jeremy Howard (@jeremyphoward)
Guillermo Rauch (@rauchg)Amjad Masad (@amasad)Aaron Levie (@levie)Garry Tan (@garrytan)
Kevin Weil (@kevinweil)Peter Steinberger (@steipete)Peter Yang (@petergyang)Dan Shipper (@danshipper)
Matt Turck (@mattturck)Nan Yu (@thenanyu)Nikunj Kothari (@nikunj)Josh Woodward (@joshwoodward)
Ryo Lu (@ryolu_)Thariq (@trq212)Aditya Agarwal (@adityaag)Madhu Guru (@realmadhuguru)
Claude (@claudeai)ClaudeDevs (@ClaudeDevs)Google Labs (@GoogleLabs)

X.com — 12 keyword searches: AI, Anthropic, OpenAI, ChatGPT, Gemini, LLM, claude code, codex, OpenClaw, prompt ai, agent ai, skill ai

🎬 Video — 13 YouTube channels (with transcript extraction)

ChannelFocus
Lex FridmanLong-form AI interviews
Dwarkesh PatelDeep interviews with AI researchers
Two Minute PapersPaper explainers
Yannic KilcherPaper deep-dives
FireshipDev news in 100 seconds
Matt WolfeAI tools & news roundups
Wes RothAI news commentary
Latent SpaceAI engineering podcast
No PriorsAI founders & investors
Sequoia CapitalTraining Data podcast
Redpoint AIUnsupervised Learning podcast
Every IncAI & work essays
Data Driven NYCData/AI talks

Weekly GitHub Trending repositories (top 25 by default), each summarized from its README and repo metadata. Configured in config.yaml under fetching.github_trending — no source entry needed.

Requirements

  • Python 3.10+
  • PostgreSQL with the pgvector extension (a free Supabase project works out of the box)
  • An OpenAI-compatible LLM API key

Quick start

pip install -r requirements.txt
playwright install chromium

cp .env.example .env        # fill in LLM key + Postgres URL

Enable pgvector once in your database:

CREATE EXTENSION IF NOT EXISTS vector;

All tables are created automatically on first run.

Add some sources (they live in the rss_sources table):

python admin_rss.py add https://openai.com/news/rss.xml "OpenAI Blog" --category news
python admin_rss.py add https://www.reddit.com/r/LocalLLaMA/ "r/LocalLLaMA" --category discussion
python admin_rss.py list

Source routing is inferred from the URL: reddit.com/r/<sub> → subreddit hot posts, reddit.com + a description starting with keyword → Reddit keyword search, x.com → X keyword search, YouTube channel feeds → transcript pipeline, everything else → RSS/Atom. GitHub Trending is enabled in config.yaml and needs no source entry.

Run:

python main.py --now   # single fetch round, good for a first test
python main.py         # scheduler mode: runs every schedule.interval_hours

Models, timeouts, engagement thresholds, per-platform quotas and retention are all in config.yaml.

Query it from your AI agent — no self-hosting

Don't want to run the pipeline yourself? The same curated feed is available as an agent skill that talks to the hosted inbrief.info API — drop it into Claude Code / Codex / OpenClaw / Antigravity and just ask in plain English. No keys, no scraping, works out of the box.

# Claude Code (other agents: swap the target dir, e.g. ~/.codex/skills, ~/.agents/skills)
git clone https://github.com/frankzch/ai-news-skill.git ~/.claude/skills/ai-news-skill

The agent turns your intent into precise filters (category, source, time range, count, summary length, language):

  • "AI videos from top KOLs in the past 5 days, exclude Fireship."
  • "Today's AI news, but drop TechReview."
  • "Reddit AI discussions from the past 3 days, long summaries."

Guests get up to 3 items per request; sign up free at inbrief.info for more. Full details: github.com/frankzch/ai-news-skill.

Notes

  • Reddit / X / YouTube fetching drives a real Chromium via Playwright. On first run you may be prompted to log in once in the opened browser window; the session persists in data/playwright_profile (git-ignored, stays on your machine).
  • Summaries are stored bilingually (*_en / *_zh columns). If you only need one language you can simply ignore the other.
  • The data/ directory holds runtime state (browser profile, cookies, daily flags) and is never committed.

FAQ

How is this different from an RSS reader? An RSS reader shows every item from every feed. This pipeline also pulls Reddit, X, Hacker News comment threads and YouTube transcripts, drops off-topic and low-engagement items, merges duplicates of the same story across sources, and writes a short summary and importance score for each one.

Which LLM does it need? Any OpenAI-compatible API; DeepSeek is the default. Set the model and key in .env / config.yaml.

Can I read the digest without self-hosting? Yes. The same feed is free to browse at inbrief.info, or query it from Claude Code / Codex with the agent skill.

License

MIT

ai
ai-digest
ai-news
ai-news-aggregator
bilingual
deepseek
hacker-news
llm
llm-summarization
news
news-aggregator
pgvector
python
reddit
rss
self-hosted
web-scraping
youtube-transcripts

frankzch/ai-news-brief

Open-source AI news aggregator & daily digest engine: pulls 94 sources (RSS, Hacker News, Reddit, X, YouTube transcripts, GitHub Trending), LLM-filters, scores, dedups (SimHash + pgvector) and writes EN/中文 summaries. Powers inbrief.info.

Python

9

3 commits

updated Sep 19, 2026

See the code

See what people are saying

README

AI News Brief

English | 中文

An open-source, self-hostable AI news aggregator that turns 94 sources — AI lab blogs, Hacker News, Reddit, 35 AI builders on X, 13 YouTube channels (via transcripts) and GitHub Trending — into a deduplicated daily AI digest with English and Chinese summaries. It is the engine behind inbrief.info; if you just want to read the digest, use the site or the agent skill instead of self-hosting.

https://github.com/user-attachments/assets/4c2a209e-46dd-422a-ae8f-82d92c73f68c

It continuously pulls AI-related content from RSS feeds, Hacker News, Reddit, X (Twitter), YouTube and GitHub Trending, then runs every item through an LLM pipeline — relevance filtering, bilingual (EN/ZH) summarization, tag extraction, importance scoring — deduplicates it in two stages, and stores the curated result in PostgreSQL.

scheduler (main.py)
  └─ pipeline_runner ─ fetchers (RSS / HN / Reddit / X / YouTube / GitHub Trending)
       └─ content_processor  (trafilatura + curl_cffi + DrissionPage fallbacks)
            └─ ai_engine     (LLM: filter / summarize / tag / score)
                 └─ dedup    (SimHash quick screen + pgvector semantic)
                      └─ PostgreSQL (articles, auto-created schema)

Features

  • Multi-source fetchers — RSS/Atom, Hugging Face weekly papers, Hacker News (with comment threads), Reddit (subreddit hot or keyword search), X.com keyword search, YouTube (with transcript extraction), GitHub Trending.
  • Anti-bot resilient scraping — layered strategy: curl_cffi TLS impersonation → DrissionPage real browser → httpx; Playwright with a persistent login profile for Reddit / X / YouTube.
  • LLM curation — per-category prompts (news / discussion / video / opensource) produce bilingual summaries, tags and an importance score; any OpenAI-compatible API works (DeepSeek by default).
  • Two-stage deduplication — SimHash fingerprint quick screening, then pgvector cosine similarity with dynamic thresholds (stricter for last-24h articles).
  • Engagement gates & delayed re-scan — low-traction posts are recorded and re-checked later instead of being fetched repeatedly.
  • Optional content moderation — Aliyun Green text moderation before storage (skipped when keys are absent).

Source catalog

The production instance at inbrief.info currently tracks 94 sources across four categories. The engine ships with an empty source table — use this catalog as a starting point and add the ones you want via admin_rss.py.

📰 News & blogs — 16 RSS feeds

SourceFeed
OpenAI Bloghttps://openai.com/news/rss.xml
Google DeepMind Bloghttps://deepmind.google/blog/rss.xml
Google Research Bloghttps://research.google/blog/rss/
Apple Machine Learninghttps://machinelearning.apple.com/rss.xml
Microsoft AI Bloghttps://blogs.microsoft.com/ai/feed/
Nvidia Deep Learning Bloghttps://blogs.nvidia.com/blog/category/deep-learning/feed/
Nvidia Developer Bloghttps://developer.nvidia.com/blog/feed/
Hugging Face Bloghttps://huggingface.co/blog/feed.xml
HF Daily Papers (community-voted, links to arXiv)https://huggingface.co/api/daily_papers
TechCrunch AIhttps://techcrunch.com/category/artificial-intelligence/feed/
The Vergehttps://www.theverge.com/rss/index.xml
MIT Technology Review AIhttps://www.technologyreview.com/topic/artificial-intelligence/feed/
VentureBeat AIhttps://venturebeat.com/category/ai/feed
MarkTechPosthttps://www.marktechpost.com/feed/
AI News (artificialintelligence-news.com)https://www.artificialintelligence-news.com/feed/
Machine Learning Masteryhttps://machinelearningmastery.com/feed/

💬 Discussion — Hacker News, 16 Reddit sources, 47 X sources

Hacker News — front page via https://news.ycombinator.com/rss, with full comment-thread extraction and engagement gates (min upvotes / comments).

Reddit — 11 subreddits (hot posts): r/OpenAI, r/artificial, r/MachineLearning, r/ChatGPT, r/ClaudeAI, r/GeminiAI, r/DeepSeek, r/PromptEngineering, r/ArtificialInteligence, r/openclaw, r/AIToolTesting

Reddit — 5 keyword searches: llm, codex, prompt ai, agent ai, skill ai

X.com — 35 KOL accounts (high-engagement posts from their timelines):

Sam Altman (@sama)Andrej Karpathy (@karpathy)Yann LeCun (@ylecun)Demis Hassabis (@demishassabis)
Fei-Fei Li (@drfeifei)François Chollet (@fchollet)John Carmack (@ID_AA_Carmack)Lilian Weng (@lilianweng)
Amanda Askell (@AmandaAskell)Alex Albert (@alexalbert__)Boris Cherny (@bcherny)Cat Wu (@_catwu)
Simon Willison (@simonw)swyx (@swyx)Riley Goodside (@goodside)Jeremy Howard (@jeremyphoward)
Guillermo Rauch (@rauchg)Amjad Masad (@amasad)Aaron Levie (@levie)Garry Tan (@garrytan)
Kevin Weil (@kevinweil)Peter Steinberger (@steipete)Peter Yang (@petergyang)Dan Shipper (@danshipper)
Matt Turck (@mattturck)Nan Yu (@thenanyu)Nikunj Kothari (@nikunj)Josh Woodward (@joshwoodward)
Ryo Lu (@ryolu_)Thariq (@trq212)Aditya Agarwal (@adityaag)Madhu Guru (@realmadhuguru)
Claude (@claudeai)ClaudeDevs (@ClaudeDevs)Google Labs (@GoogleLabs)

X.com — 12 keyword searches: AI, Anthropic, OpenAI, ChatGPT, Gemini, LLM, claude code, codex, OpenClaw, prompt ai, agent ai, skill ai

🎬 Video — 13 YouTube channels (with transcript extraction)

ChannelFocus
Lex FridmanLong-form AI interviews
Dwarkesh PatelDeep interviews with AI researchers
Two Minute PapersPaper explainers
Yannic KilcherPaper deep-dives
FireshipDev news in 100 seconds
Matt WolfeAI tools & news roundups
Wes RothAI news commentary
Latent SpaceAI engineering podcast
No PriorsAI founders & investors
Sequoia CapitalTraining Data podcast
Redpoint AIUnsupervised Learning podcast
Every IncAI & work essays
Data Driven NYCData/AI talks

Weekly GitHub Trending repositories (top 25 by default), each summarized from its README and repo metadata. Configured in config.yaml under fetching.github_trending — no source entry needed.

Requirements

  • Python 3.10+
  • PostgreSQL with the pgvector extension (a free Supabase project works out of the box)
  • An OpenAI-compatible LLM API key

Quick start

pip install -r requirements.txt
playwright install chromium

cp .env.example .env        # fill in LLM key + Postgres URL

Enable pgvector once in your database:

CREATE EXTENSION IF NOT EXISTS vector;

All tables are created automatically on first run.

Add some sources (they live in the rss_sources table):

python admin_rss.py add https://openai.com/news/rss.xml "OpenAI Blog" --category news
python admin_rss.py add https://www.reddit.com/r/LocalLLaMA/ "r/LocalLLaMA" --category discussion
python admin_rss.py list

Source routing is inferred from the URL: reddit.com/r/<sub> → subreddit hot posts, reddit.com + a description starting with keyword → Reddit keyword search, x.com → X keyword search, YouTube channel feeds → transcript pipeline, everything else → RSS/Atom. GitHub Trending is enabled in config.yaml and needs no source entry.

Run:

python main.py --now   # single fetch round, good for a first test
python main.py         # scheduler mode: runs every schedule.interval_hours

Models, timeouts, engagement thresholds, per-platform quotas and retention are all in config.yaml.

Query it from your AI agent — no self-hosting

Don't want to run the pipeline yourself? The same curated feed is available as an agent skill that talks to the hosted inbrief.info API — drop it into Claude Code / Codex / OpenClaw / Antigravity and just ask in plain English. No keys, no scraping, works out of the box.

# Claude Code (other agents: swap the target dir, e.g. ~/.codex/skills, ~/.agents/skills)
git clone https://github.com/frankzch/ai-news-skill.git ~/.claude/skills/ai-news-skill

The agent turns your intent into precise filters (category, source, time range, count, summary length, language):

  • "AI videos from top KOLs in the past 5 days, exclude Fireship."
  • "Today's AI news, but drop TechReview."
  • "Reddit AI discussions from the past 3 days, long summaries."

Guests get up to 3 items per request; sign up free at inbrief.info for more. Full details: github.com/frankzch/ai-news-skill.

Notes

  • Reddit / X / YouTube fetching drives a real Chromium via Playwright. On first run you may be prompted to log in once in the opened browser window; the session persists in data/playwright_profile (git-ignored, stays on your machine).
  • Summaries are stored bilingually (*_en / *_zh columns). If you only need one language you can simply ignore the other.
  • The data/ directory holds runtime state (browser profile, cookies, daily flags) and is never committed.

FAQ

How is this different from an RSS reader? An RSS reader shows every item from every feed. This pipeline also pulls Reddit, X, Hacker News comment threads and YouTube transcripts, drops off-topic and low-engagement items, merges duplicates of the same story across sources, and writes a short summary and importance score for each one.

Which LLM does it need? Any OpenAI-compatible API; DeepSeek is the default. Set the model and key in .env / config.yaml.

Can I read the digest without self-hosting? Yes. The same feed is free to browse at inbrief.info, or query it from Claude Code / Codex with the agent skill.

License

MIT

ai
ai-digest
ai-news
ai-news-aggregator
bilingual
deepseek
hacker-news
llm
llm-summarization
news
news-aggregator
pgvector
python
reddit
rss
self-hosted
web-scraping
youtube-transcripts