Open-source AI news aggregator & daily digest engine: pulls 94 sources (RSS, Hacker News, Reddit, X, YouTube transcripts, GitHub Trending), LLM-filters, scores, dedups (SimHash + pgvector) and writes EN/中文 summaries. Powers inbrief.info.
See the codeEnglish | 中文
An open-source, self-hostable AI news aggregator that turns 94 sources — AI lab blogs, Hacker News, Reddit, 35 AI builders on X, 13 YouTube channels (via transcripts) and GitHub Trending — into a deduplicated daily AI digest with English and Chinese summaries. It is the engine behind inbrief.info; if you just want to read the digest, use the site or the agent skill instead of self-hosting.
https://github.com/user-attachments/assets/4c2a209e-46dd-422a-ae8f-82d92c73f68c
It continuously pulls AI-related content from RSS feeds, Hacker News, Reddit, X (Twitter), YouTube and GitHub Trending, then runs every item through an LLM pipeline — relevance filtering, bilingual (EN/ZH) summarization, tag extraction, importance scoring — deduplicates it in two stages, and stores the curated result in PostgreSQL.
scheduler (main.py)
└─ pipeline_runner ─ fetchers (RSS / HN / Reddit / X / YouTube / GitHub Trending)
└─ content_processor (trafilatura + curl_cffi + DrissionPage fallbacks)
└─ ai_engine (LLM: filter / summarize / tag / score)
└─ dedup (SimHash quick screen + pgvector semantic)
└─ PostgreSQL (articles, auto-created schema)
curl_cffi TLS impersonation → DrissionPage real browser → httpx; Playwright with a persistent login profile for Reddit / X / YouTube.The production instance at inbrief.info currently tracks 94 sources across four categories. The engine ships with an empty source table — use this catalog as a starting point and add the ones you want via admin_rss.py.
| Source | Feed |
|---|---|
| OpenAI Blog | https://openai.com/news/rss.xml |
| Google DeepMind Blog | https://deepmind.google/blog/rss.xml |
| Google Research Blog | https://research.google/blog/rss/ |
| Apple Machine Learning | https://machinelearning.apple.com/rss.xml |
| Microsoft AI Blog | https://blogs.microsoft.com/ai/feed/ |
| Nvidia Deep Learning Blog | https://blogs.nvidia.com/blog/category/deep-learning/feed/ |
| Nvidia Developer Blog | https://developer.nvidia.com/blog/feed/ |
| Hugging Face Blog | https://huggingface.co/blog/feed.xml |
| HF Daily Papers (community-voted, links to arXiv) | https://huggingface.co/api/daily_papers |
| TechCrunch AI | https://techcrunch.com/category/artificial-intelligence/feed/ |
| The Verge | https://www.theverge.com/rss/index.xml |
| MIT Technology Review AI | https://www.technologyreview.com/topic/artificial-intelligence/feed/ |
| VentureBeat AI | https://venturebeat.com/category/ai/feed |
| MarkTechPost | https://www.marktechpost.com/feed/ |
| AI News (artificialintelligence-news.com) | https://www.artificialintelligence-news.com/feed/ |
| Machine Learning Mastery | https://machinelearningmastery.com/feed/ |
Hacker News — front page via https://news.ycombinator.com/rss, with full comment-thread extraction and engagement gates (min upvotes / comments).
Reddit — 11 subreddits (hot posts): r/OpenAI, r/artificial, r/MachineLearning, r/ChatGPT, r/ClaudeAI, r/GeminiAI, r/DeepSeek, r/PromptEngineering, r/ArtificialInteligence, r/openclaw, r/AIToolTesting
Reddit — 5 keyword searches: llm, codex, prompt ai, agent ai, skill ai
X.com — 35 KOL accounts (high-engagement posts from their timelines):
| Sam Altman (@sama) | Andrej Karpathy (@karpathy) | Yann LeCun (@ylecun) | Demis Hassabis (@demishassabis) |
| Fei-Fei Li (@drfeifei) | François Chollet (@fchollet) | John Carmack (@ID_AA_Carmack) | Lilian Weng (@lilianweng) |
| Amanda Askell (@AmandaAskell) | Alex Albert (@alexalbert__) | Boris Cherny (@bcherny) | Cat Wu (@_catwu) |
| Simon Willison (@simonw) | swyx (@swyx) | Riley Goodside (@goodside) | Jeremy Howard (@jeremyphoward) |
| Guillermo Rauch (@rauchg) | Amjad Masad (@amasad) | Aaron Levie (@levie) | Garry Tan (@garrytan) |
| Kevin Weil (@kevinweil) | Peter Steinberger (@steipete) | Peter Yang (@petergyang) | Dan Shipper (@danshipper) |
| Matt Turck (@mattturck) | Nan Yu (@thenanyu) | Nikunj Kothari (@nikunj) | Josh Woodward (@joshwoodward) |
| Ryo Lu (@ryolu_) | Thariq (@trq212) | Aditya Agarwal (@adityaag) | Madhu Guru (@realmadhuguru) |
| Claude (@claudeai) | ClaudeDevs (@ClaudeDevs) | Google Labs (@GoogleLabs) |
X.com — 12 keyword searches: AI, Anthropic, OpenAI, ChatGPT, Gemini, LLM, claude code, codex, OpenClaw, prompt ai, agent ai, skill ai
| Channel | Focus |
|---|---|
| Lex Fridman | Long-form AI interviews |
| Dwarkesh Patel | Deep interviews with AI researchers |
| Two Minute Papers | Paper explainers |
| Yannic Kilcher | Paper deep-dives |
| Fireship | Dev news in 100 seconds |
| Matt Wolfe | AI tools & news roundups |
| Wes Roth | AI news commentary |
| Latent Space | AI engineering podcast |
| No Priors | AI founders & investors |
| Sequoia Capital | Training Data podcast |
| Redpoint AI | Unsupervised Learning podcast |
| Every Inc | AI & work essays |
| Data Driven NYC | Data/AI talks |
Weekly GitHub Trending repositories (top 25 by default), each summarized from its README and repo metadata. Configured in config.yaml under fetching.github_trending — no source entry needed.
pip install -r requirements.txt
playwright install chromium
cp .env.example .env # fill in LLM key + Postgres URL
Enable pgvector once in your database:
CREATE EXTENSION IF NOT EXISTS vector;
All tables are created automatically on first run.
Add some sources (they live in the rss_sources table):
python admin_rss.py add https://openai.com/news/rss.xml "OpenAI Blog" --category news
python admin_rss.py add https://www.reddit.com/r/LocalLLaMA/ "r/LocalLLaMA" --category discussion
python admin_rss.py list
Source routing is inferred from the URL: reddit.com/r/<sub> → subreddit hot posts, reddit.com + a description starting with keyword → Reddit keyword search, x.com → X keyword search, YouTube channel feeds → transcript pipeline, everything else → RSS/Atom. GitHub Trending is enabled in config.yaml and needs no source entry.
Run:
python main.py --now # single fetch round, good for a first test
python main.py # scheduler mode: runs every schedule.interval_hours
Models, timeouts, engagement thresholds, per-platform quotas and retention are all in config.yaml.
Don't want to run the pipeline yourself? The same curated feed is available as an agent skill that talks to the hosted inbrief.info API — drop it into Claude Code / Codex / OpenClaw / Antigravity and just ask in plain English. No keys, no scraping, works out of the box.
# Claude Code (other agents: swap the target dir, e.g. ~/.codex/skills, ~/.agents/skills)
git clone https://github.com/frankzch/ai-news-skill.git ~/.claude/skills/ai-news-skill
The agent turns your intent into precise filters (category, source, time range, count, summary length, language):
Guests get up to 3 items per request; sign up free at inbrief.info for more. Full details: github.com/frankzch/ai-news-skill.
data/playwright_profile (git-ignored, stays on your machine).*_en / *_zh columns). If you only need one language you can simply ignore the other.data/ directory holds runtime state (browser profile, cookies, daily flags) and is never committed.How is this different from an RSS reader? An RSS reader shows every item from every feed. This pipeline also pulls Reddit, X, Hacker News comment threads and YouTube transcripts, drops off-topic and low-engagement items, merges duplicates of the same story across sources, and writes a short summary and importance score for each one.
Which LLM does it need?
Any OpenAI-compatible API; DeepSeek is the default. Set the model and key in .env / config.yaml.
Can I read the digest without self-hosting? Yes. The same feed is free to browse at inbrief.info, or query it from Claude Code / Codex with the agent skill.
Open-source AI news aggregator & daily digest engine: pulls 94 sources (RSS, Hacker News, Reddit, X, YouTube transcripts, GitHub Trending), LLM-filters, scores, dedups (SimHash + pgvector) and writes EN/中文 summaries. Powers inbrief.info.
See the codeEnglish | 中文
An open-source, self-hostable AI news aggregator that turns 94 sources — AI lab blogs, Hacker News, Reddit, 35 AI builders on X, 13 YouTube channels (via transcripts) and GitHub Trending — into a deduplicated daily AI digest with English and Chinese summaries. It is the engine behind inbrief.info; if you just want to read the digest, use the site or the agent skill instead of self-hosting.
https://github.com/user-attachments/assets/4c2a209e-46dd-422a-ae8f-82d92c73f68c
It continuously pulls AI-related content from RSS feeds, Hacker News, Reddit, X (Twitter), YouTube and GitHub Trending, then runs every item through an LLM pipeline — relevance filtering, bilingual (EN/ZH) summarization, tag extraction, importance scoring — deduplicates it in two stages, and stores the curated result in PostgreSQL.
scheduler (main.py)
└─ pipeline_runner ─ fetchers (RSS / HN / Reddit / X / YouTube / GitHub Trending)
└─ content_processor (trafilatura + curl_cffi + DrissionPage fallbacks)
└─ ai_engine (LLM: filter / summarize / tag / score)
└─ dedup (SimHash quick screen + pgvector semantic)
└─ PostgreSQL (articles, auto-created schema)
curl_cffi TLS impersonation → DrissionPage real browser → httpx; Playwright with a persistent login profile for Reddit / X / YouTube.The production instance at inbrief.info currently tracks 94 sources across four categories. The engine ships with an empty source table — use this catalog as a starting point and add the ones you want via admin_rss.py.
| Source | Feed |
|---|---|
| OpenAI Blog | https://openai.com/news/rss.xml |
| Google DeepMind Blog | https://deepmind.google/blog/rss.xml |
| Google Research Blog | https://research.google/blog/rss/ |
| Apple Machine Learning | https://machinelearning.apple.com/rss.xml |
| Microsoft AI Blog | https://blogs.microsoft.com/ai/feed/ |
| Nvidia Deep Learning Blog | https://blogs.nvidia.com/blog/category/deep-learning/feed/ |
| Nvidia Developer Blog | https://developer.nvidia.com/blog/feed/ |
| Hugging Face Blog | https://huggingface.co/blog/feed.xml |
| HF Daily Papers (community-voted, links to arXiv) | https://huggingface.co/api/daily_papers |
| TechCrunch AI | https://techcrunch.com/category/artificial-intelligence/feed/ |
| The Verge | https://www.theverge.com/rss/index.xml |
| MIT Technology Review AI | https://www.technologyreview.com/topic/artificial-intelligence/feed/ |
| VentureBeat AI | https://venturebeat.com/category/ai/feed |
| MarkTechPost | https://www.marktechpost.com/feed/ |
| AI News (artificialintelligence-news.com) | https://www.artificialintelligence-news.com/feed/ |
| Machine Learning Mastery | https://machinelearningmastery.com/feed/ |
Hacker News — front page via https://news.ycombinator.com/rss, with full comment-thread extraction and engagement gates (min upvotes / comments).
Reddit — 11 subreddits (hot posts): r/OpenAI, r/artificial, r/MachineLearning, r/ChatGPT, r/ClaudeAI, r/GeminiAI, r/DeepSeek, r/PromptEngineering, r/ArtificialInteligence, r/openclaw, r/AIToolTesting
Reddit — 5 keyword searches: llm, codex, prompt ai, agent ai, skill ai
X.com — 35 KOL accounts (high-engagement posts from their timelines):
| Sam Altman (@sama) | Andrej Karpathy (@karpathy) | Yann LeCun (@ylecun) | Demis Hassabis (@demishassabis) |
| Fei-Fei Li (@drfeifei) | François Chollet (@fchollet) | John Carmack (@ID_AA_Carmack) | Lilian Weng (@lilianweng) |
| Amanda Askell (@AmandaAskell) | Alex Albert (@alexalbert__) | Boris Cherny (@bcherny) | Cat Wu (@_catwu) |
| Simon Willison (@simonw) | swyx (@swyx) | Riley Goodside (@goodside) | Jeremy Howard (@jeremyphoward) |
| Guillermo Rauch (@rauchg) | Amjad Masad (@amasad) | Aaron Levie (@levie) | Garry Tan (@garrytan) |
| Kevin Weil (@kevinweil) | Peter Steinberger (@steipete) | Peter Yang (@petergyang) | Dan Shipper (@danshipper) |
| Matt Turck (@mattturck) | Nan Yu (@thenanyu) | Nikunj Kothari (@nikunj) | Josh Woodward (@joshwoodward) |
| Ryo Lu (@ryolu_) | Thariq (@trq212) | Aditya Agarwal (@adityaag) | Madhu Guru (@realmadhuguru) |
| Claude (@claudeai) | ClaudeDevs (@ClaudeDevs) | Google Labs (@GoogleLabs) |
X.com — 12 keyword searches: AI, Anthropic, OpenAI, ChatGPT, Gemini, LLM, claude code, codex, OpenClaw, prompt ai, agent ai, skill ai
| Channel | Focus |
|---|---|
| Lex Fridman | Long-form AI interviews |
| Dwarkesh Patel | Deep interviews with AI researchers |
| Two Minute Papers | Paper explainers |
| Yannic Kilcher | Paper deep-dives |
| Fireship | Dev news in 100 seconds |
| Matt Wolfe | AI tools & news roundups |
| Wes Roth | AI news commentary |
| Latent Space | AI engineering podcast |
| No Priors | AI founders & investors |
| Sequoia Capital | Training Data podcast |
| Redpoint AI | Unsupervised Learning podcast |
| Every Inc | AI & work essays |
| Data Driven NYC | Data/AI talks |
Weekly GitHub Trending repositories (top 25 by default), each summarized from its README and repo metadata. Configured in config.yaml under fetching.github_trending — no source entry needed.
pip install -r requirements.txt
playwright install chromium
cp .env.example .env # fill in LLM key + Postgres URL
Enable pgvector once in your database:
CREATE EXTENSION IF NOT EXISTS vector;
All tables are created automatically on first run.
Add some sources (they live in the rss_sources table):
python admin_rss.py add https://openai.com/news/rss.xml "OpenAI Blog" --category news
python admin_rss.py add https://www.reddit.com/r/LocalLLaMA/ "r/LocalLLaMA" --category discussion
python admin_rss.py list
Source routing is inferred from the URL: reddit.com/r/<sub> → subreddit hot posts, reddit.com + a description starting with keyword → Reddit keyword search, x.com → X keyword search, YouTube channel feeds → transcript pipeline, everything else → RSS/Atom. GitHub Trending is enabled in config.yaml and needs no source entry.
Run:
python main.py --now # single fetch round, good for a first test
python main.py # scheduler mode: runs every schedule.interval_hours
Models, timeouts, engagement thresholds, per-platform quotas and retention are all in config.yaml.
Don't want to run the pipeline yourself? The same curated feed is available as an agent skill that talks to the hosted inbrief.info API — drop it into Claude Code / Codex / OpenClaw / Antigravity and just ask in plain English. No keys, no scraping, works out of the box.
# Claude Code (other agents: swap the target dir, e.g. ~/.codex/skills, ~/.agents/skills)
git clone https://github.com/frankzch/ai-news-skill.git ~/.claude/skills/ai-news-skill
The agent turns your intent into precise filters (category, source, time range, count, summary length, language):
Guests get up to 3 items per request; sign up free at inbrief.info for more. Full details: github.com/frankzch/ai-news-skill.
data/playwright_profile (git-ignored, stays on your machine).*_en / *_zh columns). If you only need one language you can simply ignore the other.data/ directory holds runtime state (browser profile, cookies, daily flags) and is never committed.How is this different from an RSS reader? An RSS reader shows every item from every feed. This pipeline also pulls Reddit, X, Hacker News comment threads and YouTube transcripts, drops off-topic and low-engagement items, merges duplicates of the same story across sources, and writes a short summary and importance score for each one.
Which LLM does it need?
Any OpenAI-compatible API; DeepSeek is the default. Set the model and key in .env / config.yaml.
Can I read the digest without self-hosting? Yes. The same feed is free to browse at inbrief.info, or query it from Claude Code / Codex with the agent skill.