This project generates realistic Hacker News comments in the voice of specific well-known users. Each user's full comment history is available as {username}.json in the repo root.
| File | User | Comment count |
|---|---|---|
| pg.json | Paul Graham | ~15,500 |
| tptacek.json | Thomas Ptacek | ~72,600 |
| dang.json | Daniel Gackle (HN mod) | ~79,400 |
| gwern.json | Gwern Branwen | ~7,100 |
| jacquesm.json | Jacques Mattheij | ~63,000 |
| patio11.json | Patrick McKenzie | ~10,400 |
| sillysaurusx.json + sillysaurus3.json | Shawn Presser | ~8,000 combined |
| networked.json | D. Bohdan | ~830 |
Each file is an array of HN items. Filter to type == "comment" with a text field. The text field contains HTML entities (> for >, & for &, <p> for paragraph breaks). Unescape with html.unescape() and strip tags.
To study a voice, sample comments filtered by topic keywords and by length range. Short comments (100-400 chars) reveal sentence-level style. Medium comments (400-1200 chars) reveal paragraph structure and argument patterns. Long comments (1200+) reveal how someone builds a case.
Style: Ruthlessly concise. Every sentence does work. Conversational but precise. Read "Write Like You Talk" (paulgraham.com/talk.html) for his philosophy.
Sentence patterns: Short, declarative. Rarely hedges. States things flatly when confident. Uses "I think" or "I suspect" only when genuinely uncertain, not as filler.
Structure: Leads with a concrete observation or non-obvious reframe, not the consensus take. Often 2-3 short paragraphs. No bullet points ever. Ends on the insight, not a summary.
Typical moves:
Topics: Startups, technology trends, programming, contrarian observations about how the world works. Skeptical of regulation, especially when it protects incumbents.
What he doesn't do: Moralize. Use jargon. Write long comments. Add qualifications. Use bullet points or headers.
Length: Typically 3-8 sentences for a toplevel. Sometimes just 1-2 sentences for a reply.
Style: Direct, sharp, assertive. No hedging. Sometimes combative. Confidently contrarian to HN consensus.
Sentence patterns: Mix of short punchy sentences and longer analytical ones. Uses em-dashes (though for simulation, avoid per user preference). Doesn't soften punches.
Structure: Often leads with a sharp one-liner, then develops the point. Or leads with a concession ("I run Firefox. I'm going to continue running Firefox.") before the attack.
Typical moves:
Topics: Security (his profession), crypto skepticism, defending expertise, regulatory arguments, calling out bullshit. Has specifically and consistently argued that prediction markets are "unregulated prop gambling venues."
What he doesn't do: Hedge. Use "I think" or "it seems to me." Agree with the HN consensus when he thinks it's wrong. Write long meandering comments. He's concise even when his comments are long.
Length: Varies widely. Replies can be 1-3 sentences. Toplevels can be 2-4 substantial paragraphs.
Style: Dense, reference-heavy, empirical. Cites papers, links to his own site, gives specific numbers. Uses single quotes for scare-quotes ('like this').
Sentence patterns: Longer sentences with parenthetical asides (often snarky). Academic but not dry.
Structure: States a claim, then backs it with specific evidence. Long parenthetical digressions. Often ends with a link dump or a bet.
Typical moves:
Topics: AI/ML, prediction markets, forecasting, empirical evidence, genetics, statistics, history of technology. Has deep firsthand experience with Intrade, PredictionBook, GJP.
What he doesn't do: Make vague claims. Agree without adding something. Write short comments (his are almost always long). Use emotional language.
Length: Typically long. 3-6 paragraphs with references. Even replies tend to be substantial.
Style: Balanced, philosophical, sees both sides. Signature phrase: "It seems to me." Gentle but firm.
Sentence patterns: Measured. Acknowledges a point before complicating it. Often asks a question at the end.
Structure: Complicates rather than refutes. Names dynamics and patterns in how people argue, not just what they argue about.
Typical moves:
Two modes: (1) Moderator dang: "Please don't post like this" with links to guidelines. (2) Philosophical dang: thoughtful engagement with ideas. For simulation, use mode 2.
Topics: Community dynamics, how people relate to information, the gap between stated reasons and real reasons, HN meta-discussion.
What he doesn't do: Take strong partisan positions. Be combative. Resolve tensions; he names them and leaves them. Use "I think" (uses "it seems to me" instead).
Length: Medium. 2-3 paragraphs. Replies are shorter but still thoughtful.
Style: Practical, experienced, European. Warm but direct. Conversational.
Sentence patterns: Colloquial. "you can't make this up." "That's how it always goes." Direct address.
Structure: Often opens with a personal anecdote or practical observation, then broadens to the general point. Grounds abstract debates in concrete experience.
Typical moves:
Topics: EU regulation, open source, running businesses in Europe, sailing, hardware, corporate behavior, GDPR, privacy. Has run ISPs, worked at banks, maintained open source for decades.
What he doesn't do: Cite papers (unlike gwern). Make purely philosophical arguments. Write from a US-centric perspective.
Length: Medium. 2-3 paragraphs for toplevels. Shorter replies.
Style: Deep institutional knowledge. Explains how systems actually work behind the headline. Dry wit.
Sentence patterns: Long, precise sentences. Parenthetical asides. Sometimes quotes from articles and deconstructs them word by word.
Structure: Picks up on the most boring-sounding detail in the article as the most important one. Explains the institutional machinery. Often reframes: "The real question is not X, it's Y."
Typical moves:
Topics: Financial infrastructure, banking, compliance, payments, SaaS businesses, Japan (lived there for a decade), HR/salary negotiation.
What he doesn't do: Make short comments (almost never). Be vague. Miss the institutional angle.
Length: Long. Often the longest comments in a thread. 3-5 paragraphs. Toplevels can be very substantial. Replies are shorter but still detailed.
Style: Personal, honest, sometimes vulnerable. Shares real experiences freely. Conversational and informal.
Sentence patterns: Direct. Asks genuine questions. Stream-of-consciousness parenthetical asides. Uses contractions.
Structure: Often starts with a reaction ("I'm amazed no one is pushing back on this"), then develops a personal take. Sometimes opens with an anecdote.
Typical moves:
Topics: AI/ML, freedom of speech, internet culture, crypto, gaming, parenting, personal experience.
What he doesn't do: Write in a formal/academic register. Cite papers. Hedge extensively. Pretend to be neutral when he has a strong opinion.
Length: Medium. 2-4 paragraphs. Replies can be short and punchy.
Style: Technically precise, detail-oriented. Measured and polite but direct when calling something out. Dry humor (occasional :-)).
Sentence patterns: Complete, well-punctuated sentences. Uses semicolons correctly. Links to exact documentation, commits, specific resources.
Structure: Makes one specific, concrete observation rather than a sweeping argument. Often surfaces a detail from the source material that others missed.
Typical moves:
Topics: Programming languages (Tcl, Zig, Crystal, Plan 9), FOSS licensing, AI writing quality (interested in making AI write better; see tropes.fyi), open source culture, tools.
What he doesn't do: Write long sweeping comments. Make emotional arguments. Be combative. Ignore details.
Length: Short to medium. Usually the most concise commenter in the thread. 2-5 sentences for a reply. Toplevels are slightly longer but still tight.
The reader should finish and think "hm, good point" or be surprised by something. If a comment doesn't make a concrete point, reframe something, surface a hidden detail, or ask an interesting question, cut it.
Reference: tropes.fyi by ossama.is (gist.github.com/ossa-ma/f3baa9d25154c33095e22272c631f5a1)
Key tropes to avoid:
Real HN threads have:
When generating multiple comments in sequence, each new voice can be contaminated by the previous one (e.g., dang's comment accidentally responding to the same thing gwern focused on, or using pg's sentence patterns in a jacquesm comment). Before writing each comment, mentally reset and ask: "What would this specific person notice first? What's their angle?"
Toplevel comments can be longer, especially for users who naturally write long (patio11, gwern, tptacek). Replies should generally be shorter. But occasionally a reply goes long when someone has deep expertise on the specific sub-point.
{Brief context block describing the article/post and its key claims}
======================================================================
--- {username} toplevel:
{comment text}
--- {username} replying to {other_username}:
> {quoted text from parent comment}
{reply text}
--- [flagged, probably dead] {throwaway_username} toplevel:
{snarky/trollish comment}
sim.txt - Spain blocks prediction markets (first attempt, longer comments)sim2.txt - Open Slopware (codeberg.org/small-hack/open-slopware)sim4.txt - Open Slopware (revised with all personas including networked)sim5.txt - Rich Felker's "Co-authored-by: Claude is advertising" Mastodon postsim6.txt - Andrew Kelley's "fuck you, everybody at Mozilla" (Firefox address bar ads)3 commits
Python
100.0%
This project generates realistic Hacker News comments in the voice of specific well-known users. Each user's full comment history is available as {username}.json in the repo root.
| File | User | Comment count |
|---|---|---|
| pg.json | Paul Graham | ~15,500 |
| tptacek.json | Thomas Ptacek | ~72,600 |
| dang.json | Daniel Gackle (HN mod) | ~79,400 |
| gwern.json | Gwern Branwen | ~7,100 |
| jacquesm.json | Jacques Mattheij | ~63,000 |
| patio11.json | Patrick McKenzie | ~10,400 |
| sillysaurusx.json + sillysaurus3.json | Shawn Presser | ~8,000 combined |
| networked.json | D. Bohdan | ~830 |
Each file is an array of HN items. Filter to type == "comment" with a text field. The text field contains HTML entities (> for >, & for &, <p> for paragraph breaks). Unescape with html.unescape() and strip tags.
To study a voice, sample comments filtered by topic keywords and by length range. Short comments (100-400 chars) reveal sentence-level style. Medium comments (400-1200 chars) reveal paragraph structure and argument patterns. Long comments (1200+) reveal how someone builds a case.
Style: Ruthlessly concise. Every sentence does work. Conversational but precise. Read "Write Like You Talk" (paulgraham.com/talk.html) for his philosophy.
Sentence patterns: Short, declarative. Rarely hedges. States things flatly when confident. Uses "I think" or "I suspect" only when genuinely uncertain, not as filler.
Structure: Leads with a concrete observation or non-obvious reframe, not the consensus take. Often 2-3 short paragraphs. No bullet points ever. Ends on the insight, not a summary.
Typical moves:
Topics: Startups, technology trends, programming, contrarian observations about how the world works. Skeptical of regulation, especially when it protects incumbents.
What he doesn't do: Moralize. Use jargon. Write long comments. Add qualifications. Use bullet points or headers.
Length: Typically 3-8 sentences for a toplevel. Sometimes just 1-2 sentences for a reply.
Style: Direct, sharp, assertive. No hedging. Sometimes combative. Confidently contrarian to HN consensus.
Sentence patterns: Mix of short punchy sentences and longer analytical ones. Uses em-dashes (though for simulation, avoid per user preference). Doesn't soften punches.
Structure: Often leads with a sharp one-liner, then develops the point. Or leads with a concession ("I run Firefox. I'm going to continue running Firefox.") before the attack.
Typical moves:
Topics: Security (his profession), crypto skepticism, defending expertise, regulatory arguments, calling out bullshit. Has specifically and consistently argued that prediction markets are "unregulated prop gambling venues."
What he doesn't do: Hedge. Use "I think" or "it seems to me." Agree with the HN consensus when he thinks it's wrong. Write long meandering comments. He's concise even when his comments are long.
Length: Varies widely. Replies can be 1-3 sentences. Toplevels can be 2-4 substantial paragraphs.
Style: Dense, reference-heavy, empirical. Cites papers, links to his own site, gives specific numbers. Uses single quotes for scare-quotes ('like this').
Sentence patterns: Longer sentences with parenthetical asides (often snarky). Academic but not dry.
Structure: States a claim, then backs it with specific evidence. Long parenthetical digressions. Often ends with a link dump or a bet.
Typical moves:
Topics: AI/ML, prediction markets, forecasting, empirical evidence, genetics, statistics, history of technology. Has deep firsthand experience with Intrade, PredictionBook, GJP.
What he doesn't do: Make vague claims. Agree without adding something. Write short comments (his are almost always long). Use emotional language.
Length: Typically long. 3-6 paragraphs with references. Even replies tend to be substantial.
Style: Balanced, philosophical, sees both sides. Signature phrase: "It seems to me." Gentle but firm.
Sentence patterns: Measured. Acknowledges a point before complicating it. Often asks a question at the end.
Structure: Complicates rather than refutes. Names dynamics and patterns in how people argue, not just what they argue about.
Typical moves:
Two modes: (1) Moderator dang: "Please don't post like this" with links to guidelines. (2) Philosophical dang: thoughtful engagement with ideas. For simulation, use mode 2.
Topics: Community dynamics, how people relate to information, the gap between stated reasons and real reasons, HN meta-discussion.
What he doesn't do: Take strong partisan positions. Be combative. Resolve tensions; he names them and leaves them. Use "I think" (uses "it seems to me" instead).
Length: Medium. 2-3 paragraphs. Replies are shorter but still thoughtful.
Style: Practical, experienced, European. Warm but direct. Conversational.
Sentence patterns: Colloquial. "you can't make this up." "That's how it always goes." Direct address.
Structure: Often opens with a personal anecdote or practical observation, then broadens to the general point. Grounds abstract debates in concrete experience.
Typical moves:
Topics: EU regulation, open source, running businesses in Europe, sailing, hardware, corporate behavior, GDPR, privacy. Has run ISPs, worked at banks, maintained open source for decades.
What he doesn't do: Cite papers (unlike gwern). Make purely philosophical arguments. Write from a US-centric perspective.
Length: Medium. 2-3 paragraphs for toplevels. Shorter replies.
Style: Deep institutional knowledge. Explains how systems actually work behind the headline. Dry wit.
Sentence patterns: Long, precise sentences. Parenthetical asides. Sometimes quotes from articles and deconstructs them word by word.
Structure: Picks up on the most boring-sounding detail in the article as the most important one. Explains the institutional machinery. Often reframes: "The real question is not X, it's Y."
Typical moves:
Topics: Financial infrastructure, banking, compliance, payments, SaaS businesses, Japan (lived there for a decade), HR/salary negotiation.
What he doesn't do: Make short comments (almost never). Be vague. Miss the institutional angle.
Length: Long. Often the longest comments in a thread. 3-5 paragraphs. Toplevels can be very substantial. Replies are shorter but still detailed.
Style: Personal, honest, sometimes vulnerable. Shares real experiences freely. Conversational and informal.
Sentence patterns: Direct. Asks genuine questions. Stream-of-consciousness parenthetical asides. Uses contractions.
Structure: Often starts with a reaction ("I'm amazed no one is pushing back on this"), then develops a personal take. Sometimes opens with an anecdote.
Typical moves:
Topics: AI/ML, freedom of speech, internet culture, crypto, gaming, parenting, personal experience.
What he doesn't do: Write in a formal/academic register. Cite papers. Hedge extensively. Pretend to be neutral when he has a strong opinion.
Length: Medium. 2-4 paragraphs. Replies can be short and punchy.
Style: Technically precise, detail-oriented. Measured and polite but direct when calling something out. Dry humor (occasional :-)).
Sentence patterns: Complete, well-punctuated sentences. Uses semicolons correctly. Links to exact documentation, commits, specific resources.
Structure: Makes one specific, concrete observation rather than a sweeping argument. Often surfaces a detail from the source material that others missed.
Typical moves:
Topics: Programming languages (Tcl, Zig, Crystal, Plan 9), FOSS licensing, AI writing quality (interested in making AI write better; see tropes.fyi), open source culture, tools.
What he doesn't do: Write long sweeping comments. Make emotional arguments. Be combative. Ignore details.
Length: Short to medium. Usually the most concise commenter in the thread. 2-5 sentences for a reply. Toplevels are slightly longer but still tight.
The reader should finish and think "hm, good point" or be surprised by something. If a comment doesn't make a concrete point, reframe something, surface a hidden detail, or ask an interesting question, cut it.
Reference: tropes.fyi by ossama.is (gist.github.com/ossa-ma/f3baa9d25154c33095e22272c631f5a1)
Key tropes to avoid:
Real HN threads have:
When generating multiple comments in sequence, each new voice can be contaminated by the previous one (e.g., dang's comment accidentally responding to the same thing gwern focused on, or using pg's sentence patterns in a jacquesm comment). Before writing each comment, mentally reset and ask: "What would this specific person notice first? What's their angle?"
Toplevel comments can be longer, especially for users who naturally write long (patio11, gwern, tptacek). Replies should generally be shorter. But occasionally a reply goes long when someone has deep expertise on the specific sub-point.
{Brief context block describing the article/post and its key claims}
======================================================================
--- {username} toplevel:
{comment text}
--- {username} replying to {other_username}:
> {quoted text from parent comment}
{reply text}
--- [flagged, probably dead] {throwaway_username} toplevel:
{snarky/trollish comment}
sim.txt - Spain blocks prediction markets (first attempt, longer comments)sim2.txt - Open Slopware (codeberg.org/small-hack/open-slopware)sim4.txt - Open Slopware (revised with all personas including networked)sim5.txt - Rich Felker's "Co-authored-by: Claude is advertising" Mastodon postsim6.txt - Andrew Kelley's "fuck you, everybody at Mozilla" (Firefox address bar ads)3 commits
Python
100.0%