This repository collects an extensive list of awesome papers about Story Generation / Storytelling, exclusively focusing on the era of Large Language Models (LLMs).
Python
657
126 commits
updated Sep 28, 2026
A curated list of papers on story generation and storytelling in the era of large language models: long-form fiction, screenplays and drama, games, narrative world models, visual stories, and how to evaluate and co-create them. Every paper comes with a one-line summary, and papers we consider essential reading are marked with ๐.
Thank you for the stars! Contributions are very welcome: open an issue or PR for missing papers or mistakes. Contact: mayingpeng33 [AT] gmail [DOT] com
Each paper appears exactly once. Human-centered systems and studies go to Co-creation; work on other media goes to Beyond Text (including its evaluation); remaining work on written stories goes to Evaluation or to the Text Stories topic it mainly addresses. Visual work is included only when it operates at the story level (plot, script, shot planning, narrative reasoning), not when it only improves rendering quality or character consistency. Within a section, papers are sorted by year, with ๐ must-reads first.
| Year | Beyond Text | Text Stories | Evaluation | Co-creation | Total |
|---|---|---|---|---|---|
| 2023 | 16 | 6 | 7 | 5 | 34 |
| 2024 | 13 | 14 | 8 | 7 | 42 |
| 2025 | 24 | 18 | 13 | 7 | 62 |
| 2026* | 31 | 33 | 13 | 15 | 92 |
How to read an entry: venue ยท citation count (refreshed weekly) ยท ๐ must-read ยท title ยท [paper] ยท GitHub stars of the official code, when available, followed by authors and a one-line summary.
Venue colors:
Introduces a 100-environment benchmark on whether LLM narrators keep story commitments under user interventions; even GPT-5.2 survives only 42% after 20 turns.
Introduces an executable environment growing emotional seeds into full interactive story episodes, showing fluent LLMs still fail on robustness and personalization.
Proposes a multi-agent role-play framework whose scene manager selects speakers, switches scenes, and introduces roles, with training data and a benchmark.
Builds HAMLET, a multi-agent framework that turns a topic into a narrative blueprint and performs live embodied theatre with adaptive actor agents.
Proposes Playwriting-guided Generation and Plot-based Reflection to improve player immersion and agency in LLM-based interactive drama.
Introduces a simulation sandbox with character and narrator agents that produces behavior trajectories for fine-grained evaluation of LLM role-playing.
Builds an LLM interactive theater system with director, screenwriter, and actor agents that adapt the plot to player dialogue and object interactions.
Extends the director-actor paradigm so a director agent pursues high-level narrative goals by introducing events, selecting NPCs, and specifying outcomes.
Releases Open-Theatre, an open-source toolkit for LLM interactive drama with multi-agent architecture and hierarchical retrieval-based memory for coherent long-term behavior.
Proposes a plot-progression dataset and method for role-playing agents, detecting an LLM embedding trigger subspace to prompt timely plot advances.
Defines LLM-based interactive drama and trains a drama LLM using Narrative Chain control, Auto-Drama script synthesis, and Sparse Instruction Tuning.
Presents a demo system that lets users role-play a novel character in LLM-generated narrative environments with generated visuals and speech.
Introduces a process benchmark for storytelling in evolving world simulations, finding generation length, canonical consistency, and narrative richness are distinct, competing capacities.
Builds an LLM framework that generates solvable clue-driven investigative 3D game episodes around a deductive solution model guiding characters, clues, and dialogue.
Generates playable interactive fiction worlds in four incremental stages, letting LLMs make creative choices while symbolic validation keeps world state coherent.
Builds a multi-agent RPG game master with a narrative graph and tests six redirection strategies; players prefer in-world redirection over hard denials.
Builds STORY2GAME, which generates a story, populates a world, and writes action code from LLM-derived preconditions and effects for playable interactive fiction.
Builds NarrativeGenie, which turns a designer's story overview into a partially ordered event graph of narrative beats that adapts to player actions.
Builds a system with memory, validation, and a Unity plug-in that keeps LLM-generated RPG content consistent with designer rules despite free-form input.
Builds Word2World, which prompts LLMs to write a story, extract narrative elements, and place tiles to produce playable game worlds without fine-tuning.
Fine-tunes GPT-2 on a released dataset of 978 RPG quests, finding about one in five generated quest descriptions acceptable to players.
Introduces KNUDGE, a dataset from The Outer Worlds requiring lore-faithful, quest-revealing NPC dialogue trees, with supervised and in-context baselines leaving headroom.
Proposes SceneCraft, an LLM framework that automates NPC interaction scenes to unfold authored plot events in narrative-centered games.
Presents 1001 Nights, a game where spoken keywords in co-created LLM tales materialize as in-game items, proposing the notion of AI-native games.
Releases a dataset of about 25,000 real Discord D&D sessions with true game state, showing state information improves LLM game-turn generation.
Proposes an optimization approach that assigns real-world locations to AR story events and synthesizes a navigation graph across story branches.
Proposes a player-centered RPG quest and dialogue generator grounding content in a hand-crafted knowledge base and an LLM, approaching hand-crafted quest quality.
Trains a Dungeon Master model with RL that rewards guidance whose intent matches theory-of-mind predictions of player actions in D&D.
Adds an explicit state-reconstruction and planning interface to a game world model so NPCs act on the game state before being rendered.
Frames novel-to-film generation as building a persistent cinematic world model from prose, then rendering long multi-scene films from it.
Models interactive literary worlds as long-horizon co-evolution of characters and world state, with an open-schema framework and benchmark.
Decouples player control from NPC behavior in a game world model, so text prompts can steer how NPCs react to the player.
Reformulates multi-shot video generation as causal next-shot prediction, letting users steer an unfolding story in real time via streaming prompts.
Builds a generative infinite game in which players raise an autonomous character in an LLM-driven, image-generated world with open-ended, emergent mechanics.
Turns anime film characters into playable agents for open-ended life simulation, predicting multimodal game states to keep the generated world consistent.
Automates the Bechdel test and network analysis on LLM screenplays; human scripts pass more often, but all scripts show some representational bias.
Introduces a multi-horizon audio-drama benchmark showing frontier LLMs degrade over long arcs, plus N-VSSM, a Mamba-2 latent world-state model sustaining consistency.
Builds a hierarchical multi-agent pipeline turning a one-sentence idea into a short drama via debate-based scripting, 3D-grounded first frames, and reviewer loops.
Introduces the task of inferring stage layouts and movements from narrative text, with a dramaturgy-based evaluation suite and rejection-SFT plus GRPO training.
Builds an automated agent population mimicking studio roles to produce sketch comedy videos, using LLM critics aligned with YouTube viewer preferences.
Introduces a drama script continuation benchmark scoring six dimensions via rules and LLM labeling, evaluating eight LLMs on 1,103 scripts.
Decouples screenplay writing into outline-to-prose then prose-to-screenplay stages with hybrid data synthesis, winning 75% against strong baselines per professional screenwriters.
Introduces a movie script benchmark scoring dialogue coherence, character consistency, and plot reasonableness, plus an instruction-based prompting strategy for better scripts.
Proposes IBSEN, where a director agent steers actor agents and human players toward plot objectives to generate controllable drama scripts.
Builds HoLLMwood, a screenwriting framework assigning LLMs Writer, Editor, and role-playing Actor roles to enrich characters and plots in generated screenplays.
Builds Dramatron, which hierarchically prompts LLMs to co-write scripts and screenplays, evaluated in a study with 15 theatre and film professionals.
Presents a character-centric visual storytelling model trained on VIST enriched with visual and textual coreference chains, plus metrics for character richness.
Proposes a visual storytelling MLLM trained with a topic-driven narrative optimizer for data refinement and preference-based ranked story sampling for alignment.
Introduces a manga benchmark of 308 annotated panels showing VLMs interpret single panels well but fail at temporal causality and cross-panel reasoning.
Adapts large multimodal models to visual storytelling on VIST and advocates reference-free metrics RoViST and GROOVIST over BLEU-style evaluation.
Builds a system that converts manga into literary prose for visually impaired readers, introducing the Magiv3 comic-understanding model and annotated panel captions.
Proposes a human-likeness metric over visual grounding, coherence, and repetition, finding a small upgraded TAPM rivals LLaVA, yet good stories need more.
Introduces synchronized video storytelling, generating clip-aligned narrations of fitting length, with the E-SyncVidStory dataset and a storyline-guided VideoNarrator framework.
Proposes a visual storytelling model that extracts visual and linguistic topic information and uses two topic-consistency reinforcement learning rewards on VIST.
Proposes a visual storytelling framework that builds a social-commonsense plot graph from images and derives storylines via weighted shortest paths with Floyd-Warshall.
Introduces a dataset of about 2K curated movie-shot sequences with 12K character-grounded crowdsourced stories, plus a coherence-driven character-based generation baseline.
Proposes a non-autoregressive diffusion model that generates visual story narrations for fictional image sequences with bidirectional history guidance, improving speed and diversity.
Introduces Sound of Story, a dataset of 27K stories pairing image-text sequences with background audio, plus cross-modal retrieval and audio generation benchmarks.
Proposes a modular, interpretable metric for visual storytelling that measures how well stories are grounded in image entities, handling temporal misalignment.
Proposes visual storytelling that feeds images as a visual prefix to a pretrained language model and plans with question-answer blueprints.
Introduces stylized visual storytelling and a memory-augmented multitask model trained with unpaired style text to generate styled stories from photo streams.
Proposes an event-graph reasoning transformer for image-guided story ending generation, with cross-modal fusion, a multimodal injector, and incoherence detection.
Proposes a coherence-theory-inspired self-supervised loss and combined object and face features for character representation, plus a character matching metric for visual storytelling.
Proposes a multi-agent pipeline spanning scripting, shot design, character modeling, keyframes, animation, and audio for long-sequence video storytelling.
Proposes a multi-agent orchestration layer built on FilmDSL, a film-specific language making shot, continuity, and persona constraints explicit for long script-to-video generation.
Represents movies as text and bounding-box or keypoint tokens and curates Storyboard20K, letting LLMs learn movie priors to sample storyboards.
Builds a deployed storyboarding system that learns directing rules from expert examples, evolves them via attribution feedback, and releases the PROSE dataset.
Builds an authoring tool that expands brief story texts into cinematic scripts, then animated videos, guided by a three-layer design framework.
Proposes an agentic story-to-manga framework decomposing creation into planning, grounding, layout, rendering, composition, and lettering for controllable page generation.
Proposes a safety-aware multi-agent framework for end-to-end illustrated storybook generation with page-level text-image calibration and global consistency repair.
Proposes a multi-agent storyboarding framework that plans character, background, and location continuity, plus a new long-range consistency benchmark.
Proposes a multi-agent story visualization framework that grounds roles, extracts causal chains, and verifies consistency to model visual logic explicitly.
Introduces emotion-aware visual story generation and a two-stage framework combining agent-based planning with region-aware generation for emotional, subject-consistent image sequences.
Combines an MLLM narrative planner with a memory-bank control module for long consistent visual sequences and releases a 330K-image e-commerce storyboard dataset.
Proposes a multi-agent plan-execute-verify-revise loop for long audio-visual stories from short prompts, plus a reference-free evaluation protocol.
Proposes SEED-Story, an MLLM generating interleaved text and consistent images for long stories via multimodal attention sinks, and releases the StoryStream dataset.
Defines storytelling image generation with chains of visual reasoning clues and proposes an LLM-plus-text-to-image pipeline with dedicated evaluation metrics.
Proposes an outline-guided pipeline that generates and assembles executable visual novels, using vision-LLM self-correction for cross-modal consistency and script validation.
Uses LLMs to prompt text-to-image models for narrative scene illustration and releases SceneIllustrations, a dataset of pairwise human quality judgments.
Proposes AniMaker, a multi-agent animation framework using MCTS-driven multi-candidate clip generation and the AniEval evaluator to produce story-coherent long videos from text.
Introduces a benchmark annotating commonsense and discourse constraints in visual narratives, with metrics for consistency and text alignment of generated image sequences.
Proposes a multi-agent framework combining LLMs with image, speech, music, and sound tools to generate narrated storybook videos for children.
Introduces visual story ideation, arranging visual assets into storylines, with an MLLM framework using a story graph and a VTravel benchmark.
Introduces multimodal persona-based comic dialogue generation with a 54K-strip dataset and an architecture that generates next-panel dialogues.
Benchmarks outlines from seven long-form generation frameworks with an anchored LLM judge, finding no framework dominates across chapter and book granularities.
Proposes a plug-and-play MCTS planner that builds logic-validated evidence chains for retrieval-based conditional story generation to reduce incoherence and thematic drift.
Proposes PLOTTER, which runs an Evaluate-Plan-Revise cycle on event and character graphs to fix causality and structure before generating full narrative text.
Generates Chinese fiction by writing the climax first, then expanding plot backward and forward with bidirectional MCTS inspired by Freytag's Pyramid.
Proposes a lightweight reasoning projector producing continuous latent tokens that RL policies toggle, cutting reasoning length on plot-hole detection and chapter generation.
Proposes DOME, which fuses planning and writing through dynamic hierarchical outlines and uses a memory module to reduce contradictions in long stories.
Proposes Agents' Room, which splits fiction writing into subtasks for specialized agents, and releases the Tell Me A Story dataset and evaluation.
Proposes a multi-agent long story framework and uses it to build a 6,000-story dataset for fine-tuning Llama3.1-8B and GLM4-9B.
Proposes a plot-planning approach using SVO-triplet plot nodes plus interacting storyline and narrative entity knowledge graph modules for coherent story generation.
Proposes a writing agent that recursively interleaves retrieval, reasoning, and composition tasks instead of fixed outlining, evaluated on fiction and technical reports.
Proposes CogWriter, a training-free framework applying Cognitive Writing Theory via planning, parallel generation, and review agents for constrained long-form text.
Proposes WritingPath, which guides LLMs with explicit outlines reflecting user intent, and builds a blog-post dataset and evaluation framework for goal-oriented writing.
Proposes Ex3, which extracts structure from raw novels to build instruction data, fine-tunes an LLM, and expands tree-like into arbitrarily long novels.
Combines automated planning with LLM text generation, using a planning model as scaffolding to produce more logical, coherent, and believable stories.
Proposes SWAG, framing story writing as search where an auxiliary LLM picks the next action steering the generator toward engaging stories.
Introduces crosslingual story generation with planning and a dataset, finding three-act plans yield more coherent, interesting, controllable stories across languages.
Improves long-story plot coherence by generating a detailed hierarchical outline and a controller that keeps drafted passages aligned with outline details.
Grows a hypergraph world model whose unresolved elements seed latent narratives, improving plot coherence and reducing long-range factual conflicts in story generation.
Proposes a writer-memory system pairing a narratology-typed temporal state graph with hybrid retrieval, outperforming Graphiti/Zep and GraphRAG on multi-hop story questions.
Proposes a training-free scene-by-scene writer that tracks symbolic story states, checks narrative transitions, and uses uncertainty signals to repair inconsistencies.
Proposes Octopus, combining entropy regulation via narrative divergence thresholds with hierarchical memory of characters, plots, and scientific rules for long sci-fi generation.
Proposes SCORE, which tracks key item states and episode summaries and uses retrieval-augmented generation to detect and fix inconsistencies in LLM-generated stories.
Proposes MLD-EA, which uses LLMs with emotion and action cues to detect missing logic in narratives and generate sentences that restore coherence.
Proposes FACTTRACK, which decomposes events into atomic facts with time-aware validity intervals to track world state and detect contradictions in story outlines.
Proposes Temp-Lora, which stores long context in a temporary LoRA module trained during generation, improving long-text quality while cutting context-window costs.
Proposes RecurrentGPT, which simulates LSTM-style recurrence with natural-language long- and short-term memories so LLMs can interactively generate arbitrarily long text.
Replays story worlds from freeze points with and without personas, finding actor LLMs push characters toward cautious, flatter outcomes than canon.
Proposes multi-agent persona-driven story generation with shared world state plus a graph-based hallucination detector, halving hallucinations in 100-page stories.
Proposes a three-layer, perspective-bounded memory for book-based role-playing agents that prevents characters using unknown facts, with a 4,386-question knowledge-boundary benchmark.
Proposes a multi-agent framework with stratified narrative memory and role-location-plot alignment to sustain coherent, open-ended long-horizon story evolution.
Induces executable, interpretable decision trees of validated scene-conditioned behavior rules from narrative data to ground role-playing agents more reliably.
Proposes bottom-up long-form story generation in which multi-agent sandbox simulation yields emergent events that form coherent stories exceeding 10,000 words.
Builds BookWorld, which simulates multi-agent societies from established novels' characters and worldviews to generate creative stories faithful to the source books.
Proposes a cognitive agent framework where tensions between agents' beliefs and ideal worlds drive actions, steering emergent stories toward authored storylines.
Proposes StoryVerse, where authors write abstract acts that LLM narrative planning turns into character actions, balancing authorial intent with emergent game plots.
Builds a story engine encoding McKee's story theory as atomized rules inside an agent harness, improving WritingBench and consistency across four models.
Predicts 304 writing features to score stories against human and AI patterns, turning feature shifts into revision guidance that reduces AI flavor.
Proposes module-wise evolutionary search over premise components like persona, event, and twist, producing more original premises that yield better downstream stories.
Proposes a framework where sub-3B models generate premise-conditioned plots using an aspect reward model, DPO-aligned MoE generator, and cross-family jury evaluation.
Proposes blind peer review among LLM agents that exchange feedback but revise independently, avoiding homogenization, and introduces the SciFi-100 writing dataset.
Proposes generating long narratives by having LLMs stitch mostly verbatim human-written fragments, improving diversity and originality while often evading AI-text detectors.
Proposes a decoding strategy that penalizes concept- and narrative-level similarity to earlier outputs, increasing diversity across multiple story branches from one prompt.
Proposes character-centric story generation that uses text-to-image imagination of story elements and multi-writer persona selection to deepen characters and creativity.
Proposes CritiCS, where a group of LLM critics collectively revise story plans and text to make long stories more creative and expressive.
Proposes MoPS, which composes story premises from modular elements like background and persona, yielding more diverse and original premises for story generation.
Proposes a zero-shot iterative prompting planner grounded in cognitive-psychology and narratology theories of suspense to generate suspenseful stories with LLMs.
Proposes a neuro-symbolic framework that embeds conflict in stories by using commonsense defeasible inference to weaken causal links toward protagonist goals.
Proposes RENarGen, which first generates related opening and closing sentences then infills the middle, producing stories with stronger narrative closure.
Proposes CONCOCT, which trains a concreteness evaluator to guide vaguest-first outline expansion and filtering, yielding more consistent pacing in story outlines.
Proposes a decoding method combining bandit-driven dynamic beam sizing and affect-intensity reranking to generate stories with more interesting twists.
Proposes GROVE, which retrieves human-written story examples and builds an asking-why forest of evidence to add complex, credible plot details.
Proposes a bidirectional pretrained event model with RL using an optimal transport reward to generate coherent stories with flashbacks.
Proposes attribute-guided genre expansion to build a 50K, 13-genre creative writing corpus; fine-tuning on it beats training on story-centric writing data.
Shows that reinforcement learning from narrative-theory-informed AI feedback (d-RLAIF) yields more diverse, convention-aligned stories than supervised fine-tuning.
Proposes a GRPO recipe with LLM-judge rewards and injected human reference stories, letting a 9B model write long stories beyond training length.
Finds by comparing OLMo checkpoints that post-training compresses thematic, affective, and stylistic variation in fiction, most for professional literary text.
Introduces StoryRMB, a benchmark exposing weak reward models for story preferences, and StoryReward, trained on 100K preference pairs for best-of-n story selection.
Proposes a reference-free RL framework with an adaptive constraint-aware generative reward model and ACPO policy optimization, unifying long-form and short-form creative writing.
Proposes an RL framework for creative writing that branches diverse plans in long chain-of-thought and adds a group-aware diversity reward.
Proposes RL for storytelling with a reasoning generative reward model aligned to human creativity judgments and entropy-based reward shaping for training stability.
Introduces imitative novel generation and trains WriterAgent via curriculum learning with hierarchical LoRA modules to mimic an author's style, characters, and plots.
Trains an authorship-verification style judge and uses it as a GRPO reward to fine-tune an 8B model for writing like classic authors.
Proposes RL with a dynamically weighted mix of writing-quality and constraint-verification rewards in GRPO, improving both creative quality and instruction following.
Proposes adaptive curriculum RL for long-form writing with margin-aware data selection, pairwise comparison rewards, and dynamic reference scheduling, beating SFT baselines.
Proposes RL for story reasoning via Next-Chapter Prediction, rewarding plans that raise completion likelihood of real book chapters without labeled data.
Adds deviation from other same-prompt samples into DPO and ORPO objectives, increasing creative writing output diversity with minimal quality loss.
Releases reading preferences from 60 people over creative text pairs, finding tastes diverge and stated preferences poorly predict revealed ones.
Releases 1,665 Chinese creative writing triplets with reverse-engineered prompts and reasoning traces, finding process supervision helps only when mixed with general data.
Introduces Weaver, a 1.8B-34B LLM family pre-trained and aligned for creative and professional writing, with a routing agent balancing quality and cost.
Introduces MirrorStories, 1,500 LLM-generated stories personalized to reader identity, and finds they engage readers more than generic human or LLM stories.
Proposes Pearl, a personalized writing assistant whose retriever is trained to be generation-calibrated, selecting user documents that most improve personalized LLM outputs.
Introduces a benchmark of 2,480 human-labeled story comparisons and 43,827 training pairs for creative writing evaluation, benchmarking LLM judges and reward models.
Introduces a 2,000-prompt benchmark and automated checker for consistency errors in long story generation, analyzing where and which contradictions LLMs make.
Proposes a framework and annotated literary benchmark for narrative orchestration, finding frontier LLMs fail to jointly capture narrative function and structure.
Introduces a benchmark of 300 Chinese novels with human ratings and distilled reader viewpoints, plus CLEM, an 8B evaluator for book-length stories.
Finds LLM internal representations detect incoherent narratives but their ratings do not, and models notice setting violations more than character-trait violations.
Introduces a benchmark of 4,000+ Chinese web novels that scores LLM synopsis-to-story outputs on eight dimensions and ranks them against human-authored percentiles.
Introduces LongStoryEval, 600 books averaging 121K tokens with reader reviews, compares long-story evaluation methods, and trains NovelCritique, an 8B summary-based evaluator.
Introduces FlawedFictions, a benchmark built by synthesizing plot holes in human stories, to test LLM narrative reasoning via plot hole detection.
Introduces LongEval, a benchmark comparing direct and plan-based long-text generation, finding LLMs degrade with length while small long-text-trained models stay competitive.
Proposes a ten-metric macro/meso/micro evaluation framework and bilingual annotated fiction dataset, revealing a high-starting, low-ending pattern in LLM-written novels.
Introduces a benchmark for instruction-following long-form generation at 16K and 32K tokens, finding all tested LLMs struggle as output length grows.
Introduces CollabStory, a dataset of 32k stories co-written by up to five LLMs, with authorship analysis tasks and baselines for multi-LLM writing.
Introduces CS4, a benchmark that measures LLM story creativity by varying the number of prompt constraints to prevent retelling memorized stories.
Introduces StoryWars, 40k collaborative stories from 9,400 authors forming 101 understanding and generation tasks, with an instruction-tuned InstructStory baseline.
Releases 263K stories with TTCW-based review annotations and finds fine-tuning without reasoning traces outperforms reasoning-supervised training for literary review generation.
Introduces 100-Endings, measuring narrative tension by how often repeated ending predictions fail as a story unfolds, and a pipeline that raises tension.
Trains a pairwise story evaluator on self-synthesized, multi-agent-filtered chain-of-thought data and uses it as a reward model to improve story generation.
Finds conflicting evaluations of AI fiction reflect reader differences, clustering 101 annotators into surface-focused and holistic reader profiles via textual feature preferences.
Finds that LLMs rate creative texts more consistently than humans but miss nuanced, culturally specific, and context-dependent aspects of creativity.
Studies LLMs as automatic story evaluators, finding they beat existing metrics at system-level correlation with humans but struggle to explain their ratings.
Proposes character sheet representations built by LLM question-answering and entailment-based fact validation, improving masked-character prediction and measuring character-centricity.
Proposes PerSE, a LLaMA-2 based evaluator that infers reader preferences from in-context profiles to give personalized, interpretable scores for open-ended generation.
Finds that LLMs given the same instructions as human annotators produce story and adversarial-text ratings consistent with expert human evaluation.
Proposes DeltaScore, which evaluates story aspects like fluency and interestingness by measuring likelihood changes under aspect-specific perturbations.
Compares characters in LLM-generated and human-written stories along eight narratological dimensions, examining similarity and variety of character types.
Releases a 350K-story multilingual parallel corpus of LLM children's stories, finding narrative attribute distributions vary substantially across eight languages.
Finds that discourse-level narrative features alone separate human from AI fiction and attribute AI stories to specific models, independent of stylistic cues.
Finds across 28 LLMs that model story continuations carry much lower information-theoretic uncertainty than human writing, worsened by instruction tuning.
Finds LLM stories from one prompt reuse plot element combinations far more than human stories, and proposes an automatic narrative-level diversity metric.
Introduces a dataset exposing gender and cultural stereotypes in LLM children's stories, such as girls receiving more appearance-related attributes than boys.
Finds LLMs reproduce structured Jungian archetypes like the Hero well but struggle with ambiguous ones like the Shadow and Trickster.
Compares stories by 60 LLMs and 60 humans, finding LLMs lag in novelty and surprise though non-experts rate LLM stories more creative.
Finds a fine-tuned BART-large outscores average human writers on short fiction in human ratings, contrasting its linguistic traits with GPT-3.5 and GPT-4o.
Analyzes story arcs, turning points, and affect, finding LLM stories are homogeneously positive and lack tension compared with suspenseful, diverse human narratives.
Proposes the Torrance Test of Creative Writing, finding LLM stories pass 3-10X fewer expert tests than professional stories, and LLM judges misalign.
Stages a contest between novelist Patricio Pron and GPT-4, where expert critics judge the LLM far from a top human fiction author.
Introduces the Psychological Depth Scale for stories' emotional and empathic impact, automates it with LLM personas, finding GPT-4 rivals top Reddit stories.
Compares crowdworker and GPT-3.5/GPT-4 stories on identical Pygmalion prompts, finding AI stories more progressive on gender yet less imaginative.
Compares LLMs and humans on an unusual comic epic prompt, finding top commercial LLMs match humans on most criteria except creativity.
Finds through two roughly 500-participant studies that readers prefer purely LLM-generated stories over human-LLM interleaved stories.
Finds that prompted LLMs write stories rivaling human authors and beating prior generators, though they sometimes replicate real stories.
Trains LLMs with GRPO and a multi-component constructiveness reward to give story-specific feedback, finding actionable suggestions drive constructiveness most.
Builds a graph-based writing assistant for organizing plot points and exploring alternative branches, which writers found reduced structuring effort.
Designs and evaluates a narratology-based fiction writing app with 42 writers, probing auto-evaluators, plan-exposing interfaces, and cultural fit of story structures.
Designs CoNoder, a creator-centered LLM prototype for interactive narratives with node-graph editing, ripple-effect analysis, and simulated reader feedback, informed by creator interviews.
Builds an LLM authoring tool representing interactive narratives as card-based story graphs, found easier for structuring than Twine and AI Dungeon.
Builds a co-writing system with virtual reader reactions and AI attribution, finding transparency raises awareness but lowers creative agency and AI usage.
Builds a writing tool that highlights narrative strategies in example stories and lets novices apply them to drafts via strategy-steered generation.
Builds a multi-persona co-creative storytelling system based on blind variation and selective retention; experts rated co-authored stories more creative.
Builds a mixed-reality system where users author stories by manipulating virtual characters and props, which multi-agent AI turns into rearrangeable narrative beats.
Builds a screenplay refinement system whose AI agent first simulates character experience, then evaluates it to give feedback that deepens screenwriters' reflection.
Builds a tool that fills narrative gaps in video stories by generating context-aware clips that blend stylistically and narratively with captured footage.
Builds an AI authoring system turning text stories into interactive vignettes, using LLM-controlled divergence to keep NPC behavior within the intended story.
Builds an authoring tool that visualizes bundled storylines of LLM-driven interactive narratives, helping authors anticipate player-experienced stories in a 12-user study.
Studies a FigJam plugin combining node-graph story structure, LLM audience impersonation, and image/audio generation for personalized story writing and moral reflection.
Builds an authoring system deriving narrative possibility spaces from example stories, letting authors bound them and unfold them into game events.
Builds a storytelling system where users steer LLM story text by moving character symbols like toys, via a shared motion-text semantic space.
Builds a system letting writers develop characters by conversing with customizable chatbot avatars; a 14-writer study shows it supports iterative character construction.
Fine-tunes an LLM to drive creative or non-creative social robot storytelling, finding the creative robot boosts children's fluency, flexibility, and elaboration.
Finds from 27 writing sessions that deliberately imperfect intermediate AI suggestions encourage writers to rewrite, supporting creative ownership and reflection.
Builds GhostWriter, a writing design probe that implicitly learns user style while offering explicit controls, studying how it supports agency and personalization.
Introduces ID.8, an open-source system for co-creating visual stories with generative AI, with a user study highlighting enjoyment and remaining gaps.
Annotates narrative paragraphs with writing-mode labels and fine-tunes LLMs conditioned on these modes, finding authors prefer mode-controlled suggestions in collaborative fiction writing.
Analyzes how users iteratively revise story prompts in wild chatbot logs, releasing WildStories and WildEdits and an edit-type framework for benchmarking.
Analyzes 500,000 ChatGPT conversations, finding over a third involve fiction generation, dominated by power users favoring fanfiction, erotica, and repetition.
Wizard-of-Oz study of intrusive versus non-intrusive proactive AI suggestions in story outlining, revealing a creativity-agency trade-off moderated by how inspiring suggestions are.
Introduces a task and 1,300 deliberately corrupted stories to evaluate LLM writing feedback, finding models often miss the biggest writing issue.
Examines how co-writing with ChatGPT affects historical fiction writers' character design, plot outlining, and context checking processes.
Interviews 23 screenwriters on how they integrate AI across workflow stages and categorizes expected AI roles as actor, audience, expert, and executor.
Finds with 131 participants a U-shaped effect of AI scaffolding: paragraph-level suggestions improve writing quality and productivity, while sentence-level ones do not.
Interviews 19 professional writers and surveys readers on authenticity in AI co-writing, finding personalization should support writer growth beyond text production.
Finds through an experiment that people will forgo payment for AI writing help, which boosts productivity and confidence but raises ownership concerns.
Interviews 20 creative writers to identify what help they want, how they perceive supporters, and values shaping AI-versus-human support choices.
Studies 30 writers using an LLM interface based on the cognitive process model, finding LLMs most helpful for translating and reviewing.
Elicits parent, teacher, and researcher views on generative AI for children's visual storytelling and proposes AIStory, a prototype app supporting literacy.
Surveys LLM story generation and understanding through narratology, finding generation lags understanding and recommending theory-based metrics over a single quality benchmark.
Surveys LLM story generation, organizing work into autonomous generation versus author assistance and comparing methods, datasets, story types, and evaluations.
Surveys story evaluation across text-to-text, visual-to-text, and text-to-visual tasks, proposing a taxonomy of human criteria, benchmarks, and automatic metrics.
Surveys structured knowledge-enhanced story generation, offering a taxonomy of how knowledge is injected to improve coherence and grounding, plus future directions.
๐ฆ Datasets
๐ Leaderboards
๐๏ธ Venues & Workshops
๐ Related Lists
We welcome paper recommendations and corrections. Please read CONTRIBUTING.md for what the list includes and how to format an entry, then open a paper recommendation issue or a pull request.
If you find this list useful, please consider citing it:
@misc{ma2023awesomestorygeneration,
title = {Awesome-Story-Generation: A Curated List of Papers on Story Generation in the Era of Large Language Models},
author = {Ma, Yingpeng and Ma, Yan},
year = {2023},
howpublished = {\url{https://github.com/yingpengma/Awesome-Story-Generation}}
}
309 followers ยท starred Oct 2024
23 followers ยท starred Apr 2026
Python
100.0%
This repository collects an extensive list of awesome papers about Story Generation / Storytelling, exclusively focusing on the era of Large Language Models (LLMs).
Python
657
126 commits
updated Sep 28, 2026
A curated list of papers on story generation and storytelling in the era of large language models: long-form fiction, screenplays and drama, games, narrative world models, visual stories, and how to evaluate and co-create them. Every paper comes with a one-line summary, and papers we consider essential reading are marked with ๐.
Thank you for the stars! Contributions are very welcome: open an issue or PR for missing papers or mistakes. Contact: mayingpeng33 [AT] gmail [DOT] com
Each paper appears exactly once. Human-centered systems and studies go to Co-creation; work on other media goes to Beyond Text (including its evaluation); remaining work on written stories goes to Evaluation or to the Text Stories topic it mainly addresses. Visual work is included only when it operates at the story level (plot, script, shot planning, narrative reasoning), not when it only improves rendering quality or character consistency. Within a section, papers are sorted by year, with ๐ must-reads first.
| Year | Beyond Text | Text Stories | Evaluation | Co-creation | Total |
|---|---|---|---|---|---|
| 2023 | 16 | 6 | 7 | 5 | 34 |
| 2024 | 13 | 14 | 8 | 7 | 42 |
| 2025 | 24 | 18 | 13 | 7 | 62 |
| 2026* | 31 | 33 | 13 | 15 | 92 |
How to read an entry: venue ยท citation count (refreshed weekly) ยท ๐ must-read ยท title ยท [paper] ยท GitHub stars of the official code, when available, followed by authors and a one-line summary.
Venue colors:
Introduces a 100-environment benchmark on whether LLM narrators keep story commitments under user interventions; even GPT-5.2 survives only 42% after 20 turns.
Introduces an executable environment growing emotional seeds into full interactive story episodes, showing fluent LLMs still fail on robustness and personalization.
Proposes a multi-agent role-play framework whose scene manager selects speakers, switches scenes, and introduces roles, with training data and a benchmark.
Builds HAMLET, a multi-agent framework that turns a topic into a narrative blueprint and performs live embodied theatre with adaptive actor agents.
Proposes Playwriting-guided Generation and Plot-based Reflection to improve player immersion and agency in LLM-based interactive drama.
Introduces a simulation sandbox with character and narrator agents that produces behavior trajectories for fine-grained evaluation of LLM role-playing.
Builds an LLM interactive theater system with director, screenwriter, and actor agents that adapt the plot to player dialogue and object interactions.
Extends the director-actor paradigm so a director agent pursues high-level narrative goals by introducing events, selecting NPCs, and specifying outcomes.
Releases Open-Theatre, an open-source toolkit for LLM interactive drama with multi-agent architecture and hierarchical retrieval-based memory for coherent long-term behavior.
Proposes a plot-progression dataset and method for role-playing agents, detecting an LLM embedding trigger subspace to prompt timely plot advances.
Defines LLM-based interactive drama and trains a drama LLM using Narrative Chain control, Auto-Drama script synthesis, and Sparse Instruction Tuning.
Presents a demo system that lets users role-play a novel character in LLM-generated narrative environments with generated visuals and speech.
Introduces a process benchmark for storytelling in evolving world simulations, finding generation length, canonical consistency, and narrative richness are distinct, competing capacities.
Builds an LLM framework that generates solvable clue-driven investigative 3D game episodes around a deductive solution model guiding characters, clues, and dialogue.
Generates playable interactive fiction worlds in four incremental stages, letting LLMs make creative choices while symbolic validation keeps world state coherent.
Builds a multi-agent RPG game master with a narrative graph and tests six redirection strategies; players prefer in-world redirection over hard denials.
Builds STORY2GAME, which generates a story, populates a world, and writes action code from LLM-derived preconditions and effects for playable interactive fiction.
Builds NarrativeGenie, which turns a designer's story overview into a partially ordered event graph of narrative beats that adapts to player actions.
Builds a system with memory, validation, and a Unity plug-in that keeps LLM-generated RPG content consistent with designer rules despite free-form input.
Builds Word2World, which prompts LLMs to write a story, extract narrative elements, and place tiles to produce playable game worlds without fine-tuning.
Fine-tunes GPT-2 on a released dataset of 978 RPG quests, finding about one in five generated quest descriptions acceptable to players.
Introduces KNUDGE, a dataset from The Outer Worlds requiring lore-faithful, quest-revealing NPC dialogue trees, with supervised and in-context baselines leaving headroom.
Proposes SceneCraft, an LLM framework that automates NPC interaction scenes to unfold authored plot events in narrative-centered games.
Presents 1001 Nights, a game where spoken keywords in co-created LLM tales materialize as in-game items, proposing the notion of AI-native games.
Releases a dataset of about 25,000 real Discord D&D sessions with true game state, showing state information improves LLM game-turn generation.
Proposes an optimization approach that assigns real-world locations to AR story events and synthesizes a navigation graph across story branches.
Proposes a player-centered RPG quest and dialogue generator grounding content in a hand-crafted knowledge base and an LLM, approaching hand-crafted quest quality.
Trains a Dungeon Master model with RL that rewards guidance whose intent matches theory-of-mind predictions of player actions in D&D.
Adds an explicit state-reconstruction and planning interface to a game world model so NPCs act on the game state before being rendered.
Frames novel-to-film generation as building a persistent cinematic world model from prose, then rendering long multi-scene films from it.
Models interactive literary worlds as long-horizon co-evolution of characters and world state, with an open-schema framework and benchmark.
Decouples player control from NPC behavior in a game world model, so text prompts can steer how NPCs react to the player.
Reformulates multi-shot video generation as causal next-shot prediction, letting users steer an unfolding story in real time via streaming prompts.
Builds a generative infinite game in which players raise an autonomous character in an LLM-driven, image-generated world with open-ended, emergent mechanics.
Turns anime film characters into playable agents for open-ended life simulation, predicting multimodal game states to keep the generated world consistent.
Automates the Bechdel test and network analysis on LLM screenplays; human scripts pass more often, but all scripts show some representational bias.
Introduces a multi-horizon audio-drama benchmark showing frontier LLMs degrade over long arcs, plus N-VSSM, a Mamba-2 latent world-state model sustaining consistency.
Builds a hierarchical multi-agent pipeline turning a one-sentence idea into a short drama via debate-based scripting, 3D-grounded first frames, and reviewer loops.
Introduces the task of inferring stage layouts and movements from narrative text, with a dramaturgy-based evaluation suite and rejection-SFT plus GRPO training.
Builds an automated agent population mimicking studio roles to produce sketch comedy videos, using LLM critics aligned with YouTube viewer preferences.
Introduces a drama script continuation benchmark scoring six dimensions via rules and LLM labeling, evaluating eight LLMs on 1,103 scripts.
Decouples screenplay writing into outline-to-prose then prose-to-screenplay stages with hybrid data synthesis, winning 75% against strong baselines per professional screenwriters.
Introduces a movie script benchmark scoring dialogue coherence, character consistency, and plot reasonableness, plus an instruction-based prompting strategy for better scripts.
Proposes IBSEN, where a director agent steers actor agents and human players toward plot objectives to generate controllable drama scripts.
Builds HoLLMwood, a screenwriting framework assigning LLMs Writer, Editor, and role-playing Actor roles to enrich characters and plots in generated screenplays.
Builds Dramatron, which hierarchically prompts LLMs to co-write scripts and screenplays, evaluated in a study with 15 theatre and film professionals.
Presents a character-centric visual storytelling model trained on VIST enriched with visual and textual coreference chains, plus metrics for character richness.
Proposes a visual storytelling MLLM trained with a topic-driven narrative optimizer for data refinement and preference-based ranked story sampling for alignment.
Introduces a manga benchmark of 308 annotated panels showing VLMs interpret single panels well but fail at temporal causality and cross-panel reasoning.
Adapts large multimodal models to visual storytelling on VIST and advocates reference-free metrics RoViST and GROOVIST over BLEU-style evaluation.
Builds a system that converts manga into literary prose for visually impaired readers, introducing the Magiv3 comic-understanding model and annotated panel captions.
Proposes a human-likeness metric over visual grounding, coherence, and repetition, finding a small upgraded TAPM rivals LLaVA, yet good stories need more.
Introduces synchronized video storytelling, generating clip-aligned narrations of fitting length, with the E-SyncVidStory dataset and a storyline-guided VideoNarrator framework.
Proposes a visual storytelling model that extracts visual and linguistic topic information and uses two topic-consistency reinforcement learning rewards on VIST.
Proposes a visual storytelling framework that builds a social-commonsense plot graph from images and derives storylines via weighted shortest paths with Floyd-Warshall.
Introduces a dataset of about 2K curated movie-shot sequences with 12K character-grounded crowdsourced stories, plus a coherence-driven character-based generation baseline.
Proposes a non-autoregressive diffusion model that generates visual story narrations for fictional image sequences with bidirectional history guidance, improving speed and diversity.
Introduces Sound of Story, a dataset of 27K stories pairing image-text sequences with background audio, plus cross-modal retrieval and audio generation benchmarks.
Proposes a modular, interpretable metric for visual storytelling that measures how well stories are grounded in image entities, handling temporal misalignment.
Proposes visual storytelling that feeds images as a visual prefix to a pretrained language model and plans with question-answer blueprints.
Introduces stylized visual storytelling and a memory-augmented multitask model trained with unpaired style text to generate styled stories from photo streams.
Proposes an event-graph reasoning transformer for image-guided story ending generation, with cross-modal fusion, a multimodal injector, and incoherence detection.
Proposes a coherence-theory-inspired self-supervised loss and combined object and face features for character representation, plus a character matching metric for visual storytelling.
Proposes a multi-agent pipeline spanning scripting, shot design, character modeling, keyframes, animation, and audio for long-sequence video storytelling.
Proposes a multi-agent orchestration layer built on FilmDSL, a film-specific language making shot, continuity, and persona constraints explicit for long script-to-video generation.
Represents movies as text and bounding-box or keypoint tokens and curates Storyboard20K, letting LLMs learn movie priors to sample storyboards.
Builds a deployed storyboarding system that learns directing rules from expert examples, evolves them via attribution feedback, and releases the PROSE dataset.
Builds an authoring tool that expands brief story texts into cinematic scripts, then animated videos, guided by a three-layer design framework.
Proposes an agentic story-to-manga framework decomposing creation into planning, grounding, layout, rendering, composition, and lettering for controllable page generation.
Proposes a safety-aware multi-agent framework for end-to-end illustrated storybook generation with page-level text-image calibration and global consistency repair.
Proposes a multi-agent storyboarding framework that plans character, background, and location continuity, plus a new long-range consistency benchmark.
Proposes a multi-agent story visualization framework that grounds roles, extracts causal chains, and verifies consistency to model visual logic explicitly.
Introduces emotion-aware visual story generation and a two-stage framework combining agent-based planning with region-aware generation for emotional, subject-consistent image sequences.
Combines an MLLM narrative planner with a memory-bank control module for long consistent visual sequences and releases a 330K-image e-commerce storyboard dataset.
Proposes a multi-agent plan-execute-verify-revise loop for long audio-visual stories from short prompts, plus a reference-free evaluation protocol.
Proposes SEED-Story, an MLLM generating interleaved text and consistent images for long stories via multimodal attention sinks, and releases the StoryStream dataset.
Defines storytelling image generation with chains of visual reasoning clues and proposes an LLM-plus-text-to-image pipeline with dedicated evaluation metrics.
Proposes an outline-guided pipeline that generates and assembles executable visual novels, using vision-LLM self-correction for cross-modal consistency and script validation.
Uses LLMs to prompt text-to-image models for narrative scene illustration and releases SceneIllustrations, a dataset of pairwise human quality judgments.
Proposes AniMaker, a multi-agent animation framework using MCTS-driven multi-candidate clip generation and the AniEval evaluator to produce story-coherent long videos from text.
Introduces a benchmark annotating commonsense and discourse constraints in visual narratives, with metrics for consistency and text alignment of generated image sequences.
Proposes a multi-agent framework combining LLMs with image, speech, music, and sound tools to generate narrated storybook videos for children.
Introduces visual story ideation, arranging visual assets into storylines, with an MLLM framework using a story graph and a VTravel benchmark.
Introduces multimodal persona-based comic dialogue generation with a 54K-strip dataset and an architecture that generates next-panel dialogues.
Benchmarks outlines from seven long-form generation frameworks with an anchored LLM judge, finding no framework dominates across chapter and book granularities.
Proposes a plug-and-play MCTS planner that builds logic-validated evidence chains for retrieval-based conditional story generation to reduce incoherence and thematic drift.
Proposes PLOTTER, which runs an Evaluate-Plan-Revise cycle on event and character graphs to fix causality and structure before generating full narrative text.
Generates Chinese fiction by writing the climax first, then expanding plot backward and forward with bidirectional MCTS inspired by Freytag's Pyramid.
Proposes a lightweight reasoning projector producing continuous latent tokens that RL policies toggle, cutting reasoning length on plot-hole detection and chapter generation.
Proposes DOME, which fuses planning and writing through dynamic hierarchical outlines and uses a memory module to reduce contradictions in long stories.
Proposes Agents' Room, which splits fiction writing into subtasks for specialized agents, and releases the Tell Me A Story dataset and evaluation.
Proposes a multi-agent long story framework and uses it to build a 6,000-story dataset for fine-tuning Llama3.1-8B and GLM4-9B.
Proposes a plot-planning approach using SVO-triplet plot nodes plus interacting storyline and narrative entity knowledge graph modules for coherent story generation.
Proposes a writing agent that recursively interleaves retrieval, reasoning, and composition tasks instead of fixed outlining, evaluated on fiction and technical reports.
Proposes CogWriter, a training-free framework applying Cognitive Writing Theory via planning, parallel generation, and review agents for constrained long-form text.
Proposes WritingPath, which guides LLMs with explicit outlines reflecting user intent, and builds a blog-post dataset and evaluation framework for goal-oriented writing.
Proposes Ex3, which extracts structure from raw novels to build instruction data, fine-tunes an LLM, and expands tree-like into arbitrarily long novels.
Combines automated planning with LLM text generation, using a planning model as scaffolding to produce more logical, coherent, and believable stories.
Proposes SWAG, framing story writing as search where an auxiliary LLM picks the next action steering the generator toward engaging stories.
Introduces crosslingual story generation with planning and a dataset, finding three-act plans yield more coherent, interesting, controllable stories across languages.
Improves long-story plot coherence by generating a detailed hierarchical outline and a controller that keeps drafted passages aligned with outline details.
Grows a hypergraph world model whose unresolved elements seed latent narratives, improving plot coherence and reducing long-range factual conflicts in story generation.
Proposes a writer-memory system pairing a narratology-typed temporal state graph with hybrid retrieval, outperforming Graphiti/Zep and GraphRAG on multi-hop story questions.
Proposes a training-free scene-by-scene writer that tracks symbolic story states, checks narrative transitions, and uses uncertainty signals to repair inconsistencies.
Proposes Octopus, combining entropy regulation via narrative divergence thresholds with hierarchical memory of characters, plots, and scientific rules for long sci-fi generation.
Proposes SCORE, which tracks key item states and episode summaries and uses retrieval-augmented generation to detect and fix inconsistencies in LLM-generated stories.
Proposes MLD-EA, which uses LLMs with emotion and action cues to detect missing logic in narratives and generate sentences that restore coherence.
Proposes FACTTRACK, which decomposes events into atomic facts with time-aware validity intervals to track world state and detect contradictions in story outlines.
Proposes Temp-Lora, which stores long context in a temporary LoRA module trained during generation, improving long-text quality while cutting context-window costs.
Proposes RecurrentGPT, which simulates LSTM-style recurrence with natural-language long- and short-term memories so LLMs can interactively generate arbitrarily long text.
Replays story worlds from freeze points with and without personas, finding actor LLMs push characters toward cautious, flatter outcomes than canon.
Proposes multi-agent persona-driven story generation with shared world state plus a graph-based hallucination detector, halving hallucinations in 100-page stories.
Proposes a three-layer, perspective-bounded memory for book-based role-playing agents that prevents characters using unknown facts, with a 4,386-question knowledge-boundary benchmark.
Proposes a multi-agent framework with stratified narrative memory and role-location-plot alignment to sustain coherent, open-ended long-horizon story evolution.
Induces executable, interpretable decision trees of validated scene-conditioned behavior rules from narrative data to ground role-playing agents more reliably.
Proposes bottom-up long-form story generation in which multi-agent sandbox simulation yields emergent events that form coherent stories exceeding 10,000 words.
Builds BookWorld, which simulates multi-agent societies from established novels' characters and worldviews to generate creative stories faithful to the source books.
Proposes a cognitive agent framework where tensions between agents' beliefs and ideal worlds drive actions, steering emergent stories toward authored storylines.
Proposes StoryVerse, where authors write abstract acts that LLM narrative planning turns into character actions, balancing authorial intent with emergent game plots.
Builds a story engine encoding McKee's story theory as atomized rules inside an agent harness, improving WritingBench and consistency across four models.
Predicts 304 writing features to score stories against human and AI patterns, turning feature shifts into revision guidance that reduces AI flavor.
Proposes module-wise evolutionary search over premise components like persona, event, and twist, producing more original premises that yield better downstream stories.
Proposes a framework where sub-3B models generate premise-conditioned plots using an aspect reward model, DPO-aligned MoE generator, and cross-family jury evaluation.
Proposes blind peer review among LLM agents that exchange feedback but revise independently, avoiding homogenization, and introduces the SciFi-100 writing dataset.
Proposes generating long narratives by having LLMs stitch mostly verbatim human-written fragments, improving diversity and originality while often evading AI-text detectors.
Proposes a decoding strategy that penalizes concept- and narrative-level similarity to earlier outputs, increasing diversity across multiple story branches from one prompt.
Proposes character-centric story generation that uses text-to-image imagination of story elements and multi-writer persona selection to deepen characters and creativity.
Proposes CritiCS, where a group of LLM critics collectively revise story plans and text to make long stories more creative and expressive.
Proposes MoPS, which composes story premises from modular elements like background and persona, yielding more diverse and original premises for story generation.
Proposes a zero-shot iterative prompting planner grounded in cognitive-psychology and narratology theories of suspense to generate suspenseful stories with LLMs.
Proposes a neuro-symbolic framework that embeds conflict in stories by using commonsense defeasible inference to weaken causal links toward protagonist goals.
Proposes RENarGen, which first generates related opening and closing sentences then infills the middle, producing stories with stronger narrative closure.
Proposes CONCOCT, which trains a concreteness evaluator to guide vaguest-first outline expansion and filtering, yielding more consistent pacing in story outlines.
Proposes a decoding method combining bandit-driven dynamic beam sizing and affect-intensity reranking to generate stories with more interesting twists.
Proposes GROVE, which retrieves human-written story examples and builds an asking-why forest of evidence to add complex, credible plot details.
Proposes a bidirectional pretrained event model with RL using an optimal transport reward to generate coherent stories with flashbacks.
Proposes attribute-guided genre expansion to build a 50K, 13-genre creative writing corpus; fine-tuning on it beats training on story-centric writing data.
Shows that reinforcement learning from narrative-theory-informed AI feedback (d-RLAIF) yields more diverse, convention-aligned stories than supervised fine-tuning.
Proposes a GRPO recipe with LLM-judge rewards and injected human reference stories, letting a 9B model write long stories beyond training length.
Finds by comparing OLMo checkpoints that post-training compresses thematic, affective, and stylistic variation in fiction, most for professional literary text.
Introduces StoryRMB, a benchmark exposing weak reward models for story preferences, and StoryReward, trained on 100K preference pairs for best-of-n story selection.
Proposes a reference-free RL framework with an adaptive constraint-aware generative reward model and ACPO policy optimization, unifying long-form and short-form creative writing.
Proposes an RL framework for creative writing that branches diverse plans in long chain-of-thought and adds a group-aware diversity reward.
Proposes RL for storytelling with a reasoning generative reward model aligned to human creativity judgments and entropy-based reward shaping for training stability.
Introduces imitative novel generation and trains WriterAgent via curriculum learning with hierarchical LoRA modules to mimic an author's style, characters, and plots.
Trains an authorship-verification style judge and uses it as a GRPO reward to fine-tune an 8B model for writing like classic authors.
Proposes RL with a dynamically weighted mix of writing-quality and constraint-verification rewards in GRPO, improving both creative quality and instruction following.
Proposes adaptive curriculum RL for long-form writing with margin-aware data selection, pairwise comparison rewards, and dynamic reference scheduling, beating SFT baselines.
Proposes RL for story reasoning via Next-Chapter Prediction, rewarding plans that raise completion likelihood of real book chapters without labeled data.
Adds deviation from other same-prompt samples into DPO and ORPO objectives, increasing creative writing output diversity with minimal quality loss.
Releases reading preferences from 60 people over creative text pairs, finding tastes diverge and stated preferences poorly predict revealed ones.
Releases 1,665 Chinese creative writing triplets with reverse-engineered prompts and reasoning traces, finding process supervision helps only when mixed with general data.
Introduces Weaver, a 1.8B-34B LLM family pre-trained and aligned for creative and professional writing, with a routing agent balancing quality and cost.
Introduces MirrorStories, 1,500 LLM-generated stories personalized to reader identity, and finds they engage readers more than generic human or LLM stories.
Proposes Pearl, a personalized writing assistant whose retriever is trained to be generation-calibrated, selecting user documents that most improve personalized LLM outputs.
Introduces a benchmark of 2,480 human-labeled story comparisons and 43,827 training pairs for creative writing evaluation, benchmarking LLM judges and reward models.
Introduces a 2,000-prompt benchmark and automated checker for consistency errors in long story generation, analyzing where and which contradictions LLMs make.
Proposes a framework and annotated literary benchmark for narrative orchestration, finding frontier LLMs fail to jointly capture narrative function and structure.
Introduces a benchmark of 300 Chinese novels with human ratings and distilled reader viewpoints, plus CLEM, an 8B evaluator for book-length stories.
Finds LLM internal representations detect incoherent narratives but their ratings do not, and models notice setting violations more than character-trait violations.
Introduces a benchmark of 4,000+ Chinese web novels that scores LLM synopsis-to-story outputs on eight dimensions and ranks them against human-authored percentiles.
Introduces LongStoryEval, 600 books averaging 121K tokens with reader reviews, compares long-story evaluation methods, and trains NovelCritique, an 8B summary-based evaluator.
Introduces FlawedFictions, a benchmark built by synthesizing plot holes in human stories, to test LLM narrative reasoning via plot hole detection.
Introduces LongEval, a benchmark comparing direct and plan-based long-text generation, finding LLMs degrade with length while small long-text-trained models stay competitive.
Proposes a ten-metric macro/meso/micro evaluation framework and bilingual annotated fiction dataset, revealing a high-starting, low-ending pattern in LLM-written novels.
Introduces a benchmark for instruction-following long-form generation at 16K and 32K tokens, finding all tested LLMs struggle as output length grows.
Introduces CollabStory, a dataset of 32k stories co-written by up to five LLMs, with authorship analysis tasks and baselines for multi-LLM writing.
Introduces CS4, a benchmark that measures LLM story creativity by varying the number of prompt constraints to prevent retelling memorized stories.
Introduces StoryWars, 40k collaborative stories from 9,400 authors forming 101 understanding and generation tasks, with an instruction-tuned InstructStory baseline.
Releases 263K stories with TTCW-based review annotations and finds fine-tuning without reasoning traces outperforms reasoning-supervised training for literary review generation.
Introduces 100-Endings, measuring narrative tension by how often repeated ending predictions fail as a story unfolds, and a pipeline that raises tension.
Trains a pairwise story evaluator on self-synthesized, multi-agent-filtered chain-of-thought data and uses it as a reward model to improve story generation.
Finds conflicting evaluations of AI fiction reflect reader differences, clustering 101 annotators into surface-focused and holistic reader profiles via textual feature preferences.
Finds that LLMs rate creative texts more consistently than humans but miss nuanced, culturally specific, and context-dependent aspects of creativity.
Studies LLMs as automatic story evaluators, finding they beat existing metrics at system-level correlation with humans but struggle to explain their ratings.
Proposes character sheet representations built by LLM question-answering and entailment-based fact validation, improving masked-character prediction and measuring character-centricity.
Proposes PerSE, a LLaMA-2 based evaluator that infers reader preferences from in-context profiles to give personalized, interpretable scores for open-ended generation.
Finds that LLMs given the same instructions as human annotators produce story and adversarial-text ratings consistent with expert human evaluation.
Proposes DeltaScore, which evaluates story aspects like fluency and interestingness by measuring likelihood changes under aspect-specific perturbations.
Compares characters in LLM-generated and human-written stories along eight narratological dimensions, examining similarity and variety of character types.
Releases a 350K-story multilingual parallel corpus of LLM children's stories, finding narrative attribute distributions vary substantially across eight languages.
Finds that discourse-level narrative features alone separate human from AI fiction and attribute AI stories to specific models, independent of stylistic cues.
Finds across 28 LLMs that model story continuations carry much lower information-theoretic uncertainty than human writing, worsened by instruction tuning.
Finds LLM stories from one prompt reuse plot element combinations far more than human stories, and proposes an automatic narrative-level diversity metric.
Introduces a dataset exposing gender and cultural stereotypes in LLM children's stories, such as girls receiving more appearance-related attributes than boys.
Finds LLMs reproduce structured Jungian archetypes like the Hero well but struggle with ambiguous ones like the Shadow and Trickster.
Compares stories by 60 LLMs and 60 humans, finding LLMs lag in novelty and surprise though non-experts rate LLM stories more creative.
Finds a fine-tuned BART-large outscores average human writers on short fiction in human ratings, contrasting its linguistic traits with GPT-3.5 and GPT-4o.
Analyzes story arcs, turning points, and affect, finding LLM stories are homogeneously positive and lack tension compared with suspenseful, diverse human narratives.
Proposes the Torrance Test of Creative Writing, finding LLM stories pass 3-10X fewer expert tests than professional stories, and LLM judges misalign.
Stages a contest between novelist Patricio Pron and GPT-4, where expert critics judge the LLM far from a top human fiction author.
Introduces the Psychological Depth Scale for stories' emotional and empathic impact, automates it with LLM personas, finding GPT-4 rivals top Reddit stories.
Compares crowdworker and GPT-3.5/GPT-4 stories on identical Pygmalion prompts, finding AI stories more progressive on gender yet less imaginative.
Compares LLMs and humans on an unusual comic epic prompt, finding top commercial LLMs match humans on most criteria except creativity.
Finds through two roughly 500-participant studies that readers prefer purely LLM-generated stories over human-LLM interleaved stories.
Finds that prompted LLMs write stories rivaling human authors and beating prior generators, though they sometimes replicate real stories.
Trains LLMs with GRPO and a multi-component constructiveness reward to give story-specific feedback, finding actionable suggestions drive constructiveness most.
Builds a graph-based writing assistant for organizing plot points and exploring alternative branches, which writers found reduced structuring effort.
Designs and evaluates a narratology-based fiction writing app with 42 writers, probing auto-evaluators, plan-exposing interfaces, and cultural fit of story structures.
Designs CoNoder, a creator-centered LLM prototype for interactive narratives with node-graph editing, ripple-effect analysis, and simulated reader feedback, informed by creator interviews.
Builds an LLM authoring tool representing interactive narratives as card-based story graphs, found easier for structuring than Twine and AI Dungeon.
Builds a co-writing system with virtual reader reactions and AI attribution, finding transparency raises awareness but lowers creative agency and AI usage.
Builds a writing tool that highlights narrative strategies in example stories and lets novices apply them to drafts via strategy-steered generation.
Builds a multi-persona co-creative storytelling system based on blind variation and selective retention; experts rated co-authored stories more creative.
Builds a mixed-reality system where users author stories by manipulating virtual characters and props, which multi-agent AI turns into rearrangeable narrative beats.
Builds a screenplay refinement system whose AI agent first simulates character experience, then evaluates it to give feedback that deepens screenwriters' reflection.
Builds a tool that fills narrative gaps in video stories by generating context-aware clips that blend stylistically and narratively with captured footage.
Builds an AI authoring system turning text stories into interactive vignettes, using LLM-controlled divergence to keep NPC behavior within the intended story.
Builds an authoring tool that visualizes bundled storylines of LLM-driven interactive narratives, helping authors anticipate player-experienced stories in a 12-user study.
Studies a FigJam plugin combining node-graph story structure, LLM audience impersonation, and image/audio generation for personalized story writing and moral reflection.
Builds an authoring system deriving narrative possibility spaces from example stories, letting authors bound them and unfold them into game events.
Builds a storytelling system where users steer LLM story text by moving character symbols like toys, via a shared motion-text semantic space.
Builds a system letting writers develop characters by conversing with customizable chatbot avatars; a 14-writer study shows it supports iterative character construction.
Fine-tunes an LLM to drive creative or non-creative social robot storytelling, finding the creative robot boosts children's fluency, flexibility, and elaboration.
Finds from 27 writing sessions that deliberately imperfect intermediate AI suggestions encourage writers to rewrite, supporting creative ownership and reflection.
Builds GhostWriter, a writing design probe that implicitly learns user style while offering explicit controls, studying how it supports agency and personalization.
Introduces ID.8, an open-source system for co-creating visual stories with generative AI, with a user study highlighting enjoyment and remaining gaps.
Annotates narrative paragraphs with writing-mode labels and fine-tunes LLMs conditioned on these modes, finding authors prefer mode-controlled suggestions in collaborative fiction writing.
Analyzes how users iteratively revise story prompts in wild chatbot logs, releasing WildStories and WildEdits and an edit-type framework for benchmarking.
Analyzes 500,000 ChatGPT conversations, finding over a third involve fiction generation, dominated by power users favoring fanfiction, erotica, and repetition.
Wizard-of-Oz study of intrusive versus non-intrusive proactive AI suggestions in story outlining, revealing a creativity-agency trade-off moderated by how inspiring suggestions are.
Introduces a task and 1,300 deliberately corrupted stories to evaluate LLM writing feedback, finding models often miss the biggest writing issue.
Examines how co-writing with ChatGPT affects historical fiction writers' character design, plot outlining, and context checking processes.
Interviews 23 screenwriters on how they integrate AI across workflow stages and categorizes expected AI roles as actor, audience, expert, and executor.
Finds with 131 participants a U-shaped effect of AI scaffolding: paragraph-level suggestions improve writing quality and productivity, while sentence-level ones do not.
Interviews 19 professional writers and surveys readers on authenticity in AI co-writing, finding personalization should support writer growth beyond text production.
Finds through an experiment that people will forgo payment for AI writing help, which boosts productivity and confidence but raises ownership concerns.
Interviews 20 creative writers to identify what help they want, how they perceive supporters, and values shaping AI-versus-human support choices.
Studies 30 writers using an LLM interface based on the cognitive process model, finding LLMs most helpful for translating and reviewing.
Elicits parent, teacher, and researcher views on generative AI for children's visual storytelling and proposes AIStory, a prototype app supporting literacy.
Surveys LLM story generation and understanding through narratology, finding generation lags understanding and recommending theory-based metrics over a single quality benchmark.
Surveys LLM story generation, organizing work into autonomous generation versus author assistance and comparing methods, datasets, story types, and evaluations.
Surveys story evaluation across text-to-text, visual-to-text, and text-to-visual tasks, proposing a taxonomy of human criteria, benchmarks, and automatic metrics.
Surveys structured knowledge-enhanced story generation, offering a taxonomy of how knowledge is injected to improve coherence and grounding, plus future directions.
๐ฆ Datasets
๐ Leaderboards
๐๏ธ Venues & Workshops
๐ Related Lists
We welcome paper recommendations and corrections. Please read CONTRIBUTING.md for what the list includes and how to format an entry, then open a paper recommendation issue or a pull request.
If you find this list useful, please consider citing it:
@misc{ma2023awesomestorygeneration,
title = {Awesome-Story-Generation: A Curated List of Papers on Story Generation in the Era of Large Language Models},
author = {Ma, Yingpeng and Ma, Yan},
year = {2023},
howpublished = {\url{https://github.com/yingpengma/Awesome-Story-Generation}}
}
309 followers ยท starred Oct 2024
23 followers ยท starred Apr 2026
Python
100.0%