yingpengma/Awesome-Story-Generation

This repository collects an extensive list of awesome papers about Story Generation / Storytelling, exclusively focusing on the era of Large Language Models (LLMs).

Python

657

126 commits

updated Sep 28, 2026

See the code

README

Awesome Story Generation

Awesome Papers Last commit Stars PRs welcome License

News Overview Papers Contributing Citation

Maintained by Yingpeng Ma and Yan Ma

A curated list of papers on story generation and storytelling in the era of large language models: long-form fiction, screenplays and drama, games, narrative world models, visual stories, and how to evaluate and co-create them. Every paper comes with a one-line summary, and papers we consider essential reading are marked with ๐ŸŒŸ.

Thank you for the stars! Contributions are very welcome: open an issue or PR for missing papers or mistakes. Contact: mayingpeng33 [AT] gmail [DOT] com

๐Ÿ“ฐ News

  • [2026-09] ๐ŸŽ‰ Major update: 234 papers, a new taxonomy, one-line summaries and ๐ŸŒŸ must-read picks.
  • [2026-05] ๐Ÿ”ฅ Our paper on long-horizon consistency in interactive narratives is accepted to ICML 2026! See it here.

๐Ÿ—บ๏ธ Overview

The list has four sections: Beyond Text (84 papers), Text Stories (71), Evaluation (41) and Co-creation (34), plus 4 surveys.
  • Beyond Text: interactive drama, games, narrative world models, screenplays, and visual stories.
  • Text Stories: planning, coherence, characters, creativity, and training for written stories.
  • Evaluation: benchmarks, metrics, and analyses of written stories.
  • Co-creation: tools for creators, and studies of how people write with AI.
  • Surveys: overviews of the whole field.

Each paper appears exactly once. Human-centered systems and studies go to Co-creation; work on other media goes to Beyond Text (including its evaluation); remaining work on written stories goes to Evaluation or to the Text Stories topic it mainly addresses. Visual work is included only when it operates at the story level (plot, script, shot planning, narrative reasoning), not when it only improves rendering quality or character consistency. Within a section, papers are sorted by year, with ๐ŸŒŸ must-reads first.

Papers per year: 34 in 2023, 42 in 2024, 62 in 2025 and 92 in 2026 through September, excluding 4 surveys.

Data table for the chart (2026 counts through September; the 4 surveys are not shown)
YearBeyond TextText StoriesEvaluationCo-creationTotal
20231667534
202413148742
2025241813762
2026*3133131592

๐Ÿ“‘ Table of Contents

๐Ÿ“„ Papers

How to read an entry: venue ยท citation count (refreshed weekly) ยท ๐ŸŒŸ must-read ยท title ยท [paper] ยท GitHub stars of the official code, when available, followed by authors and a one-line summary.

Venue colors: NLP ML Vision & Graphics AI HCI Games arXiv Other

๐ŸŽญ Beyond Text

๐ŸŽช Interactive Drama

  • ICML 2026 ๐ŸŒŸ Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives [paper] GitHub stars
    Yingpeng Ma, Jianhao Yan, Bei-Ning Shi, Karim Kam, Runnan Wang, Xue-Bo Liu, Yulong Chen, Yue Zhang, Derek F. Wong

    Introduces a 100-environment benchmark on whether LLM narrators keep story commitments under user interventions; even GPT-5.2 survives only 42% after 20 turns.

  • ArXiv 2026 NARRA-Gym for Evaluating Interactive Narrative Agents [paper]
    Yue Huang, Yu-Chen Ma, Jiayi Ye, Wen-Jie Wang, Zi-Peng Ling, Xing Hu, Yuexing Hao, Zi-Chen Chen, Zhangchen Xu, Yun-Hong He, et al.

    Introduces an executable environment growing emotional seeds into full interactive story episodes, showing fluent LLMs still fail on robustness and personalization.

  • ACL Findings 2026 AdaMARP: An Adaptive Multi-Agent Interaction Framework for General Immersive Role-Playing [paper]
    Zhenhua Xu, Dongsheng Chen, Shuo Wang, Jian Li, Chengjie Wang, Meng Han, Ya-Biao Wang

    Proposes a multi-agent role-play framework whose scene manager selects speakers, switches scenes, and introduces roles, with training data and a benchmark.

  • ICLR 2026 HAMLET: A Hierarchical and Adaptive Multi-Agent Framework for Live Embodied Theatrics [paper] GitHub stars [dataset]
    Shu-Fan Jiang, Si-Zhou Chen, Chios Chen, Chi Zhang, Xiao-Lei Zhang, Xue-Long Li

    Builds HAMLET, a multi-agent framework that turns a topic into a narrative blueprint and performs live embodied theatre with adaptive actor agents.

  • ACL 2025 ๐ŸŒŸ Towards Enhanced Immersion and Agency for LLM-based Interactive Drama [paper] GitHub stars
    Hongqiu Wu, Weiqi Wu, Tianyang Xu, Jiameng Zhang, Hai Zhao

    Proposes Playwriting-guided Generation and Plot-based Reflection to improve player immersion and agency in LLM-based interactive drama.

  • NAACL 2025 ๐ŸŒŸ CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds [paper] GitHub stars
    Lei Wang, Jian-Xun Lian, Yi Huang, Yanqi Dai, Haoxuan Li, Xu Chen, Xing Xie, Ji-Rong Wen

    Introduces a simulation sandbox with character and narrator agents that produces behavior trajectories for fine-grained evaluation of LLM role-playing.

  • Complex & Intelligent Systems 2025 ProTriPlay: A trinity framework for professional interactive theater based on LLM [paper]
    Yinglong Yu, Hao Shen, Ming Yang, Yu Wang, Yanyu Liu

    Builds an LLM interactive theater system with director, screenwriter, and actor agents that adapt the plot to player dialogue and object interactions.

  • AIIDE 2025 CoDi: A Director-Actor Framework for Goal-Driven Interactive Story Generation with LLMs [paper]
    Honggu Kim, Taewoo Yoo, Yun-Gyung Cheong

    Extends the director-actor paradigm so a director agent pursues high-level narrative goals by introducing events, selecting NPCs, and specifying outcomes.

  • EMNLP 2025 OPEN-THEATRE: An Open-Source Toolkit for LLM-based Interactive Drama [paper]
    Tianyang Xu, Hongqiu Wu, Weiqi Wu, Hai Zhao

    Releases Open-Theatre, an open-source toolkit for LLM interactive drama with multi-agent architecture and hierarchical retrieval-based memory for coherent long-term behavior.

  • ACL 2025 RolePlot: A Systematic Framework for Evaluating and Enhancing the Plot-Progression Capabilities of Role-Playing Agents [paper]
    Pinyi Zhang, Si-Yu An, Lingfeng Qiao, Yi-Fei Yu, Jing-Yang Chen, Jie Wang, Di Yin, Xing Sun, Kai Zhang

    Proposes a plot-progression dataset and method for role-playing agents, detecting an LLM embedding trigger subspace to prompt timely plot advances.

  • ACL Findings 2024 ๐ŸŒŸ From Role-Play to Drama-Interaction: An LLM Solution [paper]
    Weiqi Wu, Hongqiu Wu, Lai Jiang, Xing-Chen Liu, Jiale Hong, Haizhen Zhao, Min Zhang

    Defines LLM-based interactive drama and trains a drama LLM using Narrative Chain control, Auto-Drama script synthesis, and Sparse Instruction Tuning.

  • AAAI 2024 NarrativePlay: An Automated System for Crafting Visual Worlds in Novels for Role-Playing [paper]
    Run-Cong Zhao, Wenjia Zhang, Jiazheng Li, Lixing Zhu, Yanran Li, Yulan He, Lin Gui

    Presents a demo system that lets users role-play a novel character in LLM-generated narrative environments with generated visuals and speech.

๐ŸŽฒ Games

  • ArXiv 2026 When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations [paper]
    Yuqi Chen, Sixuan Li, Yunfeng Cai, Xueai Li, Kaiwen Yan, Ying Li

    Introduces a process benchmark for storytelling in evolving world simulations, finding generation length, canonical consistency, and narrative richness are distinct, competing capacities.

  • FDG 2026 Generating Clue-Driven Investigative Game Narratives with Large Language Models [paper]
    Vikram Kumaran, A. Smith, Wookhee Min, Randall Spain, Bradford W. Mott, James C. Lester

    Builds an LLM framework that generates solvable clue-driven investigative 3D game episodes around a deductive solution model guiding characters, clues, and dialogue.

  • ArXiv 2026 IVIE: A Neuro-symbolic Approach to Incremental and Validated Generation of Interactive Fiction Worlds [paper]
    Micaela Vaucher, Santiago Silveira, Santiago Gรณngora, Luis Chiruzzo

    Generates playable interactive fiction worlds in four incremental stages, letting LLMs make creative choices while symbolic validation keeps world state coherent.

  • IUI 2026 Guiding, Not Railroading: Design and Evaluation of a Multi-Agent System for Narrative Redirection in Role-playing Games [paper]
    Nicolai Hejlesen Jรธrgensen, Sarmilan Tharmabalan, Ilhan Aslan, Nicolai Brodersen Hansen, Timothy Merritt

    Builds a multi-agent RPG game master with a narrative graph and tests six redirection strategies; players prefer in-world redirection over hard denials.

  • ArXiv 2025 STORY2GAME: Generating (Almost) Everything in an Interactive Fiction Game [paper]
    E. Zhou, Shreyas Basavatia, M. Siam, Zexin Chen, Mark O. Riedl

    Builds STORY2GAME, which generates a story, populates a world, and writes action code from LLM-derived preconditions and effects for playable interactive fiction.

  • AIIDE 2024 NarrativeGenie: Generating Narrative Beats and Dynamic Storytelling with Large Language Models [paper]
    Vikram Kumaran, Jonathan Rowe, James C. Lester

    Builds NarrativeGenie, which turns a designer's story overview into a partially ordered event graph of narrative beats that adapts to player actions.

  • AIIDE 2024 PANGeA: Procedural Artificial Narrative Using Generative AI for Turn-Based, Role-Playing Video Games [paper]
    Stephanie Buongiorno, Lawrence J. Klinkert, Zixin Zhuang, Tanishq Chawla, Corey Clark

    Builds a system with memory, validation, and a Unity plug-in that keeps LLM-generated RPG content consistent with designer rules despite free-form input.

  • ArXiv 2024 Word2World: Generating Stories and Worlds through Large Language Models [paper] GitHub stars
    Muhammad Umair Nasir, Steven James, Julian Togelius

    Builds Word2World, which prompts LLMs to write a story, extract narrative elements, and place tiles to produce playable game worlds without fine-tuning.

  • IEEE ToG 2024 Generating Role-Playing Game Quests With GPT Language Models [paper]
    Susanna Vรคrtinen, Perttu Hรคmรคlรคinen, C. Guckelsberger

    Fine-tunes GPT-2 on a released dataset of 978 RPG quests, finding about one in five generated quest descriptions acceptable to players.

  • EMNLP 2024 Ontologically Faithful Generation of Non-Player Character Dialogues [paper]
    Nathaniel Weir, Ryan Thomas, Randolph D'Amore, Kellie Hill, Benjamin Van Durme, Harsh Jhamtani

    Introduces KNUDGE, a dataset from The Outer Worlds requiring lore-faithful, quest-revealing NPC dialogue trees, with supervised and in-context baselines leaving headroom.

  • AIIDE 2023 ๐ŸŒŸ SceneCraft: Automating Interactive Narrative Scene Generation in Digital Games with Large Language Models [paper]
    Vikram Kumaran, Jonathan Rowe, Bradford W. Mott, James C. Lester

    Proposes SceneCraft, an LLM framework that automates NPC interaction scenes to unfold authored plot events in narrative-centered games.

  • AIIDE 2023 ๐ŸŒŸ Language as Reality: A Co-Creative Storytelling Game Experience in 1001 Nights using Generative AI [paper]
    Yuqian Sun, Zhouyi Li, Ke Fang, Chang Hee Lee, A. Asadipour

    Presents 1001 Nights, a game where spoken keywords in co-created LLM tales materialize as in-game items, proposing the notion of AI-native games.

  • ACL 2023 ๐ŸŒŸ FIREBALL: A Dataset of Dungeons and Dragons Actual-Play with Structured Game State Information [paper] GitHub stars
    Andrew Zhu, Karmanya Aggarwal, Alexander H. Feng, Lara J. Martin, Chris Callison-Burch

    Releases a dataset of about 25,000 real Discord D&D sessions with true game state, showing state information improves LLM game-turn generation.

  • CHI 2023 Location-Aware Adaptation of Augmented Reality Narratives [paper]
    Wan-Wan Li, Changyang Li, Minyoung Kim, Haikun Huang, L. Yu

    Proposes an optimization approach that assigns real-world locations to AR story events and synthesizes a navigation graph across story branches.

  • CHI 2023 Personalized Quest and Dialogue Generation in Role-Playing Games: A Knowledge Graph- and Language Model-based Approach [paper]
    Trevor Ashby, Braden K Webb, G. Knapp, John Searle, Nancy Fulda

    Proposes a player-centered RPG quest and dialogue generator grounding content in a hand-crafted knowledge base and an LLM, approaching hand-crafted quest quality.

  • ACL 2023 I Cast Detect Thoughts: Learning to Converse and Guide with Intents and Theory-of-Mind in Dungeons and Dragons [paper]
    Pei Zhou, Andrew Zhu, Jennifer Hu, J. Pujara, Xiang Ren, Chris Callison-Burch, Yejin Choi, Prithviraj Ammanabrolu

    Trains a Dungeon Master model with RL that rewards guidance whose intent matches theory-of-mind predictions of player actions in D&D.

๐ŸŒ World Models

  • ArXiv 2026 WorldMind: Decoupled Game World Model for State-Aware NPC Behavior [paper] GitHub stars
    Zhi-Yang Deng, Bo-Ran Zhang, Dan Chen, Ye-Ying Jin

    Adds an explicit state-reconstruction and planning interface to a game world model so NPCs act on the game state before being rendered.

  • ArXiv 2026 FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling [paper]
    Jia-Long Zuo, Haotong Zuo, Shiwei Zhang, Xiang Wang, Chen Li, Nong Sang, Chang-Xin Gao, Xiang Bai

    Frames novel-to-film generation as building a persistent cinematic world model from prose, then rendering long multi-scene films from it.

  • ArXiv 2026 EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World [paper] GitHub stars
    Qing Zong, Yue (Sophie) Guo, Mengxi Yang, Yiwen Guo, Yangqiu Song

    Models interactive literary worlds as long-horizon co-evolution of characters and world state, with an open-schema framework and benchmark.

  • ArXiv 2026 ReactiveGWM: Steering NPC in Reactive Game World Models [paper] GitHub stars
    Zeqing Wang, Dan Chen, Zhaohu Xing, Zizhao Tong, Yinhan Zhang, Xingyi Yang, Ye-Ying Jin

    Decouples player control from NPC behavior in a game world model, so text prompts can steer how NPCs react to the player.

  • ArXiv 2026 ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling [paper] GitHub stars
    Yawen Luo, Xiao-Yu Shi, Junhao Zhuang, Yu-Tian Chen, Quande Liu, Xintao Wang, Pengfei Wan, Tian-Fan Xue

    Reformulates multi-shot video generation as causal next-shot prediction, letting users steer an unfolding story in real time via streaming prompts.

  • ICLR 2025 ๐ŸŒŸ Unbounded: A Generative Infinite Game of Character Life Simulation [paper]
    Jialu Li, Yuanzhen Li, Neal Wadhwa, Y. Pritch, David E. Jacobs, Michael Rubinstein, Mohit Bansal, Nataniel Ruiz

    Builds a generative infinite game in which players raise an autonomous character in an LLM-driven, image-generated world with open-ended, emergent mechanics.

  • ICCV 2025 AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction [paper] GitHub stars
    Junhao Cheng, Yu-Ying Ge, Yi-Xiao Ge, Jing Liao, Shan Ying

    Turns anime film characters into playable agents for open-ended life simulation, predicting multimodal game states to keep the generated world consistent.

๐ŸŽž๏ธ Screenplays

  • FAccT 2026 Do Language Models Pass the Bechdel Test? Auditing Gender Biases in LLM-Generated Screenplays [paper]
    Megha N. Govindu, Stephanie T. Wang, Sorelle A. Friedler, D. Metaxa

    Automates the Bechdel test and network analysis on LLM screenplays; human scripts pass more often, but all scripts show some representational bias.

  • ArXiv 2026 NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama [paper]
    Logan Mann, Abdur Rahman, M. Saifullah, Taaha Kazi, Vasu Sharma

    Introduces a multi-horizon audio-drama benchmark showing frontier LLMs degrade over long arcs, plus N-VSSM, a Mamba-2 latent world-state model sustaining consistency.

  • ArXiv 2026 One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems [paper]
    Yu-Fei Shi, Wei-Long Yan, Naixuan Huang, Yucheng Chen, Chenyu Zhang, Tao He, Si Yong Yeo, Ming Li

    Builds a hierarchical multi-agent pipeline turning a one-sentence idea into a short drama via debate-based scripting, 3D-grounded first frames, and reviewer loops.

  • ArXiv 2026 Text-to-Stage: Spatial Layouts from Long-form Narratives [paper]
    Jefferson Hernandez, Swarnadeep Saha, Chenxi Whitehouse, Sanjeel Parekh, Calvin Murdock, Yuliang Li, W. O. Brimijoin, V. Ithapu, I. Ananthabhotla

    Introduces the task of inferring stage layouts and movements from narrative text, with a dramaturgy-based evaluation suite and rejection-SFT plus GRPO training.

  • ArXiv 2026 COMIC: Agentic Sketch Comedy Generation [paper] GitHub stars
    Susung Hong, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz

    Builds an automated agent population mimicking studio roles to produce sketch comedy videos, using LLM critics aligned with YouTube viewer preferences.

  • ArXiv 2025 DramaBench: A Six-Dimensional Evaluation Framework for Drama Script Continuation [paper] GitHub stars
    Shijian Ma, Yun-Chien Huang, Yan Lin

    Introduces a drama script continuation benchmark scoring six dimensions via rules and LLM labeling, evaluating eight LLMs on 1,103 scripts.

  • ArXiv 2025 Beyond Direct Generation: A Decomposed Approach to Well-Crafted Screenwriting with LLMs [paper]
    Hang Lei, Shengyi Zong, Zhaoyan Li, Ziren Zhou, Hao Liu, Liang Yu

    Decouples screenplay writing into outline-to-prose then prose-to-screenplay stages with hybrid data synthesis, winning 75% against strong baselines per professional screenwriters.

  • ArXiv 2025 CML-Bench: A Framework for Evaluating and Enhancing LLM-Powered Movie Scripts Generation [paper]
    Mingzhe Zheng, Dingjie Song, Guanyu Zhou, Jun You, Jia-Hao Zhan, Xuran Ma, Xin-Yuan Song, Ser-Nam Lim, Qi-Feng Chen, Harry Yang

    Introduces a movie script benchmark scoring dialogue coherence, character consistency, and plot reasonableness, plus an instruction-based prompting strategy for better scripts.

  • ACL 2024 ๐ŸŒŸ IBSEN: Director-Actor Agent Collaboration for Controllable and Interactive Drama Script Generation [paper] GitHub stars
    Senyu Han, Lu Chen, Li-Min Lin, Zhen Xu, Kai Yu

    Proposes IBSEN, where a director agent steers actor agents and human players toward plot objectives to generate controllable drama scripts.

  • EMNLP Findings 2024 ๐ŸŒŸ HoLLMwood: Unleashing the Creativity of Large Language Models in Screenwriting via Role Playing [paper]
    Jing Chen, Xinyu Zhu, Cheng Yang, Chufan Shi, Ya-Dong Xi, Yuxiang Zhang, Junjie Wang, Jiashu Pu, Rongsheng Zhang, Yu-Jiu Yang, et al.

    Builds HoLLMwood, a screenwriting framework assigning LLMs Writer, Editor, and role-playing Actor roles to enrich characters and plots in generated screenplays.

  • CHI 2023 ๐ŸŒŸ Co-Writing Screenplays and Theatre Scripts with Language Models: An Evaluation by Industry Professionals [paper]
    Piotr Wojciech Mirowski, K. Mathewson, Jaylen Pittman, Richard Evans

    Builds Dramatron, which hierarchically prompts LLMs to co-write scripts and screenplays, evaluated in a study with 15 theatre and film professionals.

๐Ÿ–ผ๏ธ Visual2Story

  • TACL 2026 Generating Visual Stories with Grounded and Coreferent Characters [paper] GitHub stars
    Danyang Liu, Mirella Lapata, Frank Keller

    Presents a character-centric visual storytelling model trained on VIST enriched with visual and textual coreference chains, plus metrics for character richness.

  • COLING 2025 ๐ŸŒŸ StoryLLaVA: Enhancing Visual Storytelling with Multi-Modal Large Language Models [paper]
    Li Yang, Zhihao Xiao, Wen-Xin Huang, Xian Zhong

    Proposes a visual storytelling MLLM trained with a topic-driven narrative optimizer for data refinement and preference-based ranked story sampling for alignment.

  • ICCV AISTORY Workshop 2025 Re:Verse -- Can Your VLM Read a Manga? [paper] GitHub stars
    Aaditya Baranwal, Madhav Kataria, Naitik Agarwal, Y. Rawat, Shruti Vyas

    Introduces a manga benchmark of 308 annotated panels showing VLMs interpret single panels well but fail at temporal causality and cross-panel reasoning.

  • ArXiv 2025 VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? [paper]
    M. Gado, Towhid Taliee, M. Memon, Dmitry Ignatov, R. Timofte

    Adapts large multimodal models to visual storytelling on VIST and advocates reference-free metrics RoViST and GROOVIST over BLEU-style evaluation.

  • ICCV 2025 From Panels to Prose: Generating Literary Narratives from Comics [paper] GitHub stars
    Ragav Sachdeva, Andrew Zisserman

    Builds a system that converts manga into literary prose for visually impaired readers, introducing the Magiv3 comic-understanding model and annotated panel captions.

  • EMNLP Findings 2024 Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition [paper] GitHub stars
    Aditya K Surikuchi, Raquel Fernรกndez, Sandro Pezzelle

    Proposes a human-likeness metric over visual grounding, coherence, and repetition, finding a small upgraded TAPM rivals LLaVA, yet good stories need more.

  • ACL 2024 Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline [paper]
    Dingyi Yang, Chunru Zhan, Ziheng Wang, Biao Wang, Tiezheng Ge, Bo Zheng, Qin Jin

    Introduces synchronized video storytelling, generating clip-aligned narrations of fitting length, with the E-SyncVidStory dataset and a storyline-guided VideoNarrator framework.

  • LREC-COLING 2024 TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling [paper]
    Wei-Ran Chen, Xin Li, Jiaqi Su, Guiqian Zhu, Ying Li, Yi Ji, Chunping Liu

    Proposes a visual storytelling model that extracts visual and linguistic topic information and uses two topic-consistency reinforcement learning rewards on VIST.

  • EACL 2024 SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling [paper]
    E. Wang, Caren Han, Josiah Poon

    Proposes a visual storytelling framework that builds a social-commonsense plot graph from images and derives storylines via weighted shortest paths with Floyd-Warshall.

  • TACL 2023 ๐ŸŒŸ Visual Writing Prompts: Character-Grounded Story Generation with Curated Image Sequences [paper]
    Xudong Hong, A. Sayeed, K. Mehra, Vera Demberg, B. Schiele

    Introduces a dataset of about 2K curated movie-shot sequences with 12K character-grounded crowdsourced stories, plus a coherence-driven character-based generation baseline.

  • EMNLP Findings 2023 DiffuVST: Narrating Fictional Scenes with Global-History-Guided Denoising Models [paper]
    Shengguang Wu, Mei Yuan, Qi Su

    Proposes a non-autoregressive diffusion model that generates visual story narrations for fictional image sequences with bidirectional history guidance, improving speed and diversity.

  • EMNLP Findings 2023 Sound of Story: Multi-modal Storytelling with Audio [paper]
    Jaeyeon Bae, Seokhoon Jeong, Seokun Kang, Namgi Han, Jae-Yon Lee, Hyounghun Kim, Taehwan Kim

    Introduces Sound of Story, a dataset of 27K stories pairing image-text sequences with background audio, plus cross-modal retrieval and audio generation benchmarks.

  • EMNLP 2023 GROOViST: A Metric for Grounding Objects in Visual Storytelling [paper] GitHub stars
    Aditya K Surikuchi, Sandro Pezzelle, Raquel Fernรกndez

    Proposes a modular, interpretable metric for visual storytelling that measures how well stories are grounded in image entities, handling temporal misalignment.

  • EMNLP Findings 2023 Visual Storytelling with Question-Answer Plans [paper]
    Danyang Liu, Mirella Lapata, Frank Keller

    Proposes visual storytelling that feeds images as a visual prefix to a pretrained language model and plans with question-answer blueprints.

  • ACL 2023 Attractive Storyteller: Stylized Visual Storytelling with Unpaired Text [paper]
    Dingyi Yang, Qin Jin

    Introduces stylized visual storytelling and a memory-augmented multitask model trained with unpaired style text to generate styled stories from photo streams.

  • EACL 2023 Multimodal Event Transformer for Image-guided Story Ending Generation [paper]
    Yucheng Zhou, Guodong Long

    Proposes an event-graph reasoning transformer for image-guided story ending generation, with cross-modal fusion, a multimodal injector, and incoherence detection.

  • ACL Findings 2023 Visual Coherence Loss for Coherent and Visually Grounded Story Generation [paper]
    Xudong Hong, Vera Demberg, A. Sayeed, Qiankun Zheng, B. Schiele

    Proposes a coherence-theory-inspired self-supervised loss and combined object and face features for character representation, plus a character matching metric for visual storytelling.

๐ŸŒ„ Story2Visual

  • EACL 2026 ๐ŸŒŸ MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling [paper]
    Qian Wang, Zi-Qi Huang, Ruoxi Jia, Paul E. Debevec, Ning Yu

    Proposes a multi-agent pipeline spanning scripting, shot design, character modeling, keyframes, animation, and audio for long-sequence video storytelling.

  • ECCV 2026 Better Call CineCrew: Consistent Ultra-Long Narrative-to-Film Generation [paper]
    Jiaben Chen, Si Dong, Qinhong Zhou, Raine Ma, Zhi-Yang Dou, Wojciech Matusik, Chuang Gan

    Proposes a multi-agent orchestration layer built on FilmDSL, a film-specific language making shot, continuity, and persona constraints explicit for long script-to-video generation.

  • TPAMI 2026 Learning Long-form Movie Prior via Large Language Models. [paper]
    Jin-Heng Xie, Jia-Jun Feng, M. Shou

    Represents movies as text and bounding-box or keypoint tokens and curates Storyboard20K, letting LLMs learn movie priors to sample storyboards.

  • ArXiv 2026 SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution [paper] GitHub stars
    Mao-Lin Ran, Xiaoyan Lu, Jia-Qi Liu, Jian Wang, Weiwen Liu, Jianghao Lin, Yong Yu, Weinan Zhang

    Builds a deployed storyboarding system that learns directing rules from expert examples, evolves them via attribution feedback, and releases the PROSE dataset.

  • ArXiv 2026 AniMaster: From Story Texts to Animated Videos via Cinematic Script Generation and Interactive Authoring [paper]
    Ruiqi Yu, De-Kun Qian, Jia-Le Xu, Si-Zhe Cheng, Yize Li, Xiang-yang Wu, Zhiguang Zhou, Wei Chen, Yong Wang

    Builds an authoring tool that expands brief story texts into cinematic scripts, then animated videos, guided by a three-layer design framework.

  • EMNLP 2026 MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation [paper]
    Mu-Yao Wang, Ze-Ke Xie, Yanhao Chen, Lixin Xiu, Hideki Nakayama

    Proposes an agentic story-to-manga framework decomposing creation into planning, grounding, layout, rendering, composition, and lettering for controllable page generation.

  • ACL Findings 2026 BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration [paper] GitHub stars
    Bo Gao, Chang Liu, Yu-Yang Miao, Siyuan Ma, S. Lim

    Proposes a safety-aware multi-agent framework for end-to-end illustrated storybook generation with page-level text-image calibration and global consistency repair.

  • ArXiv 2026 CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding [paper]
    I. Mondal, Yi-Wen Song, Mihir Parmar, Palash Goyal, J. Boyd-Graber, Tomas Pfister, Yale Song

    Proposes a multi-agent storyboarding framework that plans character, background, and location continuity, plus a new long-range consistency benchmark.

  • ArXiv 2026 LogiStory: A Logic-Aware Framework for Multi-Image Story Visualization [paper]
    Chutian Meng, Fan Ma, Chi Zhang, Jiaxu Miao, Yi Yang, Yue-Ting Zhuang

    Proposes a multi-agent story visualization framework that grounds roles, extracts causal chains, and verifies consistency to model visual logic explicitly.

  • ArXiv 2026 EmoStory: Emotion-Aware Story Generation [paper]
    Jing-Yuan Yang, Rucong Chen, Weibin Luo, Hui Huang

    Introduces emotion-aware visual story generation and a two-stage framework combining agent-based planning with region-aware generation for emotional, subject-consistent image sequences.

  • CVPR 2026 Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning [paper]
    Zhengjian Yao, Yong-Zhi Li, Xinyu Gao, Quan Chen, Peng Jiang, Yan Lu

    Combines an MLLM narrative planner with a memory-bank control module for long consistent visual sequences and releases a 330K-image e-commerce storyboard dataset.

  • ArXiv 2026 MUSE: A Multi-agent Framework for Unconstrained Story Envisioning via Closed-Loop Cognitive Orchestration [paper]
    Wenzhang Sun, Zhenyu Wang, Zhang-Chi Hu, Chun-Feng Wang, Hao Li, Wei Chen

    Proposes a multi-agent plan-execute-verify-revise loop for long audio-visual stories from short prompts, plus a reference-free evaluation protocol.

  • ICCV Workshop 2025 ๐ŸŒŸ SEED-Story: Multimodal Long Story Generation with Large Language Model [paper] GitHub stars
    Shuai Yang, Yu-Ying Ge, Yang Li, Yu-Kang Chen, Yixiao Ge, Shan Ying, Ying-Cong Chen

    Proposes SEED-Story, an MLLM generating interleaved text and consistent images for long stories via multimodal attention sinks, and releases the StoryStream dataset.

  • ArXiv 2025 Generating Storytelling Images with Rich Chains-of-Reasoning [paper]
    Xiujie Song, Qi Jia, Shota Watanabe, Xiao-Yi Pang, Ruijie Chen, Mengyue Wu, Ke Zhu

    Defines storytelling image generation with chains of visual reasoning clues and proposes an LLM-plus-text-to-image pipeline with dedicated evaluation metrics.

  • ACM MM 2025 From Outline to Detail: An Hierarchical End-to-end Framework for Coherent and Consistent Visual Novel Generation and Assembly [paper]
    Yilin Zhang, Yanyan Wei, Zhao Zhang, Jicong Fan, Haijun Zhang, Shui-Cheng Yan

    Proposes an outline-guided pipeline that generates and assembles executable visual novels, using vision-LLM self-correction for cross-modal consistency and script validation.

  • EMNLP 2025 LLMs Behind the Scenes: Enabling Narrative Scene Illustration [paper]
    Melissa Roemmele, John Joon Young Chung, Taewook Kim, Yuqian Sun, Alex Calderwood, Max Kreminski

    Uses LLMs to prompt text-to-image models for narrative scene illustration and releases SceneIllustrations, a dataset of pairwise human quality judgments.

  • SIGGRAPH Asia 2025 AniMaker: Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation [paper] GitHub stars
    Haoyuan Shi, Yunxin Li, Xinyu Chen, Long-Yue Wang, Bao-Tian Hu, Min Zhang

    Proposes AniMaker, a multi-agent animation framework using MCTS-driven multi-candidate clip generation and the AniEval evaluator to produce story-coherent long videos from text.

  • CVPR 2025 VinaBench: Benchmark for Faithful and Consistent Visual Narratives [paper]
    Silin Gao, Sheryl Mathew, Li Mi, Sepideh Mamooler, Mengjie Zhao, Hiromi Wakaki, Yuki Mitsufuji, Syrielle Montariol, Antoine Bosselut

    Introduces a benchmark annotating commonsense and discourse constraints in visual narratives, with metrics for consistency and text alignment of generated image sequences.

  • ArXiv 2025 MM-StoryAgent: Immersive Narrated Storybook Video Generation with a Multi-Agent Paradigm across Text, Image and Audio [paper] GitHub stars
    Xuenan Xu, Jiahao Mei, Chenliang Li, Yuning Wu, Ming Yan, Shaopeng Lai, Ji Zhang, Mengyue Wu

    Proposes a multi-agent framework combining LLMs with image, speech, music, and sound tools to generate narrated storybook videos for children.

  • ACL Findings 2025 VISIAR: Empower MLLM for Visual Story Ideation [paper]
    Zhaoyang Xia, Somdeb Sarkhel, Md Mehrab Tanjim, Stefano Petrangeli, Ishita Dasgupta, Yuxiao Chen, Jinxuan Xu, Di Liu, Saayan Mitra, Dimitris N. Metaxas

    Introduces visual story ideation, arranging visual assets into storylines, with an MLLM framework using a story graph and a VTravel benchmark.

  • ACL 2023 Multimodal Persona Based Generation of Comic Dialogs [paper]
    Harsh Agrawal, A. Mishra, Manish Gupta, M. -

    Introduces multimodal persona-based comic dialogue generation with a 54K-strip dataset and an architecture that generates next-panel dialogues.

โœ๏ธ Text Stories

๐Ÿ—บ๏ธ Planning

  • ArXiv 2026 A Multi-Framework Comparison of Outline Stages in Long-Form Generation with LLMs [paper]
    Yikang Song

    Benchmarks outlines from seven long-form generation frameworks with an anchored LLM judge, finding no framework dominates across chapter and book granularities.

  • TASLP 2026 LLM-Driven MCTS for Conditional Story Generation via Logic-Guided Evidence Tree Optimization [paper]
    Hongyan Wu, Zhiliang Tian, Zhen Huang, Nankai Lin, Yi-Ping Song, Zhihua Wen, Menglong Lu, Feng Liu, Dongsheng Li

    Proposes a plug-and-play MCTS planner that builds logic-validated evidence chains for retrieval-based conditional story generation to reduce incoherence and thematic drift.

  • ACL Findings 2026 Planning Beyond Text: Graph-based Reasoning for Complex Narrative Generation [paper]
    Hanwen Gu, Chao Guo, Junle Wang, Wen-Da Xie, Yi-Sheng Lv

    Proposes PLOTTER, which runs an Evaluate-Plan-Revise cycle on event and character graphs to fix causality and structure before generating full narrative text.

  • ArXiv 2026 BiT-MCTS: A Theme-based Bidirectional MCTS Approach to Chinese Fiction Generation [paper]
    Zhaoyi Li, Xu Zhang, Xiaojun Wan

    Generates Chinese fiction by writing the climax first, then expanding plot backward and forward with bidirectional MCTS inspired by Freytag's Pyramid.

  • TACL 2026 Lightweight Latent Reasoning for Narrative Tasks [paper]
    Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

    Proposes a lightweight reasoning projector producing continuous latent tokens that RL policies toggle, cutting reasoning length on plot-hole detection and chapter generation.

  • NAACL 2025 ๐ŸŒŸ Generating Long-form Story Using Dynamic Hierarchical Outlining with Memory-Enhancement [paper]
    Qianyue Wang, Jinwu Hu, Zhengpin Li, Yufeng Wang, Daiyuan Li, Yu Hu, Mingkui Tan

    Proposes DOME, which fuses planning and writing through dynamic hierarchical outlines and uses a memory module to reduce contradictions in long stories.

  • ICLR 2025 ๐ŸŒŸ Agents' Room: Narrative Generation through Multi-step Collaboration [paper]
    Fantine Huot, Reinald Kim Amplayo, J. Palomaki, Alice Shoshana Jakobovits, Elizabeth Clark, Mirella Lapata

    Proposes Agents' Room, which splits fiction writing into subtasks for specialized agents, and releases the Tell Me A Story dataset and evaluation.

  • CIKM 2025 StoryWriter: A Multi-Agent Framework for Long Story Generation [paper]
    Haotian Xia, Hao Peng, Yunjia Qi, Bin Xu, Juan-Zi Li, Hou Lei, Xiaozhi Wang

    Proposes a multi-agent long story framework and uses it to build a 6,000-story dataset for fine-tuning Llama3.1-8B and GLM4-9B.

  • ACL Findings 2025 STORYTELLER: An Enhanced Plot-Planning Framework for Coherent and Cohesive Story Generation [paper]
    Jiaming Li, Yu-Kun Chen, Ziqiang Liu, Minghuan Tan, Lei Zhang, Yunshui Li, Run Luo, Long-Ze Chen, Jing Luo, A. Argha, et al.

    Proposes a plot-planning approach using SVO-triplet plot nodes plus interacting storyline and narrative entity knowledge graph modules for coherent story generation.

  • EMNLP 2025 Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language Models [paper] GitHub stars
    Ruibin Xiong, Yi-Meng Chen, Dmitrii Khizbullin, Mingchen Zhuge, Jurgen Schmidhuber

    Proposes a writing agent that recursively interleaves retrieval, reasoning, and composition tasks instead of fixed outlining, evaluated on fiction and technical reports.

  • ACL Findings 2025 A Cognitive Writing Perspective for Constrained Long-Form Text Generation [paper] GitHub stars
    Kaiyang Wan, Hong-Lin Mu, Rui Hao, Haoran Luo, Tianle Gu, Xiuying Chen

    Proposes CogWriter, a training-free framework applying Cognitive Writing Theory via planning, parallel generation, and review agents for constrained long-form text.

  • NAACL 2025 Navigating the Path of Writing: Outline-guided Text Generation with Large Language Models [paper]
    Yukyung Lee, Soonwon Ka, Bokyung Son, Pilsung Kang, Jaewook Kang

    Proposes WritingPath, which guides LLMs with explicit outlines reflecting user intent, and builds a blog-post dataset and evaluation framework for goal-oriented writing.

  • ACL 2024 Ex3: Automatic Novel Writing by Extracting, Excelsior and Expanding [paper] GitHub stars
    H. Lei, Jiaming Guo, Guanhua He, Xi-Shan Zhang, Rui Zhang, Shaohui Peng, Shaoli Liu, Tianshi Chen

    Proposes Ex3, which extracts structure from raw novels to build instruction data, fine-tunes an LLM, and expands tree-like into arbitrarily long novels.

  • AAAI 2024 Does Robin Hood Use a Lightsaber?: Automated Planning for Storytelling [paper]
    Nisha Ingrid Simon

    Combines automated planning with LLM text generation, using a planning model as scaffolding to produce more logical, coherent, and believable stories.

  • EMNLP Findings 2024 SWAG: Storytelling With Action Guidance [paper] GitHub stars
    Zeeshan Patel, Karim El-Refai, Jonathan Pei, Tianle Li

    Proposes SWAG, framing story writing as search where an auxiliary LLM picks the next action steering the generator toward engaging stories.

  • LREC-COLING 2024 Little Red Riding Hood Goes around the Globe: Crosslingual Story Planning and Generation with Large Language Models [paper]
    E. Razumovskaia, Joshua Maynez, Annie Louis, Mirella Lapata, Shashi Narayan

    Introduces crosslingual story generation with planning and a dataset, finding three-act plans yield more coherent, interesting, controllable stories across languages.

  • ACL 2023 ๐ŸŒŸ DOC: Improving Long Story Coherence With Detailed Outline Control [paper]
    Kevin Yang, D. Klein, Nanyun Peng, Yuan-Dong Tian

    Improves long-story plot coherence by generating a detailed hierarchical outline and a controller that keeps drafted passages aligned with outline details.

๐Ÿงต Coherence

  • IJCAI 2026 FossilWriter: Learning Hypergraph World Models with Latent Narratives for Creative Story Generation [paper]
    Heng Zhang, Yi-Hao Zhong, Lubin Gan, Zhihe Chen, Tianyi Zhang, Jing Liu, Jin Huang

    Grows a hypergraph world model whose unresolved elements seed latent narratives, improving plot coherence and reducing long-range factual conflicts in story generation.

  • ArXiv 2026 Narrative World Model: Narratology-Grounded Writer Memory for Long-Form Fiction [paper]
    M. Saifullah, Thomas Kornmaier, Taaha Kazi, Vasu Sharma, A. Kanade, Aanand Kumar Yadav

    Proposes a writer-memory system pairing a narratology-typed temporal state graph with hybrid retrieval, outperforming Graphiti/Zep and GraphRAG on multi-hop story questions.

  • EMNLP Findings 2026 ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control [paper] GitHub stars
    Jindong Li, Yang Yang, Zihao Liu, Yutao Yue, Meng-Lin Yang

    Proposes a training-free scene-by-scene writer that tracks symbolic story states, checks narrative transitions, and uses uncertainty signals to repair inconsistencies.

  • AAAI 2026 Octopus: Entropy-Controlled Science Fiction Literature Generation with Persistent Memory-Context Binding [paper]
    Xu Wang, Jiaju Kang, Puyu Han, Zeyu Ai, Luqi Gong

    Proposes Octopus, combining entropy regulation via narrative divergence thresholds with hierarchical memory of characters, plots, and scientific rules for long sci-fi generation.

  • WWW 2025 SCORE: Story Coherence and Retrieval Enhancement for AI Narratives [paper]
    Qiang Yi, Yang-Fan He, Jian-Hui Wang, Xin-Yuan Song, Shi-Yao Qian, Xin-Hang Yuan, Yi Xin, Yi-Jin Wang, Jingqun Tang, Yuchen Li, et al.

    Proposes SCORE, which tracks key item states and episode summaries and uses retrieval-augmented generation to detect and fix inconsistencies in LLM-generated stories.

  • COLING 2025 MLD-EA: Check and Complete Narrative Coherence by Introducing Emotions and Actions [paper]
    Jin-Ming Zhang, Yun-Fei Long

    Proposes MLD-EA, which uses LLMs with emotion and action cues to detect missing logic in narratives and generate sentences that restore coherence.

  • NAACL 2025 FACTTRACK: Time-Aware World State Tracking in Story Outlines [paper]
    Zhiheng Lyu, Kevin Yang, Lingpeng Kong, Daniel Klein

    Proposes FACTTRACK, which decomposes events into atomic facts with time-aware validity intervals to track world state and detect contradictions in story outlines.

  • COLM 2024 With Greater Text Comes Greater Necessity: Inference-Time Training Helps Long Text Generation [paper] GitHub stars
    Yan Wang, D. Ma, Deng Cai

    Proposes Temp-Lora, which stores long context in a temporary LoRA module trained during generation, improving long-text quality while cutting context-window costs.

  • ArXiv 2023 ๐ŸŒŸ RecurrentGPT: Interactive Generation of (Arbitrarily) Long Text [paper] GitHub stars
    Wangchunshu Zhou, Y. Jiang, Peng Cui, Tiannan Wang, Zhenxin Xiao, Yifan Hou, Ryan Cotterell, Mrinmaya Sachan

    Proposes RecurrentGPT, which simulates LSTM-style recurrence with natural-language long- and short-term memories so LLMs can interactively generate arbitrarily long text.

๐Ÿง‘โ€๐Ÿคโ€๐Ÿง‘ Characters

  • ArXiv 2026 ANIMASK: What the Model Contributes to Role Play in Simulated Story Worlds [paper]
    Xiu-Cheng Zhang, Zhuo-Ning Xu, Han-Jun Luo, Yankai Chen, Hanan Salam, Xue Liu

    Replays story worlds from freeze points with and without personas, finding actor LLMs push characters toward cautious, flatter outcomes than canon.

  • ArXiv 2026 From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives [paper]
    Aayush Aluru, C. Ho, Muhammad Hammouri, Kerry Luo, Myra Malik, Ryan Lagasse, Arjun Bahuguna, Vasu Sharma

    Proposes multi-agent persona-driven story generation with shared world state plus a graph-based hallucination detector, halving hallucinations in 100-page stories.

  • ArXiv 2026 Staying In Character: Perspective-Bounded Memory For Book-Based Role-Playing Agents [paper]
    Xushuo Tang, Junhe Zhang, Zi-Han Yang, Yi-Fu Tang, Sichao Li, Longbin Lai, Zheng-Yi Yang

    Proposes a three-layer, perspective-bounded memory for book-based role-playing agents that prevents characters using unknown facts, with a 4,386-question knowledge-boundary benchmark.

  • ACL 2026 EvoSpark: Endogenous Interactive Agent Societies for Unified Long-Horizon Narrative Evolution [paper]
    Shiyu He, Min-Chi Kuang, Mengxian Wang, Bin Hu, Tingxiang Gu

    Proposes a multi-agent framework with stratified narrative memory and role-location-plot alignment to sustain coherent, open-ended long-horizon story evolution.

  • ACL 2026 Deriving Character Logic from Storyline as Codified Decision Trees [paper] GitHub stars
    Letian Peng, Kun Zhou, Longfei Yun, Yu-Peng Hou, Jingbo Shang

    Induces executable, interpretable decision trees of validated scene-conditioned behavior rules from narrative data to ground role-playing agents more reliably.

  • AAAI 2026 StoryBox: Collaborative Multi-Agent Simulation for Hybrid Bottom-Up Long-Form Story Generation Using Large Language Models [paper]
    Zehao Chen, Rong Pan, Hao-Ran Li

    Proposes bottom-up long-form story generation in which multi-agent sandbox simulation yields emergent events that form coherent stories exceeding 10,000 words.

  • ArXiv 2025 ๐ŸŒŸ BookWorld: From Novels to Interactive Agent Societies for Creative Story Generation [paper] GitHub stars
    Yi-Ting Ran, Xintao Wang, Tian Qiu, Jiaqing Liang, Yanghua Xiao, Deqing Yang

    Builds BookWorld, which simulates multi-agent societies from established novels' characters and worldviews to generate creative stories faithful to the source books.

  • AIIDE 2025 Steering Narrative Agents Through a Dynamic Cognitive Framework for Guided Emergent Storytelling [paper]
    Chen Yang, M. Gross, R. Wampfler

    Proposes a cognitive agent framework where tensions between agents' beliefs and ideal worlds drive actions, steering emergent stories toward authored storylines.

  • FDG 2024 ๐ŸŒŸ StoryVerse: Towards Co-authoring Dynamic Plot with LLM-based Character Simulation via Narrative Planning [paper]
    Yi Wang, Qian Zhou, David Ledo

    Proposes StoryVerse, where authors write abstract acts that LLM narrative planning turns into character actions, balancing authorial intent with emergent game plots.

๐ŸŽจ Creativity

  • ArXiv 2026 MUSE: A Theory-Harnessed Story Engine for Vibe Narrativizing [paper] GitHub stars
    Jian-Xiang Ma, Xiaocui Yang, Da-Ling Wang, Yue-Song Hou, Ming-Fu Zhang, Yi-Chen Gao, Jun-Zhao Huang

    Builds a story engine encoding McKee's story theory as atomized rules inside an agent harness, improving WritingBench and consistency across four models.

  • ArXiv 2026 CraftAlign: Feature-Grounded Evaluation and Revision Guidance for AI Stories [paper]
    Yang Yang, Boyun Xu, Shaofeng Liang, Yun Han, Zining Zhong, Songning Lai, Kaishen Yuan, Yutao Yue

    Predicts 304 writing features to score stories against human and AI patterns, turning feature shifts into revision guidance that reduces AI flavor.

  • ArXiv 2026 StorySpark: Module-wise Evolutionary Search for Story Premise Generation [paper]
    Yang Yang, Zining Zhong, Qian Cao, Jindong Li, Boyun Xu, Kaishen Yuan, Meng-Lin Yang, Yutao Yue

    Proposes module-wise evolutionary search over premise components like persona, event, and twist, producing more original premises that yield better downstream stories.

  • ArXiv 2026 PlotTwist: A Creative Plot Generation Framework with Small Language Models [paper]
    A. Thorat, Ravi Kolla, Jyotin Goel, Madhav Kataria, N. Pedanekar

    Proposes a framework where sub-3B models generate premise-conditioned plots using an aspect reward model, DPO-aligned MoE generator, and cross-family jury evaluation.

  • ArXiv 2026 LLM Review: Enhancing Creative Writing via Blind Peer Review Feedback [paper] GitHub stars
    Weiyue Li, Mingxiao Song, Zhenda Shen, Dachuan Zhao, Yunfan Long, Yi Li, Yongce Li, Ruyi Yang, Meng-Yu Wang

    Proposes blind peer review among LLM agents that exchange feedback but revise independently, avoiding homogenization, and introduces the SciFi-100 writing dataset.

  • ACL 2026 Frankentext: Stitching random text fragments into long-form narratives [paper] GitHub stars
    Chau Minh Pham, Jenna Russell, Dzung Pham, Mohit Iyyer

    Proposes generating long narratives by having LLMs stitch mostly verbatim human-written fragments, improving diversity and originality while often evading AI-text detectors.

  • EMNLP 2025 Avoidance Decoding for Diverse Multi-Branch Story Generation [paper]
    Kyeongman Park, Nakyeong Yang, Kyomin Jung

    Proposes a decoding strategy that penalizes concept- and narrative-level similarity to earlier outputs, increasing diversity across multiple story branches from one prompt.

  • ACL Findings 2025 A Character-Centric Creative Story Generation via Imagination [paper]
    Kyeongman Park, Minbeom Kim, Kyomin Jung

    Proposes character-centric story generation that uses text-to-image imagination of story elements and multi-writer persona selection to deepen characters and creativity.

  • EMNLP 2024 ๐ŸŒŸ Collective Critics for Creative Story Generation [paper] GitHub stars
    Minwook Bae, Hyounghun Kim

    Proposes CritiCS, where a group of LLM critics collectively revise story plans and text to make long stories more creative and expressive.

  • ACL 2024 ๐ŸŒŸ MoPS: Modular Story Premise Synthesis for Open-Ended Automatic Story Generation [paper] GitHub stars
    Yan Ma, Yu Qiao, Pengfei Liu

    Proposes MoPS, which composes story premises from modular elements like background and persona, yielding more diverse and original premises for story generation.

  • EACL 2024 ๐ŸŒŸ Creating Suspenseful Stories: Iterative Planning with Large Language Models [paper]
    Kaige Xie, Mark Riedl

    Proposes a zero-shot iterative prompting planner grounded in cognitive-psychology and narratology theories of suspense to generate suspenseful stories with LLMs.

  • IJCAI 2024 A Conflict-Embedded Narrative Generation Using Commonsense Reasoning [paper]
    Youngrok Song, Gunhee Cho, Hyun-Jee Kim, Youngjune Kim, Byung-Chull Bae, Yun-Gyung Cheong

    Proposes a neuro-symbolic framework that embeds conflict in stories by using commonsense defeasible inference to weaken causal links toward protagonist goals.

  • NAACL 2024 Returning to the Start: Generating Narratives with Related Endpoints [paper] GitHub stars
    A. Brei, Chao Zhao, Snigdha Chaturvedi

    Proposes RENarGen, which first generates related opening and closing sentences then infills the middle, producing stories with stronger narrative closure.

  • EMNLP Findings 2023 Improving Pacing in Long-Form Story Planning [paper] GitHub stars
    Yichen Wang, Kevin Yang, Xiaoming Liu, Dan Klein

    Proposes CONCOCT, which trains a concreteness evaluator to guide vaguest-first outline expansion and filtering, yielding more consistent pacing in story outlines.

  • EMNLP Findings 2023 Affective and Dynamic Beam Search for Story Generation [paper]
    Tenghao Huang, Ehsan Qasemi, Bangzheng Li, He Wang, Faeze Brahman, Muhao Chen, Snigdha Chaturvedi

    Proposes a decoding method combining bandit-driven dynamic beam sizing and affect-intensity reranking to generate stories with more interesting twists.

  • EMNLP Findings 2023 GROVE: A Retrieval-augmented Complex Story Generation Framework with A Forest of Evidence [paper]
    Zhihua Wen, Zhi-Liang Tian, Wei Wu, Yu-Xin Yang, Yanqi Shi, Zhen Huang, Dongsheng Li

    Proposes GROVE, which retrieves human-written story examples and builds an asking-why forest of evidence to add complex, credible plot details.

  • EMNLP Findings 2023 Narrative Order Aware Story Generation via Bidirectional Pretraining Model with Optimal Transport Reward [paper]
    Zhicong Lu, Li Jin, Guang-Luan Xu, Linmei Hu, Nayu Liu, Xiaoyu Li, Xian Sun, Zequn Zhang, Kaiwen Wei

    Proposes a bidirectional pretrained event model with RL using an optimal transport reward to generate coherent stories with flashbacks.

๐ŸŽฏ Training

  • ArXiv 2026 Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion [paper]
    Hwan Chang, Yongil Kim, Heuiyeen Yeen, Yireun Kim, Jinsik Lee, Hwanhee Lee

    Proposes attribute-guided genre expansion to build a 50K, 13-genre creative writing corpus; fine-tuning on it beats training on story-centric writing data.

  • ArXiv 2026 Retell, Reward, Repeat: Reinforcement Learning for Narrative Theory-Informed Story Generation [paper]
    David Y. Liu, Xanthe Muston, A. Joshi, Sebastian Sequoiah-Grayson

    Shows that reinforcement learning from narrative-theory-informed AI feedback (d-RLAIF) yields more diverse, convention-aligned stories than supervised fine-tuning.

  • EMNLP 2026 POLARIS: Guiding Small Models to Write Long Stories [paper]
    Rishanth Rajendhran, Jenna Russell, Mohit Iyyer, J. Wieting

    Proposes a GRPO recipe with LLM-judge rewards and injected human reference stories, letting a 9B model write long stories beyond training length.

  • EMNLP 2026 Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction [paper]
    Ze-Han Li, Yu-Tong Zhu, Si-Yang Wu, Honglin Bao, James A. Evans

    Finds by comparing OLMo checkpoints that post-training compresses thematic, affective, and stylistic variation in fiction, most for professional literary text.

  • ArXiv 2026 StoryAlign: Evaluating and Training Reward Models for Story Generation [paper]
    Hao-Tian Xia, Hao Peng, Yunjia Qi, Xiaozhi Wang, Bin Xu, Lei Hou, Juan-Zi Li

    Introduces StoryRMB, a benchmark exposing weak reward models for story preferences, and StoryReward, trained on 100K preference pairs for best-of-n story selection.

  • ACL Findings 2026 UniCreative: Unifying Long-form Logic and Short-form Sparkle via Reference-Free Reinforcement Learning [paper]
    Xiao-Long Wei, Zerun Zhu, Simin Niu, Xingyu Zhang, Peiying Yu, Chang Xiao, Yu-Chen Li, Ji-Cheng Yang, Zhejun Zhao, Chong Meng, et al.

    Proposes a reference-free RL framework with an adaptive constraint-aware generative reward model and ACPO policy optimization, unifying long-form and short-form creative writing.

  • ACL 2026 DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative Writing [paper] GitHub stars
    Qian Cao, Yahui Liu, Wei Bi, Yi Zhao, Rui-jie Song, Xiting Wang, Rui-Ming Tang, Guo-Rui Zhou, Han Li

    Proposes an RL framework for creative writing that branches diverse plans in long chain-of-thought and adds a group-aware diversity reward.

  • ArXiv 2026 Rewarding Creativity: A Human-Aligned Generative Reward Model for Reinforcement Learning in Storytelling [paper]
    Zhaoyan Li, Hang Lei, Yuji Wang, Lan Liu, Hao Liu, Liang Yu

    Proposes RL for storytelling with a reasoning generative reward model aligned to human creativity judgments and entropy-based reward shaping for training stability.

  • ACL Findings 2026 From Style to Story: A Curriculum Learning Approach for Imitative Novel Generation [paper]
    Xueran Han, Yuhan Liu, Mingzhe Li, Wei Liu, Sen Hu, Rui Yan, Zhiqiang Xu, Xiuying Chen

    Introduces imitative novel generation and trains WriterAgent via curriculum learning with hierarchical LoRA modules to mimic an author's style, characters, and plots.

  • CoNLL 2026 Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning [paper] GitHub stars
    Jinlong Liu, Mohammed Bahja, Venelin Kovatchev, Mark Lee

    Trains an authorship-verification style judge and uses it as a GRPO reward to fine-tune an 8B model for writing like classic authors.

  • AAAI 2026 RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing [paper]
    Jian-Xing Liao, Tian Zhang, Xiao Feng, Yusong Zhang, Rui Yang, Hao-Rui Wang, Bosi Wen, Ziyi Wang, Run-Zhi Shi

    Proposes RL with a dynamically weighted mix of writing-quality and constraint-verification rewards in GRPO, improving both creative quality and instruction following.

  • ACL 2026 Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning [paper] GitHub stars
    Xuanyu Lei, Chen-Liang Li, Yuning Wu, Kai Liu, Weizhou Shen, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Yang Liu

    Proposes adaptive curriculum RL for long-form writing with margin-aware data selection, pairwise comparison rewards, and dynamic reference scheduling, beating SFT baselines.

  • ArXiv 2025 ๐ŸŒŸ Learning to Reason for Long-Form Story Generation [paper] GitHub stars
    Alexander Gurung, Mirella Lapata

    Proposes RL for story reasoning via Next-Chapter Prediction, rewarding plans that raise completion likelihood of real book chapters without labeled data.

  • ArXiv 2025 ๐ŸŒŸ Modifying Large Language Model Post-Training for Diverse Creative Writing [paper] GitHub stars
    John Joon Young Chung, Vishakh Padmakumar, Melissa Roemmele, Yuqian Sun, Max Kreminski

    Adds deviation from other same-prompt samples into DPO and ORPO objectives, increasing creative writing output diversity with minimal quality loss.

  • ArXiv 2025 LiteraryTaste: A Preference Dataset for Creative Writing Personalization [paper]
    John Joon Young Chung, Vishakh Padmakumar, Melissa Roemmele, Yi Wang, Yuqian Sun, Tiffany Wang, S. Almeda, Brett A. Halperin, Yuwen Lu, Max Kreminski

    Releases reading preferences from 60 people over creative text pairs, finding tastes diverge and stated preferences poorly predict revealed ones.

  • ArXiv 2025 COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes [paper] GitHub stars
    Yunwen Li, Shuangshuang Ying, Xingwei Qu, Xin Li, Sheng Jin, Minghao Liu, Zhoufutu Wen, Tianyu Zheng, Xeron Du, Qiguang Chen, et al.

    Releases 1,665 Chinese creative writing triplets with reverse-engineered prompts and reasoning traces, finding process supervision helps only when mixed with general data.

  • ArXiv 2024 ๐ŸŒŸ Weaver: Foundation Models for Creative Writing [paper]
    Tiannan Wang, Jiamin Chen, Qi Jia, Shuai Wang, Ruoyu Fang, Huilin Wang, Zhaowei Gao, Chunzhao Xie, Chuou Xu, Jihong Dai, et al.

    Introduces Weaver, a 1.8B-34B LLM family pre-trained and aligned for creative and professional writing, with a routing agent balancing quality and cost.

  • EMNLP 2024 MirrorStories: Reflecting Diversity through Personalized Narrative Generation with Large Language Models [paper]
    Sarfaroz Yunusov, Hamza Sidat, Ali Emami

    Introduces MirrorStories, 1,500 LLM-generated stories personalized to reader identity, and finds they engage readers more than generic human or LLM stories.

  • EMNLP Workshop 2024 PEARL: Personalizing Large Language Model Writing Assistants with Generation-Calibrated Retrievers [paper]
    Sheshera Mysore, Zhuoran Lu, Meng-Ting Wan, Long-Qi Yang, Steve Menezes, Tina Baghaee, E. Gonzalez, Jennifer Neville, Tara Safavi

    Proposes Pearl, a personalized writing assistant whose retriever is trained to be generation-calibrated, selecting user documents that most improve personalized LLM outputs.

๐Ÿ“ Evaluation

๐Ÿงช Benchmarks

  • EACL 2026 ๐ŸŒŸ LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing [paper]
    Daniel Fein, S. Russo, Violet Xiang, Kabir Jolly, Rafael Rafailov, Nick Haber

    Introduces a benchmark of 2,480 human-labeled story comparisons and 43,827 training pairs for creative writing evaluation, benchmarking LLM judges and reward models.

  • ACL Findings 2026 Lost in Stories: Consistency Bugs in Long Story Generation by LLMs [paper] GitHub stars
    Junjie Li, Xinru Guo, Yuhao Wu, Roy Ka-Wei Lee, Hong-Zhi Li, Yutao Xie

    Introduces a 2,000-prompt benchmark and automated checker for consistency errors in long story generation, analyzing where and which contradictions LLMs make.

  • ACL 2026 LitVISTA: A Benchmark for Narrative Orchestration in Literary Text [paper]
    Mingzhe Lu, Yiwen Wang, Yanbing Liu, Qi You, Chong Liu, Ruize Qin, Haoyu Dong, Wenyu Zhang, Jia-Rui Zhang, Yue Hu, et al.

    Proposes a framework and annotated literary benchmark for narrative orchestration, finding frontier LLMs fail to jointly capture narrative function and structure.

  • ACL Findings 2026 ChangJuan: A Comprehensive Benchmark for Book-Length Chinese Story Evaluation [paper] GitHub stars
    Dingyi Yang, Mingshuo Wang, Qin Jin

    Introduces a benchmark of 300 Chinese novels with human ratings and distilled reader viewpoints, plus CLEM, an 8B evaluator for book-length stories.

  • EACL 2026 Mary, the Cheeseburger-Eating Vegetarian: Do LLMs Recognize Incoherence in Narratives? [paper]
    Karin de Langis, Pรผren ร–ncel, Ryan Peters, Andrew Elfenbein, Laura K. Allen, Andreas Schramm, Dongyeop Kang

    Finds LLM internal representations detect incoherent narratives but their ratings do not, and models notice setting violations more than character-trait violations.

  • EACL Findings 2026 WebNovelBench: Placing LLM Novelists on the Web Novel Distribution [paper] GitHub stars
    Leon Lin, Jun Zheng, Haidong Wang

    Introduces a benchmark of 4,000+ Chinese web novels that scores LLM synopsis-to-story outputs on eight dimensions and ranks them against human-authored percentiles.

  • ACL 2025 What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation [paper] GitHub stars
    Dingyi Yang, Qin Jin

    Introduces LongStoryEval, 600 books averaging 121K tokens with reader reviews, compares long-story evaluation methods, and trains NovelCritique, an 8B summary-based evaluator.

  • ArXiv 2025 Finding Flawed Fictions: Evaluating Complex Reasoning in Language Models via Plot Hole Detection [paper]
    Kabir Ahuja, Melanie Sclar, Yulia Tsvetkov

    Introduces FlawedFictions, a benchmark built by synthesizing plot holes in human stories, to test LLM narrative reasoning via plot hole detection.

  • ArXiv 2025 LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm [paper] GitHub stars
    Siwei Wu, Yizhi Li, Xingwei Qu, R. Ravikumar, Yucheng Li, Tyler Loakman, Shanghaoran Quan, Xiao-Yong Wei, R. Batista-Navarro, Cheng-Hua Lin

    Introduces LongEval, a benchmark comparing direct and plan-based long-text generation, finding LLMs degrade with length while small long-text-trained models stay competitive.

  • ACL Findings 2025 Towards A "Novel" Benchmark: Evaluating Literary Fiction with Large Language Models [paper]
    Wenqing Wang, Mingqi Gao, Xinyu Hu, Xiaojun Wan

    Proposes a ten-metric macro/meso/micro evaluation framework and bilingual annotated fiction dataset, revealing a high-starting, low-ending pattern in LLM-written novels.

  • ICLR 2025 LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs [paper] GitHub stars
    Yuhao Wu, Ming Shan Hee, Zhiqing Hu, Roy Ka-Wei Lee

    Introduces a benchmark for instruction-following long-form generation at 16K and 32K tokens, finding all tested LLMs struggle as output length grows.

  • NAACL Findings 2025 CollabStory: Multi-LLM Collaborative Story Generation and Authorship Analysis [paper] GitHub stars
    Saranya Venkatraman, N. Tripto, Dongwon Lee

    Introduces CollabStory, a dataset of 32k stories co-written by up to five LLMs, with authorship analysis tasks and baselines for multi-LLM writing.

  • ArXiv 2024 CS4: Measuring the Creativity of Large Language Models Automatically by Controlling the Number of Story-Writing Constraints [paper] GitHub stars
    Anirudh Atmakuru, Jatin Nainani, Rohith Siddhartha Reddy Bheemreddy, Anirudh Lakkaraju, Zonghai Yao, Hamed Zamani, Haw-Shiuan Chang

    Introduces CS4, a benchmark that measures LLM story creativity by varying the number of prompt constraints to prevent retelling memorized stories.

  • ACL 2023 StoryWars: A Dataset and Instruction Tuning Baselines for Collaborative Story Understanding and Generation [paper]
    Yulun Du, Lydia B. Chilton

    Introduces StoryWars, 40k collaborative stories from 9,400 authors forming 101 understanding and generation tasks, with an instruction-tuned InstructStory baseline.

๐Ÿ“ Metrics

  • ArXiv 2026 When Reasoning Supervision Hurts: TTCW-Based Long-Form Literary Review Generation [paper] GitHub stars
    Jinlong Liu, Mohammed Bahja, Mark D. Lee

    Releases 263K stories with TTCW-based review annotations and finds fine-tuning without reasoning traces outperforms reasoning-supervised training for literary review generation.

  • ArXiv 2026 Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling [paper]
    Pei-Qi Sui, Yu-Tong Zhu, Tianyi Cheng, Peter West, R. So, Hoyt Long, Ari Holtzman

    Introduces 100-Endings, measuring narrative tension by how often repeated ending predictions fail as a story unfolds, and a pipeline that raises tension.

  • ACL 2026 EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation [paper]
    Xinda Wang, Zhengxu Hou, Yangshijie Zhang, Bingren Yan, Jia-Lin Liu, Chen-Zhuo Zhao, Zhi-Bo Yang, Bin-Bin Yang, Feng Xiao

    Trains a pairwise story evaluator on self-synthesized, multi-agent-filtered chain-of-thought data and uses it as a reward model to improve story generation.

  • ACL Findings 2025 The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing [paper] GitHub stars
    Guillermo Marco, Julio Gonzalo, Vรญctor Fresno-Fernรกndez

    Finds conflicting evaluations of AI fiction reflect reader differences, clustering 101 annotators into surface-focused and holistic reader profiles via textual feature preferences.

  • Applied Sciences 2025 Evaluating Creativity: Can LLMs Be Good Evaluators in Creative Writing Tasks? [paper]
    Sungeun Kim, Dongsuk Oh

    Finds that LLMs rate creative texts more consistently than humans but miss nuanced, culturally specific, and context-dependent aspects of creativity.

  • TACL 2024 ๐ŸŒŸ Do Language Models Enjoy Their Own Stories? Prompting Large Language Models for Automatic Story Evaluation [paper]
    Cyril Chhun, Fabian M. Suchanek, Chloรฉ Clavel

    Studies LLMs as automatic story evaluators, finding they beat existing metrics at system-level correlation with humans but struggle to explain their ratings.

  • EMNLP Findings 2024 CHIRON: Rich Character Representations in Long-Form Narratives [paper]
    Alexander Gurung, Mirella Lapata

    Proposes character sheet representations built by LLM question-answering and entailment-based fact validation, improving masked-character prediction and measuring character-centricity.

  • EMNLP 2024 Learning Personalized Alignment for Evaluating Open-ended Text Generation [paper]
    Danqing Wang, Kevin Yang, Hanlin Zhu, Xiaomeng Yang, Andrew Cohen, Lei Li, Yuandong Tian

    Proposes PerSE, a LLaMA-2 based evaluator that infers reader preferences from in-context profiles to give personalized, interpretable scores for open-ended generation.

  • ACL 2023 ๐ŸŒŸ Can Large Language Models Be an Alternative to Human Evaluations? [paper]
    Cheng-Han Chiang, Hung-yi Lee

    Finds that LLMs given the same instructions as human annotators produce story and adversarial-text ratings consistent with expert human evaluation.

  • EMNLP Findings 2023 DeltaScore: Evaluating Story Generation with Differentiating Perturbations [paper]
    Zhuohan Xie, Miao Li, Trevor Cohn, Jey Han Lau

    Proposes DeltaScore, which evaluates story aspects like fluency and interestingness by measuring likelihood changes under aspect-specific perturbations.

๐Ÿ” Analyses

  • ACL 2026 CASPER in the Machine: Insights into Character Variety in LLM-Generated Stories [paper]
    A. Brei, Abhisheik Sharma, Nicholas Sanaie, Lu Wang, Snigdha Chaturvedi

    Compares characters in LLM-generated and human-written stories along eight narratological dimensions, examining similarity and variety of character types.

  • ACL Findings 2026 BIASEDTALES-ML: A Multilingual Dataset for Analyzing Narrative Attribute Distributions in LLM-Generated Stories [paper]
    Yuxuan Ouyang, Yi Luo, Jing-Bo Zhu, Tong Xiao

    Releases a 350K-story multilingual parallel corpus of LLM children's stories, finding narrative attribute distributions vary substantially across eight languages.

  • ArXiv 2026 StoryScope: Investigating idiosyncrasies in AI fiction [paper]
    Jenna Russell, Rishanth Rajendhran, Chau Minh Pham, Mohit Iyyer, J. Wieting

    Finds that discourse-level narrative features alone separate human from AI fiction and attribute AI stories to specific models, independent of stylistic cues.

  • ArXiv 2026 LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers [paper]
    Pei-Qi Sui

    Finds across 28 LLMs that model story continuations carry much lower information-theoretic uncertainty than human writing, worsened by instruction tuning.

  • PNAS 2025 ๐ŸŒŸ Echoes in AI: Quantifying Lack of Plot Diversity in LLM Outputs [paper]
    Wei-Jia Xu, Nebojsa Jojic, Sudha Rao, C. Brockett, Bill Dolan

    Finds LLM stories from one prompt reuse plot element combinations far more than human stories, and proposes an automatic narrative-level diversity metric.

  • EMNLP 2025 Biased Tales: Cultural and Topic Bias in Generating Children's Stories [paper]
    Donya Rooein, Vilรฉm Zouhar, Debora Nozza, Dirk Hovy

    Introduces a dataset exposing gender and cultural stereotypes in LLM children's stories, such as girls receiving more appearance-related attributes than boys.

  • Information 2025 AI Narrative Modeling: How Machines' Intelligence Reproduces Archetypal Storytelling [paper]
    I. Kabashkin, Olga Zervina, Boriss Misnevs

    Finds LLMs reproduce structured Jungian archetypes like the Hero well but struggle with ambiguous ones like the Shadow and Trickster.

  • ICCC 2025 Evaluating Creative Short Story Generation in Humans and Large Language Models [paper] GitHub stars
    Mete Ismayilzada, Claire E. Stevenson, Lonneke van der Plas

    Compares stories by 60 LLMs and 60 humans, finding LLMs lag in novelty and surprise though non-experts rate LLM stories more creative.

  • COLING 2025 Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs [paper] GitHub stars
    Guillermo Marco, Luz Rello, Julio Gonzalo

    Finds a fine-tuned BART-large outscores average human writers on short fiction in human ratings, contrasting its linguistic traits with GPT-3.5 and GPT-4o.

  • EMNLP 2024 ๐ŸŒŸ Are Large Language Models Capable of Generating Human-Level Narratives? [paper]
    Yufei Tian, Tenghao Huang, Miri Liu, Derek Jiang, Alexander Spangher, Muhao Chen, Jonathan May, Nan-Yun Peng

    Analyzes story arcs, turning points, and affect, finding LLM stories are homogeneously positive and lack tension compared with suspenseful, diverse human narratives.

  • CHI 2024 ๐ŸŒŸ Art or Artifice? Large Language Models and the False Promise of Creativity [paper]
    Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, S. Muresan, Chien-Sheng Wu

    Proposes the Torrance Test of Creative Writing, finding LLM stories pass 3-10X fewer expert tests than professional stories, and LLM judges misalign.

  • EMNLP 2024 Pron vs Prompt: Can Large Language Models already Challenge a World-Class Fiction Author at Creative Text Writing? [paper]
    Guillermo Marco, Julio Gonzalo, M.Teresa Mateo-Girona, Ramรณn Santos

    Stages a contest between novelist Patricio Pron and GPT-4, where expert critics judge the LLM far from a top human fiction author.

  • EMNLP 2024 Measuring Psychological Depth in Language Models [paper]
    Fabrice Y. Harel-Canada, Hanyu Zhou, Sreya Muppalla, Zeynep Yildiz, Miryung Kim, Amit Sahai, Nan-Yun Peng

    Introduces the Psychological Depth Scale for stories' emotional and empathic impact, automates it with LLM personas, finding GPT-4 rivals top Reddit stories.

  • HSSC 2023 Experimental narratives: A comparison of human crowdsourced storytelling and AI storytelling [paper]
    Nina Beguลก

    Compares crowdworker and GPT-3.5/GPT-4 stories on identical Pygmalion prompts, finding AI stories more progressive on gender yet less imaginative.

  • EMNLP Findings 2023 A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing [paper]
    Carlos G'omez-Rodr'iguez, Paul Williams

    Compares LLMs and humans on an unusual comic epic prompt, finding top commercial LLMs match humans on most criteria except creativity.

  • C&C 2023 More human than human: LLM-generated narratives outperform human-LLM interleaved narratives [paper]
    Z. Zhao, Sophie Song, Bridget Duah, J. Macbeth, Scott A. Carter, Monica P. Van, N. Bravo, M. Klenk, Kate Sick, Alexandre L. S. Filipowicz

    Finds through two roughly 500-participant studies that readers prefer purely LLM-generated stories over human-LLM interleaved stories.

  • INLG 2023 The Next Chapter: A Study of Large Language Models in Storytelling [paper]
    Zhuohan Xie, Trevor Cohn, Jey Han Lau

    Finds that prompted LLMs write stories rivaling human authors and beating prior generators, though they sometimes replicate real stories.

๐Ÿค Co-creation

๐Ÿ› ๏ธ Tools

  • EMNLP Findings 2026 Generating Constructive Feedback on Stories via Reinforcement Learning [paper]
    Maja Stahl, Timon Ziegenbein, Henning Wachsmuth

    Trains LLMs with GRPO and a multi-component constructiveness reward to give story-specific feedback, finding actionable suggestions drive constructiveness most.

  • ArXiv 2026 GraphStory: Collaborative Story Writing through Event-Based Narrative Editing [paper]
    X. Lรช, Minh-Loi Nguyen, Khanh-Duy Le, Minh-Triet Tran, Trung-Nghia Le

    Builds a graph-based writing assistant for organizing plot points and exploring alternative branches, which writers found reduced structuring effort.

  • ArXiv 2026 Fabula: Building a Narrative Storytelling Sidekick with the Writers' Community [paper]
    Piotr Mirowski, Benjamin D. Wedin, Reinald Kim Amplayo, Rich Galt, Duncan Williams, Rida Qadri, Jaume Sanchez-Elias, Erin Drake-Kajioka, Sian Gooding, Lucรญa Lรณpez-Rivilla, et al.

    Designs and evaluates a narratology-based fiction writing app with 42 writers, probing auto-evaluators, plan-exposing interfaces, and cultural fit of story structures.

  • CHI 2026 Exploring Creator-Centric Methods for LLM-Assisted Interactive Storytelling [paper]
    Yue-Lu Li, Siyi Wu, Lu-Jin Zhang, Zhihan Guo, Wenchuan Lu, David Yip

    Designs CoNoder, a creator-centered LLM prototype for interactive narratives with node-graph editing, ripple-effect analysis, and simulated reader feedback, informed by creator interviews.

  • CHI 2026 Orchid-Creator: An Authoring Tool Supporting LLM-Driven Interactive Narrative Creation [paper]
    Zhen Wu, Serkan Kumyol, Zhengyang Ma, T. Braud

    Builds an LLM authoring tool representing interactive narratives as card-based story graphs, found easier for structuring than Twine and AI Dungeon.

  • CHI 2026 Plotania: Exploring Transparency Trade-offs in AI Co-Writing Through Virtual Readers and Transparent Attribution [paper]
    Yu-Feng Hu, Jinyi Zhang, Ze-Hua Wang, Chun Yu

    Builds a co-writing system with virtual reader reactions and AI attribution, finding transparency raises awareness but lowers creative agency and AI usage.

  • CHI 2026 Narrix: Remixing Narrative Strategies from Examples for Story Writing [paper]
    Chao Zhang, Shunan Guo, Abe Davis, Eunyee Koh

    Builds a writing tool that highlights narrative strategies in example stories and lets novices apply them to drafts via strategy-steered generation.

  • CHI 2026 NarrativeLoom: Enhancing Creative Storytelling through Multi-Persona Collaborative Improvisation [paper]
    Yuxi Ma, Yongqian Peng, Fengyuan Yang, Siyu Zha, Chi Zhang, Zi-Xia Jia, Zilong Zheng, Yixin Zhu

    Builds a multi-persona co-creative storytelling system based on blind variation and selective retention; experts rated co-authored stories more creative.

  • CHI 2026 PlayWrite: A Multimodal System for AI Supported Narrative Co-Authoring Through Play in XR [paper]
    Esen K. Tutuncu, Qian Zhou, Frederik Brudy, George W. Fitzmaurice, Fraser Anderson

    Builds a mixed-reality system where users author stories by manipulating virtual characters and props, which multi-agent AI turns into rearrangeable narrative beats.

  • CHI 2026 DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection [paper]
    Yuying Tang, Xinyi Chen, Haotian Li, Xing Xie, Xiaojuan Ma, Huamin Qu

    Builds a screenplay refinement system whose AI agent first simulates character experience, then evaluates it to give feedback that deepens screenwriters' reflection.

  • CHI 2026 Vidmento: Creating Video Stories through Context-Aware Expansion with Generative Video [paper]
    Catherine Yeh, Anh Truong, Mira Dontcheva, Bryan Wang

    Builds a tool that fills narrative gaps in video stories by generating context-aware clips that blend stylistically and narratively with captured footage.

  • CHI 2026 DiaryPlay: AI-Assisted Creation of Interactive Story Vignettes for Everyday Storytelling [paper]
    Jiangnan Xu, Haeseul Cha, Gosu Choi, Gyu-cheol Lee, Y. Yoon, Zucheul Lee, Konstantinos Papangelis, D. Kim, Juho Kim

    Builds an AI authoring system turning text stories into interactive vignettes, using LLM-controlled divergence to keep NPC behavior within the intended story.

  • ArXiv 2025 Elsewise: Authoring Open-ended Interactive Narrative with Possibility Space Visualization [paper]
    Yi Wang, John Joon Young Chung, Melissa Roemmele, Yuqian Sun, Tiffany Wang, S. Almeda, Brett A. Halperin, Yuwen Lu, Max Kreminski

    Builds an authoring tool that visualizes bundled storylines of LLM-driven interactive narratives, helping authors anticipate player-experienced stories in a 12-user study.

  • CHI 2025 Toward Personalizable AI Node Graph Creative Writing Support: Insights on Preferences for Generative AI Features and Information Presentation Across Story Writing Processes [paper]
    Hua-Xuan Qin, Guangzhi Zhu, Mingming Fan, Pan Hui

    Studies a FigJam plugin combining node-graph story structure, LLM audience impersonation, and image/audio generation for personalized story writing and moral reflection.

  • CHI 2025 WhatELSE: Shaping Narrative Spaces at Configurable Level of Abstraction for AI-bridged Interactive Storytelling [paper]
    Zhuoran Lu, Qian Zhou, Yi Wang

    Builds an authoring system deriving narrative possibility spaces from example stories, letting authors bound them and unfold them into game events.

  • CHI 2025 Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols [paper]
    John Joon Young Chung, Melissa Roemmele, Max Kreminski

    Builds a storytelling system where users steer LLM story text by moving character symbols like toys, via a shared motion-text semantic space.

  • CHI 2024 ๐ŸŒŸ CharacterMeet: Supporting Creative Writers' Entire Story Character Construction Processes Through Conversation with LLM-Powered Chatbot Avatars [paper]
    Hua-Xuan Qin, Shan Jin, Ze Gao, Mingming Fan, Pan Hui

    Builds a system letting writers develop characters by conversing with customizable chatbot avatars; a 14-writer study shows it supports iterative character construction.

  • Frontiers in Robotics and AI 2024 Fostering childrenโ€™s creativity through LLM-driven storytelling with a social robot [paper]
    Maha Elgarf, Hanan Salam, Christopher Peters

    Fine-tunes an LLM to drive creative or non-creative social robot storytelling, finding the creative robot boosts children's fluency, flexibility, and elaboration.

  • C&C 2024 Ai.llude: Encouraging Rewriting AI-Generated Text to Support Creative Expression [paper]
    David Zhou, S. Sterman

    Finds from 27 writing sessions that deliberately imperfect intermediate AI suggestions encourage writers to rewrite, supporting creative ownership and reflection.

  • ArXiv 2024 GhostWriter: Augmenting Collaborative Human-AI Writing Experiences Through Personalization and Agency [paper]
    Catherine Yeh, Gonzalo Ramos, Rachel Ng, Andy Huntington, R. Banks

    Builds GhostWriter, a writing design probe that implicitly learns user style while offering explicit controls, studying how it supports agency and personalization.

  • TiiS 2023 ๐ŸŒŸ ID.8: Co-Creating Visual Stories with Generative AI [paper] GitHub stars
    Victor Antony, Chien-Ming Huang

    Introduces ID.8, an open-source system for co-creating visual stories with generative AI, with a user study highlighting enjoyment and remaining gaps.

  • EACL 2023 Fiction-Writing Mode: An Effective Control for Human-Machine Collaborative Writing [paper]
    Wenjie Zhong, Jason Naradowsky, Hiroya Takamura, Ichiro Kobayashi, Yusuke Miyao

    Annotates narrative paragraphs with writing-mode labels and fine-tunes LLMs conditioned on these modes, finding authors prefer mode-controlled suggestions in collaborative fiction writing.

๐Ÿ‘ฅ User Studies

  • COLM 2026 The Garden of Forking Prompts: How Users Explore Narrative Space in Story Generation [paper]
    Advait Deshmukh, N. Benedict, Melanie Walsh, Maria Antoniak

    Analyzes how users iteratively revise story prompts in wild chatbot logs, releasing WildStories and WildEdits and an edit-type framework for benchmarking.

  • ArXiv 2026 AI Fiction in the Wild [paper]
    Neel Gupta, Maria Antoniak, Melanie Walsh

    Analyzes 500,000 ChatGPT conversations, finding over a third involve fiction generation, dominated by power users favoring fanfiction, erotica, and repetition.

  • CHI 2026 Proactive AI as a Catalyst for Creativity? Balancing Human Agency and AI Contribution in Collaborative Story Writing [paper]
    Yiwen Yin, Ming-Ze Wu, R. Huang, X. Tong, Jun Zhou, Chun Yu, Yuanchun Shi

    Wizard-of-Oz study of intrusive versus non-intrusive proactive AI suggestions in story outlining, revealing a creativity-agency trade-off moderated by how inspiring suggestions are.

  • ACL 2025 Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback [paper]
    Hannah Rashkin, Elizabeth Clark, Fantine Huot, Mirella Lapata

    Introduces a task and 1,300 deliberately corrupted stories to evaluate LLM writing feedback, finding models often miss the biggest writing issue.

  • Interacciรณn 2025 Once More with (the Right) Feeling: How Historical Fiction Writing Processes of Character Design, Plot Outline, and Context Checking Are Affected by Co-Writing with ChatGPT [paper]
    Yun Chen, Yiwei Wang, Antoni B. Chan, Jixing Li, LC Ray

    Examines how co-writing with ChatGPT affects historical fiction writers' character design, plot outlining, and context checking processes.

  • CHI 2025 Understanding Screenwriters' Practices, Attitudes, and Future Expectations in Human-AI Co-Creation [paper]
    Yuying Tang, Haotian Li, Minghe Lan, Xiao-Juan Ma, Huamin Qu

    Interviews 23 screenwriters on how they integrate AI across workflow stages and categorizes expected AI roles as actor, audience, expert, and executor.

  • CHI 2024 ๐ŸŒŸ Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models [paper]
    Paramveer S. Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub, Shaochun Zheng, L. P. Robert

    Finds with 131 participants a U-shaped effect of AI scaffolding: paragraph-level suggestions improve writing quality and productivity, while sentence-level ones do not.

  • CSCW 2024 'It was 80% me, 20% AI': Seeking Authenticity in Co-Writing with Large Language Models [paper]
    Angel Hsing-Chi Hwang, Q. Liao, Su Lin Blodgett, Alexandra Olteanu, Adam Trischler

    Interviews 19 professional writers and surveys readers on authenticity in AI co-writing, finding personalization should support writer growth beyond text production.

  • CHI 2024 The Value, Benefits, and Concerns of Generative AI-Powered Assistance in Writing [paper]
    Zhuo-Yan Li, Chen Liang, Jing Peng, Ming Yin

    Finds through an experiment that people will forgo payment for AI writing help, which boosts productivity and confidence but raises ownership concerns.

  • CHI 2023 ๐ŸŒŸ Social Dynamics of AI Support in Creative Writing [paper]
    Katy Ilonka Gero, Tao Long, Lydia B. Chilton

    Interviews 20 creative writers to identify what help they want, how they perceive supporters, and values shaping AI-versus-human support choices.

  • ArXiv 2023 Creativity Support in the Age of Large Language Models: An Empirical Study Involving Emerging Writers [paper]
    Tuhin Chakrabarty, Vishakh Padmakumar, Faeze Brahman, S. Muresan

    Studies 30 writers using an LLM interface based on the cognitive process model, finding LLMs most helpful for translating and reviewing.

  • IDC 2023 Design implications of generative AI systems for visual storytelling for young learners [paper]
    Ariel Han, Zhenyao Cai

    Elicits parent, teacher, and researcher views on generative AI for children's visual storytelling and proposes AIStory, a prototype app supporting literacy.

๐Ÿ“š Surveys

  • ArXiv 2026 Narrative Theory-Driven LLM Methods for Automatic Story Generation and Understanding: A Survey [paper]
    David Y. Liu, A. Joshi, Paul Dawson

    Surveys LLM story generation and understanding through narratology, finding generation lags understanding and recommending theory-based metrics over a single quality benchmark.

  • EMNLP Findings 2025 ๐ŸŒŸ A Survey on LLMs for Story Generation [paper]
    Maria Teleki, Vedangi Bengali, Xiangjue Dong, Sai Janjur, Haoran Liu, Tian Liu, Cong Wang, Ting-Yiu Liu, Yin Zhang, Frank Shipman, et al.

    Surveys LLM story generation, organizing work into autonomous generation versus author assistance and comparing methods, datasets, story types, and evaluations.

  • ArXiv 2024 ๐ŸŒŸ What Makes a Good Story and How Can We Measure It? A Comprehensive Survey of Story Evaluation [paper]
    Dingyi Yang, Qin Jin

    Surveys story evaluation across text-to-text, visual-to-text, and text-to-visual tasks, proposing a taxonomy of human criteria, benchmarks, and automatic metrics.

  • Neurocomputing 2023 Open-world Story Generation with Structured Knowledge Enhancement: A Comprehensive Survey [paper]
    Yuxin Wang, Jieru Lin, Zhiwei Yu, Wei Hu, Bรถrje F. Karlsson

    Surveys structured knowledge-enhanced story generation, offering a taxonomy of how knowledge is injected to improve coherence and grounding, plus future directions.

๐Ÿงฐ Public Resources

๐Ÿ“ฆ Datasets

  • WritingPrompts: about 300K human-written stories paired with Reddit writing prompts; the most widely used dataset for open-ended story generation.
  • VIST: photo sequences paired with human-written stories; the standard dataset for Visual2Story.
  • PG-19: full-length books from Project Gutenberg, a common source of long-form fiction.

๐Ÿ† Leaderboards

๐Ÿ›๏ธ Venues & Workshops

  • ICIDS: International Conference on Interactive Digital Storytelling.
  • AIIDE: AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment.
  • WNU: Workshop on Narrative Understanding, co-located with ACL conferences.
  • In2Writing: Workshop on Intelligent and Interactive Writing Assistants.
  • Wordplay: When Language Meets Games workshop.

๐Ÿ”— Related Lists

๐Ÿค Contributing

We welcome paper recommendations and corrections. Please read CONTRIBUTING.md for what the list includes and how to format an entry, then open a paper recommendation issue or a pull request.

๐Ÿ“ Citation

If you find this list useful, please consider citing it:

@misc{ma2023awesomestorygeneration,
  title        = {Awesome-Story-Generation: A Curated List of Papers on Story Generation in the Era of Large Language Models},
  author       = {Ma, Yingpeng and Ma, Yan},
  year         = {2023},
  howpublished = {\url{https://github.com/yingpengma/Awesome-Story-Generation}}
}

โญ Star History

Star History Chart

language-generation
large-language-model
narrative
natural-language-processing
nlp
novel
open-ended
story
story-generation
storytelling

Significant stargazers

Fredrik Norรฉn

309 followers ยท starred Oct 2024

flashmouse

23 followers ยท starred Apr 2026

yingpengma/Awesome-Story-Generation

This repository collects an extensive list of awesome papers about Story Generation / Storytelling, exclusively focusing on the era of Large Language Models (LLMs).

Python

657

126 commits

updated Sep 28, 2026

See the code

README

Awesome Story Generation

Awesome Papers Last commit Stars PRs welcome License

News Overview Papers Contributing Citation

Maintained by Yingpeng Ma and Yan Ma

A curated list of papers on story generation and storytelling in the era of large language models: long-form fiction, screenplays and drama, games, narrative world models, visual stories, and how to evaluate and co-create them. Every paper comes with a one-line summary, and papers we consider essential reading are marked with ๐ŸŒŸ.

Thank you for the stars! Contributions are very welcome: open an issue or PR for missing papers or mistakes. Contact: mayingpeng33 [AT] gmail [DOT] com

๐Ÿ“ฐ News

  • [2026-09] ๐ŸŽ‰ Major update: 234 papers, a new taxonomy, one-line summaries and ๐ŸŒŸ must-read picks.
  • [2026-05] ๐Ÿ”ฅ Our paper on long-horizon consistency in interactive narratives is accepted to ICML 2026! See it here.

๐Ÿ—บ๏ธ Overview

The list has four sections: Beyond Text (84 papers), Text Stories (71), Evaluation (41) and Co-creation (34), plus 4 surveys.
  • Beyond Text: interactive drama, games, narrative world models, screenplays, and visual stories.
  • Text Stories: planning, coherence, characters, creativity, and training for written stories.
  • Evaluation: benchmarks, metrics, and analyses of written stories.
  • Co-creation: tools for creators, and studies of how people write with AI.
  • Surveys: overviews of the whole field.

Each paper appears exactly once. Human-centered systems and studies go to Co-creation; work on other media goes to Beyond Text (including its evaluation); remaining work on written stories goes to Evaluation or to the Text Stories topic it mainly addresses. Visual work is included only when it operates at the story level (plot, script, shot planning, narrative reasoning), not when it only improves rendering quality or character consistency. Within a section, papers are sorted by year, with ๐ŸŒŸ must-reads first.

Papers per year: 34 in 2023, 42 in 2024, 62 in 2025 and 92 in 2026 through September, excluding 4 surveys.

Data table for the chart (2026 counts through September; the 4 surveys are not shown)
YearBeyond TextText StoriesEvaluationCo-creationTotal
20231667534
202413148742
2025241813762
2026*3133131592

๐Ÿ“‘ Table of Contents

๐Ÿ“„ Papers

How to read an entry: venue ยท citation count (refreshed weekly) ยท ๐ŸŒŸ must-read ยท title ยท [paper] ยท GitHub stars of the official code, when available, followed by authors and a one-line summary.

Venue colors: NLP ML Vision & Graphics AI HCI Games arXiv Other

๐ŸŽญ Beyond Text

๐ŸŽช Interactive Drama

  • ICML 2026 ๐ŸŒŸ Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives [paper] GitHub stars
    Yingpeng Ma, Jianhao Yan, Bei-Ning Shi, Karim Kam, Runnan Wang, Xue-Bo Liu, Yulong Chen, Yue Zhang, Derek F. Wong

    Introduces a 100-environment benchmark on whether LLM narrators keep story commitments under user interventions; even GPT-5.2 survives only 42% after 20 turns.

  • ArXiv 2026 NARRA-Gym for Evaluating Interactive Narrative Agents [paper]
    Yue Huang, Yu-Chen Ma, Jiayi Ye, Wen-Jie Wang, Zi-Peng Ling, Xing Hu, Yuexing Hao, Zi-Chen Chen, Zhangchen Xu, Yun-Hong He, et al.

    Introduces an executable environment growing emotional seeds into full interactive story episodes, showing fluent LLMs still fail on robustness and personalization.

  • ACL Findings 2026 AdaMARP: An Adaptive Multi-Agent Interaction Framework for General Immersive Role-Playing [paper]
    Zhenhua Xu, Dongsheng Chen, Shuo Wang, Jian Li, Chengjie Wang, Meng Han, Ya-Biao Wang

    Proposes a multi-agent role-play framework whose scene manager selects speakers, switches scenes, and introduces roles, with training data and a benchmark.

  • ICLR 2026 HAMLET: A Hierarchical and Adaptive Multi-Agent Framework for Live Embodied Theatrics [paper] GitHub stars [dataset]
    Shu-Fan Jiang, Si-Zhou Chen, Chios Chen, Chi Zhang, Xiao-Lei Zhang, Xue-Long Li

    Builds HAMLET, a multi-agent framework that turns a topic into a narrative blueprint and performs live embodied theatre with adaptive actor agents.

  • ACL 2025 ๐ŸŒŸ Towards Enhanced Immersion and Agency for LLM-based Interactive Drama [paper] GitHub stars
    Hongqiu Wu, Weiqi Wu, Tianyang Xu, Jiameng Zhang, Hai Zhao

    Proposes Playwriting-guided Generation and Plot-based Reflection to improve player immersion and agency in LLM-based interactive drama.

  • NAACL 2025 ๐ŸŒŸ CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds [paper] GitHub stars
    Lei Wang, Jian-Xun Lian, Yi Huang, Yanqi Dai, Haoxuan Li, Xu Chen, Xing Xie, Ji-Rong Wen

    Introduces a simulation sandbox with character and narrator agents that produces behavior trajectories for fine-grained evaluation of LLM role-playing.

  • Complex & Intelligent Systems 2025 ProTriPlay: A trinity framework for professional interactive theater based on LLM [paper]
    Yinglong Yu, Hao Shen, Ming Yang, Yu Wang, Yanyu Liu

    Builds an LLM interactive theater system with director, screenwriter, and actor agents that adapt the plot to player dialogue and object interactions.

  • AIIDE 2025 CoDi: A Director-Actor Framework for Goal-Driven Interactive Story Generation with LLMs [paper]
    Honggu Kim, Taewoo Yoo, Yun-Gyung Cheong

    Extends the director-actor paradigm so a director agent pursues high-level narrative goals by introducing events, selecting NPCs, and specifying outcomes.

  • EMNLP 2025 OPEN-THEATRE: An Open-Source Toolkit for LLM-based Interactive Drama [paper]
    Tianyang Xu, Hongqiu Wu, Weiqi Wu, Hai Zhao

    Releases Open-Theatre, an open-source toolkit for LLM interactive drama with multi-agent architecture and hierarchical retrieval-based memory for coherent long-term behavior.

  • ACL 2025 RolePlot: A Systematic Framework for Evaluating and Enhancing the Plot-Progression Capabilities of Role-Playing Agents [paper]
    Pinyi Zhang, Si-Yu An, Lingfeng Qiao, Yi-Fei Yu, Jing-Yang Chen, Jie Wang, Di Yin, Xing Sun, Kai Zhang

    Proposes a plot-progression dataset and method for role-playing agents, detecting an LLM embedding trigger subspace to prompt timely plot advances.

  • ACL Findings 2024 ๐ŸŒŸ From Role-Play to Drama-Interaction: An LLM Solution [paper]
    Weiqi Wu, Hongqiu Wu, Lai Jiang, Xing-Chen Liu, Jiale Hong, Haizhen Zhao, Min Zhang

    Defines LLM-based interactive drama and trains a drama LLM using Narrative Chain control, Auto-Drama script synthesis, and Sparse Instruction Tuning.

  • AAAI 2024 NarrativePlay: An Automated System for Crafting Visual Worlds in Novels for Role-Playing [paper]
    Run-Cong Zhao, Wenjia Zhang, Jiazheng Li, Lixing Zhu, Yanran Li, Yulan He, Lin Gui

    Presents a demo system that lets users role-play a novel character in LLM-generated narrative environments with generated visuals and speech.

๐ŸŽฒ Games

  • ArXiv 2026 When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations [paper]
    Yuqi Chen, Sixuan Li, Yunfeng Cai, Xueai Li, Kaiwen Yan, Ying Li

    Introduces a process benchmark for storytelling in evolving world simulations, finding generation length, canonical consistency, and narrative richness are distinct, competing capacities.

  • FDG 2026 Generating Clue-Driven Investigative Game Narratives with Large Language Models [paper]
    Vikram Kumaran, A. Smith, Wookhee Min, Randall Spain, Bradford W. Mott, James C. Lester

    Builds an LLM framework that generates solvable clue-driven investigative 3D game episodes around a deductive solution model guiding characters, clues, and dialogue.

  • ArXiv 2026 IVIE: A Neuro-symbolic Approach to Incremental and Validated Generation of Interactive Fiction Worlds [paper]
    Micaela Vaucher, Santiago Silveira, Santiago Gรณngora, Luis Chiruzzo

    Generates playable interactive fiction worlds in four incremental stages, letting LLMs make creative choices while symbolic validation keeps world state coherent.

  • IUI 2026 Guiding, Not Railroading: Design and Evaluation of a Multi-Agent System for Narrative Redirection in Role-playing Games [paper]
    Nicolai Hejlesen Jรธrgensen, Sarmilan Tharmabalan, Ilhan Aslan, Nicolai Brodersen Hansen, Timothy Merritt

    Builds a multi-agent RPG game master with a narrative graph and tests six redirection strategies; players prefer in-world redirection over hard denials.

  • ArXiv 2025 STORY2GAME: Generating (Almost) Everything in an Interactive Fiction Game [paper]
    E. Zhou, Shreyas Basavatia, M. Siam, Zexin Chen, Mark O. Riedl

    Builds STORY2GAME, which generates a story, populates a world, and writes action code from LLM-derived preconditions and effects for playable interactive fiction.

  • AIIDE 2024 NarrativeGenie: Generating Narrative Beats and Dynamic Storytelling with Large Language Models [paper]
    Vikram Kumaran, Jonathan Rowe, James C. Lester

    Builds NarrativeGenie, which turns a designer's story overview into a partially ordered event graph of narrative beats that adapts to player actions.

  • AIIDE 2024 PANGeA: Procedural Artificial Narrative Using Generative AI for Turn-Based, Role-Playing Video Games [paper]
    Stephanie Buongiorno, Lawrence J. Klinkert, Zixin Zhuang, Tanishq Chawla, Corey Clark

    Builds a system with memory, validation, and a Unity plug-in that keeps LLM-generated RPG content consistent with designer rules despite free-form input.

  • ArXiv 2024 Word2World: Generating Stories and Worlds through Large Language Models [paper] GitHub stars
    Muhammad Umair Nasir, Steven James, Julian Togelius

    Builds Word2World, which prompts LLMs to write a story, extract narrative elements, and place tiles to produce playable game worlds without fine-tuning.

  • IEEE ToG 2024 Generating Role-Playing Game Quests With GPT Language Models [paper]
    Susanna Vรคrtinen, Perttu Hรคmรคlรคinen, C. Guckelsberger

    Fine-tunes GPT-2 on a released dataset of 978 RPG quests, finding about one in five generated quest descriptions acceptable to players.

  • EMNLP 2024 Ontologically Faithful Generation of Non-Player Character Dialogues [paper]
    Nathaniel Weir, Ryan Thomas, Randolph D'Amore, Kellie Hill, Benjamin Van Durme, Harsh Jhamtani

    Introduces KNUDGE, a dataset from The Outer Worlds requiring lore-faithful, quest-revealing NPC dialogue trees, with supervised and in-context baselines leaving headroom.

  • AIIDE 2023 ๐ŸŒŸ SceneCraft: Automating Interactive Narrative Scene Generation in Digital Games with Large Language Models [paper]
    Vikram Kumaran, Jonathan Rowe, Bradford W. Mott, James C. Lester

    Proposes SceneCraft, an LLM framework that automates NPC interaction scenes to unfold authored plot events in narrative-centered games.

  • AIIDE 2023 ๐ŸŒŸ Language as Reality: A Co-Creative Storytelling Game Experience in 1001 Nights using Generative AI [paper]
    Yuqian Sun, Zhouyi Li, Ke Fang, Chang Hee Lee, A. Asadipour

    Presents 1001 Nights, a game where spoken keywords in co-created LLM tales materialize as in-game items, proposing the notion of AI-native games.

  • ACL 2023 ๐ŸŒŸ FIREBALL: A Dataset of Dungeons and Dragons Actual-Play with Structured Game State Information [paper] GitHub stars
    Andrew Zhu, Karmanya Aggarwal, Alexander H. Feng, Lara J. Martin, Chris Callison-Burch

    Releases a dataset of about 25,000 real Discord D&D sessions with true game state, showing state information improves LLM game-turn generation.

  • CHI 2023 Location-Aware Adaptation of Augmented Reality Narratives [paper]
    Wan-Wan Li, Changyang Li, Minyoung Kim, Haikun Huang, L. Yu

    Proposes an optimization approach that assigns real-world locations to AR story events and synthesizes a navigation graph across story branches.

  • CHI 2023 Personalized Quest and Dialogue Generation in Role-Playing Games: A Knowledge Graph- and Language Model-based Approach [paper]
    Trevor Ashby, Braden K Webb, G. Knapp, John Searle, Nancy Fulda

    Proposes a player-centered RPG quest and dialogue generator grounding content in a hand-crafted knowledge base and an LLM, approaching hand-crafted quest quality.

  • ACL 2023 I Cast Detect Thoughts: Learning to Converse and Guide with Intents and Theory-of-Mind in Dungeons and Dragons [paper]
    Pei Zhou, Andrew Zhu, Jennifer Hu, J. Pujara, Xiang Ren, Chris Callison-Burch, Yejin Choi, Prithviraj Ammanabrolu

    Trains a Dungeon Master model with RL that rewards guidance whose intent matches theory-of-mind predictions of player actions in D&D.

๐ŸŒ World Models

  • ArXiv 2026 WorldMind: Decoupled Game World Model for State-Aware NPC Behavior [paper] GitHub stars
    Zhi-Yang Deng, Bo-Ran Zhang, Dan Chen, Ye-Ying Jin

    Adds an explicit state-reconstruction and planning interface to a game world model so NPCs act on the game state before being rendered.

  • ArXiv 2026 FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling [paper]
    Jia-Long Zuo, Haotong Zuo, Shiwei Zhang, Xiang Wang, Chen Li, Nong Sang, Chang-Xin Gao, Xiang Bai

    Frames novel-to-film generation as building a persistent cinematic world model from prose, then rendering long multi-scene films from it.

  • ArXiv 2026 EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World [paper] GitHub stars
    Qing Zong, Yue (Sophie) Guo, Mengxi Yang, Yiwen Guo, Yangqiu Song

    Models interactive literary worlds as long-horizon co-evolution of characters and world state, with an open-schema framework and benchmark.

  • ArXiv 2026 ReactiveGWM: Steering NPC in Reactive Game World Models [paper] GitHub stars
    Zeqing Wang, Dan Chen, Zhaohu Xing, Zizhao Tong, Yinhan Zhang, Xingyi Yang, Ye-Ying Jin

    Decouples player control from NPC behavior in a game world model, so text prompts can steer how NPCs react to the player.

  • ArXiv 2026 ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling [paper] GitHub stars
    Yawen Luo, Xiao-Yu Shi, Junhao Zhuang, Yu-Tian Chen, Quande Liu, Xintao Wang, Pengfei Wan, Tian-Fan Xue

    Reformulates multi-shot video generation as causal next-shot prediction, letting users steer an unfolding story in real time via streaming prompts.

  • ICLR 2025 ๐ŸŒŸ Unbounded: A Generative Infinite Game of Character Life Simulation [paper]
    Jialu Li, Yuanzhen Li, Neal Wadhwa, Y. Pritch, David E. Jacobs, Michael Rubinstein, Mohit Bansal, Nataniel Ruiz

    Builds a generative infinite game in which players raise an autonomous character in an LLM-driven, image-generated world with open-ended, emergent mechanics.

  • ICCV 2025 AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction [paper] GitHub stars
    Junhao Cheng, Yu-Ying Ge, Yi-Xiao Ge, Jing Liao, Shan Ying

    Turns anime film characters into playable agents for open-ended life simulation, predicting multimodal game states to keep the generated world consistent.

๐ŸŽž๏ธ Screenplays

  • FAccT 2026 Do Language Models Pass the Bechdel Test? Auditing Gender Biases in LLM-Generated Screenplays [paper]
    Megha N. Govindu, Stephanie T. Wang, Sorelle A. Friedler, D. Metaxa

    Automates the Bechdel test and network analysis on LLM screenplays; human scripts pass more often, but all scripts show some representational bias.

  • ArXiv 2026 NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama [paper]
    Logan Mann, Abdur Rahman, M. Saifullah, Taaha Kazi, Vasu Sharma

    Introduces a multi-horizon audio-drama benchmark showing frontier LLMs degrade over long arcs, plus N-VSSM, a Mamba-2 latent world-state model sustaining consistency.

  • ArXiv 2026 One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems [paper]
    Yu-Fei Shi, Wei-Long Yan, Naixuan Huang, Yucheng Chen, Chenyu Zhang, Tao He, Si Yong Yeo, Ming Li

    Builds a hierarchical multi-agent pipeline turning a one-sentence idea into a short drama via debate-based scripting, 3D-grounded first frames, and reviewer loops.

  • ArXiv 2026 Text-to-Stage: Spatial Layouts from Long-form Narratives [paper]
    Jefferson Hernandez, Swarnadeep Saha, Chenxi Whitehouse, Sanjeel Parekh, Calvin Murdock, Yuliang Li, W. O. Brimijoin, V. Ithapu, I. Ananthabhotla

    Introduces the task of inferring stage layouts and movements from narrative text, with a dramaturgy-based evaluation suite and rejection-SFT plus GRPO training.

  • ArXiv 2026 COMIC: Agentic Sketch Comedy Generation [paper] GitHub stars
    Susung Hong, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz

    Builds an automated agent population mimicking studio roles to produce sketch comedy videos, using LLM critics aligned with YouTube viewer preferences.

  • ArXiv 2025 DramaBench: A Six-Dimensional Evaluation Framework for Drama Script Continuation [paper] GitHub stars
    Shijian Ma, Yun-Chien Huang, Yan Lin

    Introduces a drama script continuation benchmark scoring six dimensions via rules and LLM labeling, evaluating eight LLMs on 1,103 scripts.

  • ArXiv 2025 Beyond Direct Generation: A Decomposed Approach to Well-Crafted Screenwriting with LLMs [paper]
    Hang Lei, Shengyi Zong, Zhaoyan Li, Ziren Zhou, Hao Liu, Liang Yu

    Decouples screenplay writing into outline-to-prose then prose-to-screenplay stages with hybrid data synthesis, winning 75% against strong baselines per professional screenwriters.

  • ArXiv 2025 CML-Bench: A Framework for Evaluating and Enhancing LLM-Powered Movie Scripts Generation [paper]
    Mingzhe Zheng, Dingjie Song, Guanyu Zhou, Jun You, Jia-Hao Zhan, Xuran Ma, Xin-Yuan Song, Ser-Nam Lim, Qi-Feng Chen, Harry Yang

    Introduces a movie script benchmark scoring dialogue coherence, character consistency, and plot reasonableness, plus an instruction-based prompting strategy for better scripts.

  • ACL 2024 ๐ŸŒŸ IBSEN: Director-Actor Agent Collaboration for Controllable and Interactive Drama Script Generation [paper] GitHub stars
    Senyu Han, Lu Chen, Li-Min Lin, Zhen Xu, Kai Yu

    Proposes IBSEN, where a director agent steers actor agents and human players toward plot objectives to generate controllable drama scripts.

  • EMNLP Findings 2024 ๐ŸŒŸ HoLLMwood: Unleashing the Creativity of Large Language Models in Screenwriting via Role Playing [paper]
    Jing Chen, Xinyu Zhu, Cheng Yang, Chufan Shi, Ya-Dong Xi, Yuxiang Zhang, Junjie Wang, Jiashu Pu, Rongsheng Zhang, Yu-Jiu Yang, et al.

    Builds HoLLMwood, a screenwriting framework assigning LLMs Writer, Editor, and role-playing Actor roles to enrich characters and plots in generated screenplays.

  • CHI 2023 ๐ŸŒŸ Co-Writing Screenplays and Theatre Scripts with Language Models: An Evaluation by Industry Professionals [paper]
    Piotr Wojciech Mirowski, K. Mathewson, Jaylen Pittman, Richard Evans

    Builds Dramatron, which hierarchically prompts LLMs to co-write scripts and screenplays, evaluated in a study with 15 theatre and film professionals.

๐Ÿ–ผ๏ธ Visual2Story

  • TACL 2026 Generating Visual Stories with Grounded and Coreferent Characters [paper] GitHub stars
    Danyang Liu, Mirella Lapata, Frank Keller

    Presents a character-centric visual storytelling model trained on VIST enriched with visual and textual coreference chains, plus metrics for character richness.

  • COLING 2025 ๐ŸŒŸ StoryLLaVA: Enhancing Visual Storytelling with Multi-Modal Large Language Models [paper]
    Li Yang, Zhihao Xiao, Wen-Xin Huang, Xian Zhong

    Proposes a visual storytelling MLLM trained with a topic-driven narrative optimizer for data refinement and preference-based ranked story sampling for alignment.

  • ICCV AISTORY Workshop 2025 Re:Verse -- Can Your VLM Read a Manga? [paper] GitHub stars
    Aaditya Baranwal, Madhav Kataria, Naitik Agarwal, Y. Rawat, Shruti Vyas

    Introduces a manga benchmark of 308 annotated panels showing VLMs interpret single panels well but fail at temporal causality and cross-panel reasoning.

  • ArXiv 2025 VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs? [paper]
    M. Gado, Towhid Taliee, M. Memon, Dmitry Ignatov, R. Timofte

    Adapts large multimodal models to visual storytelling on VIST and advocates reference-free metrics RoViST and GROOVIST over BLEU-style evaluation.

  • ICCV 2025 From Panels to Prose: Generating Literary Narratives from Comics [paper] GitHub stars
    Ragav Sachdeva, Andrew Zisserman

    Builds a system that converts manga into literary prose for visually impaired readers, introducing the Magiv3 comic-understanding model and annotated panel captions.

  • EMNLP Findings 2024 Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition [paper] GitHub stars
    Aditya K Surikuchi, Raquel Fernรกndez, Sandro Pezzelle

    Proposes a human-likeness metric over visual grounding, coherence, and repetition, finding a small upgraded TAPM rivals LLaVA, yet good stories need more.

  • ACL 2024 Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline [paper]
    Dingyi Yang, Chunru Zhan, Ziheng Wang, Biao Wang, Tiezheng Ge, Bo Zheng, Qin Jin

    Introduces synchronized video storytelling, generating clip-aligned narrations of fitting length, with the E-SyncVidStory dataset and a storyline-guided VideoNarrator framework.

  • LREC-COLING 2024 TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling [paper]
    Wei-Ran Chen, Xin Li, Jiaqi Su, Guiqian Zhu, Ying Li, Yi Ji, Chunping Liu

    Proposes a visual storytelling model that extracts visual and linguistic topic information and uses two topic-consistency reinforcement learning rewards on VIST.

  • EACL 2024 SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling [paper]
    E. Wang, Caren Han, Josiah Poon

    Proposes a visual storytelling framework that builds a social-commonsense plot graph from images and derives storylines via weighted shortest paths with Floyd-Warshall.

  • TACL 2023 ๐ŸŒŸ Visual Writing Prompts: Character-Grounded Story Generation with Curated Image Sequences [paper]
    Xudong Hong, A. Sayeed, K. Mehra, Vera Demberg, B. Schiele

    Introduces a dataset of about 2K curated movie-shot sequences with 12K character-grounded crowdsourced stories, plus a coherence-driven character-based generation baseline.

  • EMNLP Findings 2023 DiffuVST: Narrating Fictional Scenes with Global-History-Guided Denoising Models [paper]
    Shengguang Wu, Mei Yuan, Qi Su

    Proposes a non-autoregressive diffusion model that generates visual story narrations for fictional image sequences with bidirectional history guidance, improving speed and diversity.

  • EMNLP Findings 2023 Sound of Story: Multi-modal Storytelling with Audio [paper]
    Jaeyeon Bae, Seokhoon Jeong, Seokun Kang, Namgi Han, Jae-Yon Lee, Hyounghun Kim, Taehwan Kim

    Introduces Sound of Story, a dataset of 27K stories pairing image-text sequences with background audio, plus cross-modal retrieval and audio generation benchmarks.

  • EMNLP 2023 GROOViST: A Metric for Grounding Objects in Visual Storytelling [paper] GitHub stars
    Aditya K Surikuchi, Sandro Pezzelle, Raquel Fernรกndez

    Proposes a modular, interpretable metric for visual storytelling that measures how well stories are grounded in image entities, handling temporal misalignment.

  • EMNLP Findings 2023 Visual Storytelling with Question-Answer Plans [paper]
    Danyang Liu, Mirella Lapata, Frank Keller

    Proposes visual storytelling that feeds images as a visual prefix to a pretrained language model and plans with question-answer blueprints.

  • ACL 2023 Attractive Storyteller: Stylized Visual Storytelling with Unpaired Text [paper]
    Dingyi Yang, Qin Jin

    Introduces stylized visual storytelling and a memory-augmented multitask model trained with unpaired style text to generate styled stories from photo streams.

  • EACL 2023 Multimodal Event Transformer for Image-guided Story Ending Generation [paper]
    Yucheng Zhou, Guodong Long

    Proposes an event-graph reasoning transformer for image-guided story ending generation, with cross-modal fusion, a multimodal injector, and incoherence detection.

  • ACL Findings 2023 Visual Coherence Loss for Coherent and Visually Grounded Story Generation [paper]
    Xudong Hong, Vera Demberg, A. Sayeed, Qiankun Zheng, B. Schiele

    Proposes a coherence-theory-inspired self-supervised loss and combined object and face features for character representation, plus a character matching metric for visual storytelling.

๐ŸŒ„ Story2Visual

  • EACL 2026 ๐ŸŒŸ MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling [paper]
    Qian Wang, Zi-Qi Huang, Ruoxi Jia, Paul E. Debevec, Ning Yu

    Proposes a multi-agent pipeline spanning scripting, shot design, character modeling, keyframes, animation, and audio for long-sequence video storytelling.

  • ECCV 2026 Better Call CineCrew: Consistent Ultra-Long Narrative-to-Film Generation [paper]
    Jiaben Chen, Si Dong, Qinhong Zhou, Raine Ma, Zhi-Yang Dou, Wojciech Matusik, Chuang Gan

    Proposes a multi-agent orchestration layer built on FilmDSL, a film-specific language making shot, continuity, and persona constraints explicit for long script-to-video generation.

  • TPAMI 2026 Learning Long-form Movie Prior via Large Language Models. [paper]
    Jin-Heng Xie, Jia-Jun Feng, M. Shou

    Represents movies as text and bounding-box or keypoint tokens and curates Storyboard20K, letting LLMs learn movie priors to sample storyboards.

  • ArXiv 2026 SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution [paper] GitHub stars
    Mao-Lin Ran, Xiaoyan Lu, Jia-Qi Liu, Jian Wang, Weiwen Liu, Jianghao Lin, Yong Yu, Weinan Zhang

    Builds a deployed storyboarding system that learns directing rules from expert examples, evolves them via attribution feedback, and releases the PROSE dataset.

  • ArXiv 2026 AniMaster: From Story Texts to Animated Videos via Cinematic Script Generation and Interactive Authoring [paper]
    Ruiqi Yu, De-Kun Qian, Jia-Le Xu, Si-Zhe Cheng, Yize Li, Xiang-yang Wu, Zhiguang Zhou, Wei Chen, Yong Wang

    Builds an authoring tool that expands brief story texts into cinematic scripts, then animated videos, guided by a three-layer design framework.

  • EMNLP 2026 MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation [paper]
    Mu-Yao Wang, Ze-Ke Xie, Yanhao Chen, Lixin Xiu, Hideki Nakayama

    Proposes an agentic story-to-manga framework decomposing creation into planning, grounding, layout, rendering, composition, and lettering for controllable page generation.

  • ACL Findings 2026 BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration [paper] GitHub stars
    Bo Gao, Chang Liu, Yu-Yang Miao, Siyuan Ma, S. Lim

    Proposes a safety-aware multi-agent framework for end-to-end illustrated storybook generation with page-level text-image calibration and global consistency repair.

  • ArXiv 2026 CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding [paper]
    I. Mondal, Yi-Wen Song, Mihir Parmar, Palash Goyal, J. Boyd-Graber, Tomas Pfister, Yale Song

    Proposes a multi-agent storyboarding framework that plans character, background, and location continuity, plus a new long-range consistency benchmark.

  • ArXiv 2026 LogiStory: A Logic-Aware Framework for Multi-Image Story Visualization [paper]
    Chutian Meng, Fan Ma, Chi Zhang, Jiaxu Miao, Yi Yang, Yue-Ting Zhuang

    Proposes a multi-agent story visualization framework that grounds roles, extracts causal chains, and verifies consistency to model visual logic explicitly.

  • ArXiv 2026 EmoStory: Emotion-Aware Story Generation [paper]
    Jing-Yuan Yang, Rucong Chen, Weibin Luo, Hui Huang

    Introduces emotion-aware visual story generation and a two-stage framework combining agent-based planning with region-aware generation for emotional, subject-consistent image sequences.

  • CVPR 2026 Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning [paper]
    Zhengjian Yao, Yong-Zhi Li, Xinyu Gao, Quan Chen, Peng Jiang, Yan Lu

    Combines an MLLM narrative planner with a memory-bank control module for long consistent visual sequences and releases a 330K-image e-commerce storyboard dataset.

  • ArXiv 2026 MUSE: A Multi-agent Framework for Unconstrained Story Envisioning via Closed-Loop Cognitive Orchestration [paper]
    Wenzhang Sun, Zhenyu Wang, Zhang-Chi Hu, Chun-Feng Wang, Hao Li, Wei Chen

    Proposes a multi-agent plan-execute-verify-revise loop for long audio-visual stories from short prompts, plus a reference-free evaluation protocol.

  • ICCV Workshop 2025 ๐ŸŒŸ SEED-Story: Multimodal Long Story Generation with Large Language Model [paper] GitHub stars
    Shuai Yang, Yu-Ying Ge, Yang Li, Yu-Kang Chen, Yixiao Ge, Shan Ying, Ying-Cong Chen

    Proposes SEED-Story, an MLLM generating interleaved text and consistent images for long stories via multimodal attention sinks, and releases the StoryStream dataset.

  • ArXiv 2025 Generating Storytelling Images with Rich Chains-of-Reasoning [paper]
    Xiujie Song, Qi Jia, Shota Watanabe, Xiao-Yi Pang, Ruijie Chen, Mengyue Wu, Ke Zhu

    Defines storytelling image generation with chains of visual reasoning clues and proposes an LLM-plus-text-to-image pipeline with dedicated evaluation metrics.

  • ACM MM 2025 From Outline to Detail: An Hierarchical End-to-end Framework for Coherent and Consistent Visual Novel Generation and Assembly [paper]
    Yilin Zhang, Yanyan Wei, Zhao Zhang, Jicong Fan, Haijun Zhang, Shui-Cheng Yan

    Proposes an outline-guided pipeline that generates and assembles executable visual novels, using vision-LLM self-correction for cross-modal consistency and script validation.

  • EMNLP 2025 LLMs Behind the Scenes: Enabling Narrative Scene Illustration [paper]
    Melissa Roemmele, John Joon Young Chung, Taewook Kim, Yuqian Sun, Alex Calderwood, Max Kreminski

    Uses LLMs to prompt text-to-image models for narrative scene illustration and releases SceneIllustrations, a dataset of pairwise human quality judgments.

  • SIGGRAPH Asia 2025 AniMaker: Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation [paper] GitHub stars
    Haoyuan Shi, Yunxin Li, Xinyu Chen, Long-Yue Wang, Bao-Tian Hu, Min Zhang

    Proposes AniMaker, a multi-agent animation framework using MCTS-driven multi-candidate clip generation and the AniEval evaluator to produce story-coherent long videos from text.

  • CVPR 2025 VinaBench: Benchmark for Faithful and Consistent Visual Narratives [paper]
    Silin Gao, Sheryl Mathew, Li Mi, Sepideh Mamooler, Mengjie Zhao, Hiromi Wakaki, Yuki Mitsufuji, Syrielle Montariol, Antoine Bosselut

    Introduces a benchmark annotating commonsense and discourse constraints in visual narratives, with metrics for consistency and text alignment of generated image sequences.

  • ArXiv 2025 MM-StoryAgent: Immersive Narrated Storybook Video Generation with a Multi-Agent Paradigm across Text, Image and Audio [paper] GitHub stars
    Xuenan Xu, Jiahao Mei, Chenliang Li, Yuning Wu, Ming Yan, Shaopeng Lai, Ji Zhang, Mengyue Wu

    Proposes a multi-agent framework combining LLMs with image, speech, music, and sound tools to generate narrated storybook videos for children.

  • ACL Findings 2025 VISIAR: Empower MLLM for Visual Story Ideation [paper]
    Zhaoyang Xia, Somdeb Sarkhel, Md Mehrab Tanjim, Stefano Petrangeli, Ishita Dasgupta, Yuxiao Chen, Jinxuan Xu, Di Liu, Saayan Mitra, Dimitris N. Metaxas

    Introduces visual story ideation, arranging visual assets into storylines, with an MLLM framework using a story graph and a VTravel benchmark.

  • ACL 2023 Multimodal Persona Based Generation of Comic Dialogs [paper]
    Harsh Agrawal, A. Mishra, Manish Gupta, M. -

    Introduces multimodal persona-based comic dialogue generation with a 54K-strip dataset and an architecture that generates next-panel dialogues.

โœ๏ธ Text Stories

๐Ÿ—บ๏ธ Planning

  • ArXiv 2026 A Multi-Framework Comparison of Outline Stages in Long-Form Generation with LLMs [paper]
    Yikang Song

    Benchmarks outlines from seven long-form generation frameworks with an anchored LLM judge, finding no framework dominates across chapter and book granularities.

  • TASLP 2026 LLM-Driven MCTS for Conditional Story Generation via Logic-Guided Evidence Tree Optimization [paper]
    Hongyan Wu, Zhiliang Tian, Zhen Huang, Nankai Lin, Yi-Ping Song, Zhihua Wen, Menglong Lu, Feng Liu, Dongsheng Li

    Proposes a plug-and-play MCTS planner that builds logic-validated evidence chains for retrieval-based conditional story generation to reduce incoherence and thematic drift.

  • ACL Findings 2026 Planning Beyond Text: Graph-based Reasoning for Complex Narrative Generation [paper]
    Hanwen Gu, Chao Guo, Junle Wang, Wen-Da Xie, Yi-Sheng Lv

    Proposes PLOTTER, which runs an Evaluate-Plan-Revise cycle on event and character graphs to fix causality and structure before generating full narrative text.

  • ArXiv 2026 BiT-MCTS: A Theme-based Bidirectional MCTS Approach to Chinese Fiction Generation [paper]
    Zhaoyi Li, Xu Zhang, Xiaojun Wan

    Generates Chinese fiction by writing the climax first, then expanding plot backward and forward with bidirectional MCTS inspired by Freytag's Pyramid.

  • TACL 2026 Lightweight Latent Reasoning for Narrative Tasks [paper]
    Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

    Proposes a lightweight reasoning projector producing continuous latent tokens that RL policies toggle, cutting reasoning length on plot-hole detection and chapter generation.

  • NAACL 2025 ๐ŸŒŸ Generating Long-form Story Using Dynamic Hierarchical Outlining with Memory-Enhancement [paper]
    Qianyue Wang, Jinwu Hu, Zhengpin Li, Yufeng Wang, Daiyuan Li, Yu Hu, Mingkui Tan

    Proposes DOME, which fuses planning and writing through dynamic hierarchical outlines and uses a memory module to reduce contradictions in long stories.

  • ICLR 2025 ๐ŸŒŸ Agents' Room: Narrative Generation through Multi-step Collaboration [paper]
    Fantine Huot, Reinald Kim Amplayo, J. Palomaki, Alice Shoshana Jakobovits, Elizabeth Clark, Mirella Lapata

    Proposes Agents' Room, which splits fiction writing into subtasks for specialized agents, and releases the Tell Me A Story dataset and evaluation.

  • CIKM 2025 StoryWriter: A Multi-Agent Framework for Long Story Generation [paper]
    Haotian Xia, Hao Peng, Yunjia Qi, Bin Xu, Juan-Zi Li, Hou Lei, Xiaozhi Wang

    Proposes a multi-agent long story framework and uses it to build a 6,000-story dataset for fine-tuning Llama3.1-8B and GLM4-9B.

  • ACL Findings 2025 STORYTELLER: An Enhanced Plot-Planning Framework for Coherent and Cohesive Story Generation [paper]
    Jiaming Li, Yu-Kun Chen, Ziqiang Liu, Minghuan Tan, Lei Zhang, Yunshui Li, Run Luo, Long-Ze Chen, Jing Luo, A. Argha, et al.

    Proposes a plot-planning approach using SVO-triplet plot nodes plus interacting storyline and narrative entity knowledge graph modules for coherent story generation.

  • EMNLP 2025 Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language Models [paper] GitHub stars
    Ruibin Xiong, Yi-Meng Chen, Dmitrii Khizbullin, Mingchen Zhuge, Jurgen Schmidhuber

    Proposes a writing agent that recursively interleaves retrieval, reasoning, and composition tasks instead of fixed outlining, evaluated on fiction and technical reports.

  • ACL Findings 2025 A Cognitive Writing Perspective for Constrained Long-Form Text Generation [paper] GitHub stars
    Kaiyang Wan, Hong-Lin Mu, Rui Hao, Haoran Luo, Tianle Gu, Xiuying Chen

    Proposes CogWriter, a training-free framework applying Cognitive Writing Theory via planning, parallel generation, and review agents for constrained long-form text.

  • NAACL 2025 Navigating the Path of Writing: Outline-guided Text Generation with Large Language Models [paper]
    Yukyung Lee, Soonwon Ka, Bokyung Son, Pilsung Kang, Jaewook Kang

    Proposes WritingPath, which guides LLMs with explicit outlines reflecting user intent, and builds a blog-post dataset and evaluation framework for goal-oriented writing.

  • ACL 2024 Ex3: Automatic Novel Writing by Extracting, Excelsior and Expanding [paper] GitHub stars
    H. Lei, Jiaming Guo, Guanhua He, Xi-Shan Zhang, Rui Zhang, Shaohui Peng, Shaoli Liu, Tianshi Chen

    Proposes Ex3, which extracts structure from raw novels to build instruction data, fine-tunes an LLM, and expands tree-like into arbitrarily long novels.

  • AAAI 2024 Does Robin Hood Use a Lightsaber?: Automated Planning for Storytelling [paper]
    Nisha Ingrid Simon

    Combines automated planning with LLM text generation, using a planning model as scaffolding to produce more logical, coherent, and believable stories.

  • EMNLP Findings 2024 SWAG: Storytelling With Action Guidance [paper] GitHub stars
    Zeeshan Patel, Karim El-Refai, Jonathan Pei, Tianle Li

    Proposes SWAG, framing story writing as search where an auxiliary LLM picks the next action steering the generator toward engaging stories.

  • LREC-COLING 2024 Little Red Riding Hood Goes around the Globe: Crosslingual Story Planning and Generation with Large Language Models [paper]
    E. Razumovskaia, Joshua Maynez, Annie Louis, Mirella Lapata, Shashi Narayan

    Introduces crosslingual story generation with planning and a dataset, finding three-act plans yield more coherent, interesting, controllable stories across languages.

  • ACL 2023 ๐ŸŒŸ DOC: Improving Long Story Coherence With Detailed Outline Control [paper]
    Kevin Yang, D. Klein, Nanyun Peng, Yuan-Dong Tian

    Improves long-story plot coherence by generating a detailed hierarchical outline and a controller that keeps drafted passages aligned with outline details.

๐Ÿงต Coherence

  • IJCAI 2026 FossilWriter: Learning Hypergraph World Models with Latent Narratives for Creative Story Generation [paper]
    Heng Zhang, Yi-Hao Zhong, Lubin Gan, Zhihe Chen, Tianyi Zhang, Jing Liu, Jin Huang

    Grows a hypergraph world model whose unresolved elements seed latent narratives, improving plot coherence and reducing long-range factual conflicts in story generation.

  • ArXiv 2026 Narrative World Model: Narratology-Grounded Writer Memory for Long-Form Fiction [paper]
    M. Saifullah, Thomas Kornmaier, Taaha Kazi, Vasu Sharma, A. Kanade, Aanand Kumar Yadav

    Proposes a writer-memory system pairing a narratology-typed temporal state graph with hybrid retrieval, outperforming Graphiti/Zep and GraphRAG on multi-hop story questions.

  • EMNLP Findings 2026 ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control [paper] GitHub stars
    Jindong Li, Yang Yang, Zihao Liu, Yutao Yue, Meng-Lin Yang

    Proposes a training-free scene-by-scene writer that tracks symbolic story states, checks narrative transitions, and uses uncertainty signals to repair inconsistencies.

  • AAAI 2026 Octopus: Entropy-Controlled Science Fiction Literature Generation with Persistent Memory-Context Binding [paper]
    Xu Wang, Jiaju Kang, Puyu Han, Zeyu Ai, Luqi Gong

    Proposes Octopus, combining entropy regulation via narrative divergence thresholds with hierarchical memory of characters, plots, and scientific rules for long sci-fi generation.

  • WWW 2025 SCORE: Story Coherence and Retrieval Enhancement for AI Narratives [paper]
    Qiang Yi, Yang-Fan He, Jian-Hui Wang, Xin-Yuan Song, Shi-Yao Qian, Xin-Hang Yuan, Yi Xin, Yi-Jin Wang, Jingqun Tang, Yuchen Li, et al.

    Proposes SCORE, which tracks key item states and episode summaries and uses retrieval-augmented generation to detect and fix inconsistencies in LLM-generated stories.

  • COLING 2025 MLD-EA: Check and Complete Narrative Coherence by Introducing Emotions and Actions [paper]
    Jin-Ming Zhang, Yun-Fei Long

    Proposes MLD-EA, which uses LLMs with emotion and action cues to detect missing logic in narratives and generate sentences that restore coherence.

  • NAACL 2025 FACTTRACK: Time-Aware World State Tracking in Story Outlines [paper]
    Zhiheng Lyu, Kevin Yang, Lingpeng Kong, Daniel Klein

    Proposes FACTTRACK, which decomposes events into atomic facts with time-aware validity intervals to track world state and detect contradictions in story outlines.

  • COLM 2024 With Greater Text Comes Greater Necessity: Inference-Time Training Helps Long Text Generation [paper] GitHub stars
    Yan Wang, D. Ma, Deng Cai

    Proposes Temp-Lora, which stores long context in a temporary LoRA module trained during generation, improving long-text quality while cutting context-window costs.

  • ArXiv 2023 ๐ŸŒŸ RecurrentGPT: Interactive Generation of (Arbitrarily) Long Text [paper] GitHub stars
    Wangchunshu Zhou, Y. Jiang, Peng Cui, Tiannan Wang, Zhenxin Xiao, Yifan Hou, Ryan Cotterell, Mrinmaya Sachan

    Proposes RecurrentGPT, which simulates LSTM-style recurrence with natural-language long- and short-term memories so LLMs can interactively generate arbitrarily long text.

๐Ÿง‘โ€๐Ÿคโ€๐Ÿง‘ Characters

  • ArXiv 2026 ANIMASK: What the Model Contributes to Role Play in Simulated Story Worlds [paper]
    Xiu-Cheng Zhang, Zhuo-Ning Xu, Han-Jun Luo, Yankai Chen, Hanan Salam, Xue Liu

    Replays story worlds from freeze points with and without personas, finding actor LLMs push characters toward cautious, flatter outcomes than canon.

  • ArXiv 2026 From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives [paper]
    Aayush Aluru, C. Ho, Muhammad Hammouri, Kerry Luo, Myra Malik, Ryan Lagasse, Arjun Bahuguna, Vasu Sharma

    Proposes multi-agent persona-driven story generation with shared world state plus a graph-based hallucination detector, halving hallucinations in 100-page stories.

  • ArXiv 2026 Staying In Character: Perspective-Bounded Memory For Book-Based Role-Playing Agents [paper]
    Xushuo Tang, Junhe Zhang, Zi-Han Yang, Yi-Fu Tang, Sichao Li, Longbin Lai, Zheng-Yi Yang

    Proposes a three-layer, perspective-bounded memory for book-based role-playing agents that prevents characters using unknown facts, with a 4,386-question knowledge-boundary benchmark.

  • ACL 2026 EvoSpark: Endogenous Interactive Agent Societies for Unified Long-Horizon Narrative Evolution [paper]
    Shiyu He, Min-Chi Kuang, Mengxian Wang, Bin Hu, Tingxiang Gu

    Proposes a multi-agent framework with stratified narrative memory and role-location-plot alignment to sustain coherent, open-ended long-horizon story evolution.

  • ACL 2026 Deriving Character Logic from Storyline as Codified Decision Trees [paper] GitHub stars
    Letian Peng, Kun Zhou, Longfei Yun, Yu-Peng Hou, Jingbo Shang

    Induces executable, interpretable decision trees of validated scene-conditioned behavior rules from narrative data to ground role-playing agents more reliably.

  • AAAI 2026 StoryBox: Collaborative Multi-Agent Simulation for Hybrid Bottom-Up Long-Form Story Generation Using Large Language Models [paper]
    Zehao Chen, Rong Pan, Hao-Ran Li

    Proposes bottom-up long-form story generation in which multi-agent sandbox simulation yields emergent events that form coherent stories exceeding 10,000 words.

  • ArXiv 2025 ๐ŸŒŸ BookWorld: From Novels to Interactive Agent Societies for Creative Story Generation [paper] GitHub stars
    Yi-Ting Ran, Xintao Wang, Tian Qiu, Jiaqing Liang, Yanghua Xiao, Deqing Yang

    Builds BookWorld, which simulates multi-agent societies from established novels' characters and worldviews to generate creative stories faithful to the source books.

  • AIIDE 2025 Steering Narrative Agents Through a Dynamic Cognitive Framework for Guided Emergent Storytelling [paper]
    Chen Yang, M. Gross, R. Wampfler

    Proposes a cognitive agent framework where tensions between agents' beliefs and ideal worlds drive actions, steering emergent stories toward authored storylines.

  • FDG 2024 ๐ŸŒŸ StoryVerse: Towards Co-authoring Dynamic Plot with LLM-based Character Simulation via Narrative Planning [paper]
    Yi Wang, Qian Zhou, David Ledo

    Proposes StoryVerse, where authors write abstract acts that LLM narrative planning turns into character actions, balancing authorial intent with emergent game plots.

๐ŸŽจ Creativity

  • ArXiv 2026 MUSE: A Theory-Harnessed Story Engine for Vibe Narrativizing [paper] GitHub stars
    Jian-Xiang Ma, Xiaocui Yang, Da-Ling Wang, Yue-Song Hou, Ming-Fu Zhang, Yi-Chen Gao, Jun-Zhao Huang

    Builds a story engine encoding McKee's story theory as atomized rules inside an agent harness, improving WritingBench and consistency across four models.

  • ArXiv 2026 CraftAlign: Feature-Grounded Evaluation and Revision Guidance for AI Stories [paper]
    Yang Yang, Boyun Xu, Shaofeng Liang, Yun Han, Zining Zhong, Songning Lai, Kaishen Yuan, Yutao Yue

    Predicts 304 writing features to score stories against human and AI patterns, turning feature shifts into revision guidance that reduces AI flavor.

  • ArXiv 2026 StorySpark: Module-wise Evolutionary Search for Story Premise Generation [paper]
    Yang Yang, Zining Zhong, Qian Cao, Jindong Li, Boyun Xu, Kaishen Yuan, Meng-Lin Yang, Yutao Yue

    Proposes module-wise evolutionary search over premise components like persona, event, and twist, producing more original premises that yield better downstream stories.

  • ArXiv 2026 PlotTwist: A Creative Plot Generation Framework with Small Language Models [paper]
    A. Thorat, Ravi Kolla, Jyotin Goel, Madhav Kataria, N. Pedanekar

    Proposes a framework where sub-3B models generate premise-conditioned plots using an aspect reward model, DPO-aligned MoE generator, and cross-family jury evaluation.

  • ArXiv 2026 LLM Review: Enhancing Creative Writing via Blind Peer Review Feedback [paper] GitHub stars
    Weiyue Li, Mingxiao Song, Zhenda Shen, Dachuan Zhao, Yunfan Long, Yi Li, Yongce Li, Ruyi Yang, Meng-Yu Wang

    Proposes blind peer review among LLM agents that exchange feedback but revise independently, avoiding homogenization, and introduces the SciFi-100 writing dataset.

  • ACL 2026 Frankentext: Stitching random text fragments into long-form narratives [paper] GitHub stars
    Chau Minh Pham, Jenna Russell, Dzung Pham, Mohit Iyyer

    Proposes generating long narratives by having LLMs stitch mostly verbatim human-written fragments, improving diversity and originality while often evading AI-text detectors.

  • EMNLP 2025 Avoidance Decoding for Diverse Multi-Branch Story Generation [paper]
    Kyeongman Park, Nakyeong Yang, Kyomin Jung

    Proposes a decoding strategy that penalizes concept- and narrative-level similarity to earlier outputs, increasing diversity across multiple story branches from one prompt.

  • ACL Findings 2025 A Character-Centric Creative Story Generation via Imagination [paper]
    Kyeongman Park, Minbeom Kim, Kyomin Jung

    Proposes character-centric story generation that uses text-to-image imagination of story elements and multi-writer persona selection to deepen characters and creativity.

  • EMNLP 2024 ๐ŸŒŸ Collective Critics for Creative Story Generation [paper] GitHub stars
    Minwook Bae, Hyounghun Kim

    Proposes CritiCS, where a group of LLM critics collectively revise story plans and text to make long stories more creative and expressive.

  • ACL 2024 ๐ŸŒŸ MoPS: Modular Story Premise Synthesis for Open-Ended Automatic Story Generation [paper] GitHub stars
    Yan Ma, Yu Qiao, Pengfei Liu

    Proposes MoPS, which composes story premises from modular elements like background and persona, yielding more diverse and original premises for story generation.

  • EACL 2024 ๐ŸŒŸ Creating Suspenseful Stories: Iterative Planning with Large Language Models [paper]
    Kaige Xie, Mark Riedl

    Proposes a zero-shot iterative prompting planner grounded in cognitive-psychology and narratology theories of suspense to generate suspenseful stories with LLMs.

  • IJCAI 2024 A Conflict-Embedded Narrative Generation Using Commonsense Reasoning [paper]
    Youngrok Song, Gunhee Cho, Hyun-Jee Kim, Youngjune Kim, Byung-Chull Bae, Yun-Gyung Cheong

    Proposes a neuro-symbolic framework that embeds conflict in stories by using commonsense defeasible inference to weaken causal links toward protagonist goals.

  • NAACL 2024 Returning to the Start: Generating Narratives with Related Endpoints [paper] GitHub stars
    A. Brei, Chao Zhao, Snigdha Chaturvedi

    Proposes RENarGen, which first generates related opening and closing sentences then infills the middle, producing stories with stronger narrative closure.

  • EMNLP Findings 2023 Improving Pacing in Long-Form Story Planning [paper] GitHub stars
    Yichen Wang, Kevin Yang, Xiaoming Liu, Dan Klein

    Proposes CONCOCT, which trains a concreteness evaluator to guide vaguest-first outline expansion and filtering, yielding more consistent pacing in story outlines.

  • EMNLP Findings 2023 Affective and Dynamic Beam Search for Story Generation [paper]
    Tenghao Huang, Ehsan Qasemi, Bangzheng Li, He Wang, Faeze Brahman, Muhao Chen, Snigdha Chaturvedi

    Proposes a decoding method combining bandit-driven dynamic beam sizing and affect-intensity reranking to generate stories with more interesting twists.

  • EMNLP Findings 2023 GROVE: A Retrieval-augmented Complex Story Generation Framework with A Forest of Evidence [paper]
    Zhihua Wen, Zhi-Liang Tian, Wei Wu, Yu-Xin Yang, Yanqi Shi, Zhen Huang, Dongsheng Li

    Proposes GROVE, which retrieves human-written story examples and builds an asking-why forest of evidence to add complex, credible plot details.

  • EMNLP Findings 2023 Narrative Order Aware Story Generation via Bidirectional Pretraining Model with Optimal Transport Reward [paper]
    Zhicong Lu, Li Jin, Guang-Luan Xu, Linmei Hu, Nayu Liu, Xiaoyu Li, Xian Sun, Zequn Zhang, Kaiwen Wei

    Proposes a bidirectional pretrained event model with RL using an optimal transport reward to generate coherent stories with flashbacks.

๐ŸŽฏ Training

  • ArXiv 2026 Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion [paper]
    Hwan Chang, Yongil Kim, Heuiyeen Yeen, Yireun Kim, Jinsik Lee, Hwanhee Lee

    Proposes attribute-guided genre expansion to build a 50K, 13-genre creative writing corpus; fine-tuning on it beats training on story-centric writing data.

  • ArXiv 2026 Retell, Reward, Repeat: Reinforcement Learning for Narrative Theory-Informed Story Generation [paper]
    David Y. Liu, Xanthe Muston, A. Joshi, Sebastian Sequoiah-Grayson

    Shows that reinforcement learning from narrative-theory-informed AI feedback (d-RLAIF) yields more diverse, convention-aligned stories than supervised fine-tuning.

  • EMNLP 2026 POLARIS: Guiding Small Models to Write Long Stories [paper]
    Rishanth Rajendhran, Jenna Russell, Mohit Iyyer, J. Wieting

    Proposes a GRPO recipe with LLM-judge rewards and injected human reference stories, letting a 9B model write long stories beyond training length.

  • EMNLP 2026 Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction [paper]
    Ze-Han Li, Yu-Tong Zhu, Si-Yang Wu, Honglin Bao, James A. Evans

    Finds by comparing OLMo checkpoints that post-training compresses thematic, affective, and stylistic variation in fiction, most for professional literary text.

  • ArXiv 2026 StoryAlign: Evaluating and Training Reward Models for Story Generation [paper]
    Hao-Tian Xia, Hao Peng, Yunjia Qi, Xiaozhi Wang, Bin Xu, Lei Hou, Juan-Zi Li

    Introduces StoryRMB, a benchmark exposing weak reward models for story preferences, and StoryReward, trained on 100K preference pairs for best-of-n story selection.

  • ACL Findings 2026 UniCreative: Unifying Long-form Logic and Short-form Sparkle via Reference-Free Reinforcement Learning [paper]
    Xiao-Long Wei, Zerun Zhu, Simin Niu, Xingyu Zhang, Peiying Yu, Chang Xiao, Yu-Chen Li, Ji-Cheng Yang, Zhejun Zhao, Chong Meng, et al.

    Proposes a reference-free RL framework with an adaptive constraint-aware generative reward model and ACPO policy optimization, unifying long-form and short-form creative writing.

  • ACL 2026 DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative Writing [paper] GitHub stars
    Qian Cao, Yahui Liu, Wei Bi, Yi Zhao, Rui-jie Song, Xiting Wang, Rui-Ming Tang, Guo-Rui Zhou, Han Li

    Proposes an RL framework for creative writing that branches diverse plans in long chain-of-thought and adds a group-aware diversity reward.

  • ArXiv 2026 Rewarding Creativity: A Human-Aligned Generative Reward Model for Reinforcement Learning in Storytelling [paper]
    Zhaoyan Li, Hang Lei, Yuji Wang, Lan Liu, Hao Liu, Liang Yu

    Proposes RL for storytelling with a reasoning generative reward model aligned to human creativity judgments and entropy-based reward shaping for training stability.

  • ACL Findings 2026 From Style to Story: A Curriculum Learning Approach for Imitative Novel Generation [paper]
    Xueran Han, Yuhan Liu, Mingzhe Li, Wei Liu, Sen Hu, Rui Yan, Zhiqiang Xu, Xiuying Chen

    Introduces imitative novel generation and trains WriterAgent via curriculum learning with hierarchical LoRA modules to mimic an author's style, characters, and plots.

  • CoNLL 2026 Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning [paper] GitHub stars
    Jinlong Liu, Mohammed Bahja, Venelin Kovatchev, Mark Lee

    Trains an authorship-verification style judge and uses it as a GRPO reward to fine-tune an 8B model for writing like classic authors.

  • AAAI 2026 RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing [paper]
    Jian-Xing Liao, Tian Zhang, Xiao Feng, Yusong Zhang, Rui Yang, Hao-Rui Wang, Bosi Wen, Ziyi Wang, Run-Zhi Shi

    Proposes RL with a dynamically weighted mix of writing-quality and constraint-verification rewards in GRPO, improving both creative quality and instruction following.

  • ACL 2026 Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning [paper] GitHub stars
    Xuanyu Lei, Chen-Liang Li, Yuning Wu, Kai Liu, Weizhou Shen, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Yang Liu

    Proposes adaptive curriculum RL for long-form writing with margin-aware data selection, pairwise comparison rewards, and dynamic reference scheduling, beating SFT baselines.

  • ArXiv 2025 ๐ŸŒŸ Learning to Reason for Long-Form Story Generation [paper] GitHub stars
    Alexander Gurung, Mirella Lapata

    Proposes RL for story reasoning via Next-Chapter Prediction, rewarding plans that raise completion likelihood of real book chapters without labeled data.

  • ArXiv 2025 ๐ŸŒŸ Modifying Large Language Model Post-Training for Diverse Creative Writing [paper] GitHub stars
    John Joon Young Chung, Vishakh Padmakumar, Melissa Roemmele, Yuqian Sun, Max Kreminski

    Adds deviation from other same-prompt samples into DPO and ORPO objectives, increasing creative writing output diversity with minimal quality loss.

  • ArXiv 2025 LiteraryTaste: A Preference Dataset for Creative Writing Personalization [paper]
    John Joon Young Chung, Vishakh Padmakumar, Melissa Roemmele, Yi Wang, Yuqian Sun, Tiffany Wang, S. Almeda, Brett A. Halperin, Yuwen Lu, Max Kreminski

    Releases reading preferences from 60 people over creative text pairs, finding tastes diverge and stated preferences poorly predict revealed ones.

  • ArXiv 2025 COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes [paper] GitHub stars
    Yunwen Li, Shuangshuang Ying, Xingwei Qu, Xin Li, Sheng Jin, Minghao Liu, Zhoufutu Wen, Tianyu Zheng, Xeron Du, Qiguang Chen, et al.

    Releases 1,665 Chinese creative writing triplets with reverse-engineered prompts and reasoning traces, finding process supervision helps only when mixed with general data.

  • ArXiv 2024 ๐ŸŒŸ Weaver: Foundation Models for Creative Writing [paper]
    Tiannan Wang, Jiamin Chen, Qi Jia, Shuai Wang, Ruoyu Fang, Huilin Wang, Zhaowei Gao, Chunzhao Xie, Chuou Xu, Jihong Dai, et al.

    Introduces Weaver, a 1.8B-34B LLM family pre-trained and aligned for creative and professional writing, with a routing agent balancing quality and cost.

  • EMNLP 2024 MirrorStories: Reflecting Diversity through Personalized Narrative Generation with Large Language Models [paper]
    Sarfaroz Yunusov, Hamza Sidat, Ali Emami

    Introduces MirrorStories, 1,500 LLM-generated stories personalized to reader identity, and finds they engage readers more than generic human or LLM stories.

  • EMNLP Workshop 2024 PEARL: Personalizing Large Language Model Writing Assistants with Generation-Calibrated Retrievers [paper]
    Sheshera Mysore, Zhuoran Lu, Meng-Ting Wan, Long-Qi Yang, Steve Menezes, Tina Baghaee, E. Gonzalez, Jennifer Neville, Tara Safavi

    Proposes Pearl, a personalized writing assistant whose retriever is trained to be generation-calibrated, selecting user documents that most improve personalized LLM outputs.

๐Ÿ“ Evaluation

๐Ÿงช Benchmarks

  • EACL 2026 ๐ŸŒŸ LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing [paper]
    Daniel Fein, S. Russo, Violet Xiang, Kabir Jolly, Rafael Rafailov, Nick Haber

    Introduces a benchmark of 2,480 human-labeled story comparisons and 43,827 training pairs for creative writing evaluation, benchmarking LLM judges and reward models.

  • ACL Findings 2026 Lost in Stories: Consistency Bugs in Long Story Generation by LLMs [paper] GitHub stars
    Junjie Li, Xinru Guo, Yuhao Wu, Roy Ka-Wei Lee, Hong-Zhi Li, Yutao Xie

    Introduces a 2,000-prompt benchmark and automated checker for consistency errors in long story generation, analyzing where and which contradictions LLMs make.

  • ACL 2026 LitVISTA: A Benchmark for Narrative Orchestration in Literary Text [paper]
    Mingzhe Lu, Yiwen Wang, Yanbing Liu, Qi You, Chong Liu, Ruize Qin, Haoyu Dong, Wenyu Zhang, Jia-Rui Zhang, Yue Hu, et al.

    Proposes a framework and annotated literary benchmark for narrative orchestration, finding frontier LLMs fail to jointly capture narrative function and structure.

  • ACL Findings 2026 ChangJuan: A Comprehensive Benchmark for Book-Length Chinese Story Evaluation [paper] GitHub stars
    Dingyi Yang, Mingshuo Wang, Qin Jin

    Introduces a benchmark of 300 Chinese novels with human ratings and distilled reader viewpoints, plus CLEM, an 8B evaluator for book-length stories.

  • EACL 2026 Mary, the Cheeseburger-Eating Vegetarian: Do LLMs Recognize Incoherence in Narratives? [paper]
    Karin de Langis, Pรผren ร–ncel, Ryan Peters, Andrew Elfenbein, Laura K. Allen, Andreas Schramm, Dongyeop Kang

    Finds LLM internal representations detect incoherent narratives but their ratings do not, and models notice setting violations more than character-trait violations.

  • EACL Findings 2026 WebNovelBench: Placing LLM Novelists on the Web Novel Distribution [paper] GitHub stars
    Leon Lin, Jun Zheng, Haidong Wang

    Introduces a benchmark of 4,000+ Chinese web novels that scores LLM synopsis-to-story outputs on eight dimensions and ranks them against human-authored percentiles.

  • ACL 2025 What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation [paper] GitHub stars
    Dingyi Yang, Qin Jin

    Introduces LongStoryEval, 600 books averaging 121K tokens with reader reviews, compares long-story evaluation methods, and trains NovelCritique, an 8B summary-based evaluator.

  • ArXiv 2025 Finding Flawed Fictions: Evaluating Complex Reasoning in Language Models via Plot Hole Detection [paper]
    Kabir Ahuja, Melanie Sclar, Yulia Tsvetkov

    Introduces FlawedFictions, a benchmark built by synthesizing plot holes in human stories, to test LLM narrative reasoning via plot hole detection.

  • ArXiv 2025 LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm [paper] GitHub stars
    Siwei Wu, Yizhi Li, Xingwei Qu, R. Ravikumar, Yucheng Li, Tyler Loakman, Shanghaoran Quan, Xiao-Yong Wei, R. Batista-Navarro, Cheng-Hua Lin

    Introduces LongEval, a benchmark comparing direct and plan-based long-text generation, finding LLMs degrade with length while small long-text-trained models stay competitive.

  • ACL Findings 2025 Towards A "Novel" Benchmark: Evaluating Literary Fiction with Large Language Models [paper]
    Wenqing Wang, Mingqi Gao, Xinyu Hu, Xiaojun Wan

    Proposes a ten-metric macro/meso/micro evaluation framework and bilingual annotated fiction dataset, revealing a high-starting, low-ending pattern in LLM-written novels.

  • ICLR 2025 LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs [paper] GitHub stars
    Yuhao Wu, Ming Shan Hee, Zhiqing Hu, Roy Ka-Wei Lee

    Introduces a benchmark for instruction-following long-form generation at 16K and 32K tokens, finding all tested LLMs struggle as output length grows.

  • NAACL Findings 2025 CollabStory: Multi-LLM Collaborative Story Generation and Authorship Analysis [paper] GitHub stars
    Saranya Venkatraman, N. Tripto, Dongwon Lee

    Introduces CollabStory, a dataset of 32k stories co-written by up to five LLMs, with authorship analysis tasks and baselines for multi-LLM writing.

  • ArXiv 2024 CS4: Measuring the Creativity of Large Language Models Automatically by Controlling the Number of Story-Writing Constraints [paper] GitHub stars
    Anirudh Atmakuru, Jatin Nainani, Rohith Siddhartha Reddy Bheemreddy, Anirudh Lakkaraju, Zonghai Yao, Hamed Zamani, Haw-Shiuan Chang

    Introduces CS4, a benchmark that measures LLM story creativity by varying the number of prompt constraints to prevent retelling memorized stories.

  • ACL 2023 StoryWars: A Dataset and Instruction Tuning Baselines for Collaborative Story Understanding and Generation [paper]
    Yulun Du, Lydia B. Chilton

    Introduces StoryWars, 40k collaborative stories from 9,400 authors forming 101 understanding and generation tasks, with an instruction-tuned InstructStory baseline.

๐Ÿ“ Metrics

  • ArXiv 2026 When Reasoning Supervision Hurts: TTCW-Based Long-Form Literary Review Generation [paper] GitHub stars
    Jinlong Liu, Mohammed Bahja, Mark D. Lee

    Releases 263K stories with TTCW-based review annotations and finds fine-tuning without reasoning traces outperforms reasoning-supervised training for literary review generation.

  • ArXiv 2026 Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling [paper]
    Pei-Qi Sui, Yu-Tong Zhu, Tianyi Cheng, Peter West, R. So, Hoyt Long, Ari Holtzman

    Introduces 100-Endings, measuring narrative tension by how often repeated ending predictions fail as a story unfolds, and a pipeline that raises tension.

  • ACL 2026 EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation [paper]
    Xinda Wang, Zhengxu Hou, Yangshijie Zhang, Bingren Yan, Jia-Lin Liu, Chen-Zhuo Zhao, Zhi-Bo Yang, Bin-Bin Yang, Feng Xiao

    Trains a pairwise story evaluator on self-synthesized, multi-agent-filtered chain-of-thought data and uses it as a reward model to improve story generation.

  • ACL Findings 2025 The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing [paper] GitHub stars
    Guillermo Marco, Julio Gonzalo, Vรญctor Fresno-Fernรกndez

    Finds conflicting evaluations of AI fiction reflect reader differences, clustering 101 annotators into surface-focused and holistic reader profiles via textual feature preferences.

  • Applied Sciences 2025 Evaluating Creativity: Can LLMs Be Good Evaluators in Creative Writing Tasks? [paper]
    Sungeun Kim, Dongsuk Oh

    Finds that LLMs rate creative texts more consistently than humans but miss nuanced, culturally specific, and context-dependent aspects of creativity.

  • TACL 2024 ๐ŸŒŸ Do Language Models Enjoy Their Own Stories? Prompting Large Language Models for Automatic Story Evaluation [paper]
    Cyril Chhun, Fabian M. Suchanek, Chloรฉ Clavel

    Studies LLMs as automatic story evaluators, finding they beat existing metrics at system-level correlation with humans but struggle to explain their ratings.

  • EMNLP Findings 2024 CHIRON: Rich Character Representations in Long-Form Narratives [paper]
    Alexander Gurung, Mirella Lapata

    Proposes character sheet representations built by LLM question-answering and entailment-based fact validation, improving masked-character prediction and measuring character-centricity.

  • EMNLP 2024 Learning Personalized Alignment for Evaluating Open-ended Text Generation [paper]
    Danqing Wang, Kevin Yang, Hanlin Zhu, Xiaomeng Yang, Andrew Cohen, Lei Li, Yuandong Tian

    Proposes PerSE, a LLaMA-2 based evaluator that infers reader preferences from in-context profiles to give personalized, interpretable scores for open-ended generation.

  • ACL 2023 ๐ŸŒŸ Can Large Language Models Be an Alternative to Human Evaluations? [paper]
    Cheng-Han Chiang, Hung-yi Lee

    Finds that LLMs given the same instructions as human annotators produce story and adversarial-text ratings consistent with expert human evaluation.

  • EMNLP Findings 2023 DeltaScore: Evaluating Story Generation with Differentiating Perturbations [paper]
    Zhuohan Xie, Miao Li, Trevor Cohn, Jey Han Lau

    Proposes DeltaScore, which evaluates story aspects like fluency and interestingness by measuring likelihood changes under aspect-specific perturbations.

๐Ÿ” Analyses

  • ACL 2026 CASPER in the Machine: Insights into Character Variety in LLM-Generated Stories [paper]
    A. Brei, Abhisheik Sharma, Nicholas Sanaie, Lu Wang, Snigdha Chaturvedi

    Compares characters in LLM-generated and human-written stories along eight narratological dimensions, examining similarity and variety of character types.

  • ACL Findings 2026 BIASEDTALES-ML: A Multilingual Dataset for Analyzing Narrative Attribute Distributions in LLM-Generated Stories [paper]
    Yuxuan Ouyang, Yi Luo, Jing-Bo Zhu, Tong Xiao

    Releases a 350K-story multilingual parallel corpus of LLM children's stories, finding narrative attribute distributions vary substantially across eight languages.

  • ArXiv 2026 StoryScope: Investigating idiosyncrasies in AI fiction [paper]
    Jenna Russell, Rishanth Rajendhran, Chau Minh Pham, Mohit Iyyer, J. Wieting

    Finds that discourse-level narrative features alone separate human from AI fiction and attribute AI stories to specific models, independent of stylistic cues.

  • ArXiv 2026 LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers [paper]
    Pei-Qi Sui

    Finds across 28 LLMs that model story continuations carry much lower information-theoretic uncertainty than human writing, worsened by instruction tuning.

  • PNAS 2025 ๐ŸŒŸ Echoes in AI: Quantifying Lack of Plot Diversity in LLM Outputs [paper]
    Wei-Jia Xu, Nebojsa Jojic, Sudha Rao, C. Brockett, Bill Dolan

    Finds LLM stories from one prompt reuse plot element combinations far more than human stories, and proposes an automatic narrative-level diversity metric.

  • EMNLP 2025 Biased Tales: Cultural and Topic Bias in Generating Children's Stories [paper]
    Donya Rooein, Vilรฉm Zouhar, Debora Nozza, Dirk Hovy

    Introduces a dataset exposing gender and cultural stereotypes in LLM children's stories, such as girls receiving more appearance-related attributes than boys.

  • Information 2025 AI Narrative Modeling: How Machines' Intelligence Reproduces Archetypal Storytelling [paper]
    I. Kabashkin, Olga Zervina, Boriss Misnevs

    Finds LLMs reproduce structured Jungian archetypes like the Hero well but struggle with ambiguous ones like the Shadow and Trickster.

  • ICCC 2025 Evaluating Creative Short Story Generation in Humans and Large Language Models [paper] GitHub stars
    Mete Ismayilzada, Claire E. Stevenson, Lonneke van der Plas

    Compares stories by 60 LLMs and 60 humans, finding LLMs lag in novelty and surprise though non-experts rate LLM stories more creative.

  • COLING 2025 Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs [paper] GitHub stars
    Guillermo Marco, Luz Rello, Julio Gonzalo

    Finds a fine-tuned BART-large outscores average human writers on short fiction in human ratings, contrasting its linguistic traits with GPT-3.5 and GPT-4o.

  • EMNLP 2024 ๐ŸŒŸ Are Large Language Models Capable of Generating Human-Level Narratives? [paper]
    Yufei Tian, Tenghao Huang, Miri Liu, Derek Jiang, Alexander Spangher, Muhao Chen, Jonathan May, Nan-Yun Peng

    Analyzes story arcs, turning points, and affect, finding LLM stories are homogeneously positive and lack tension compared with suspenseful, diverse human narratives.

  • CHI 2024 ๐ŸŒŸ Art or Artifice? Large Language Models and the False Promise of Creativity [paper]
    Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, S. Muresan, Chien-Sheng Wu

    Proposes the Torrance Test of Creative Writing, finding LLM stories pass 3-10X fewer expert tests than professional stories, and LLM judges misalign.

  • EMNLP 2024 Pron vs Prompt: Can Large Language Models already Challenge a World-Class Fiction Author at Creative Text Writing? [paper]
    Guillermo Marco, Julio Gonzalo, M.Teresa Mateo-Girona, Ramรณn Santos

    Stages a contest between novelist Patricio Pron and GPT-4, where expert critics judge the LLM far from a top human fiction author.

  • EMNLP 2024 Measuring Psychological Depth in Language Models [paper]
    Fabrice Y. Harel-Canada, Hanyu Zhou, Sreya Muppalla, Zeynep Yildiz, Miryung Kim, Amit Sahai, Nan-Yun Peng

    Introduces the Psychological Depth Scale for stories' emotional and empathic impact, automates it with LLM personas, finding GPT-4 rivals top Reddit stories.

  • HSSC 2023 Experimental narratives: A comparison of human crowdsourced storytelling and AI storytelling [paper]
    Nina Beguลก

    Compares crowdworker and GPT-3.5/GPT-4 stories on identical Pygmalion prompts, finding AI stories more progressive on gender yet less imaginative.

  • EMNLP Findings 2023 A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing [paper]
    Carlos G'omez-Rodr'iguez, Paul Williams

    Compares LLMs and humans on an unusual comic epic prompt, finding top commercial LLMs match humans on most criteria except creativity.

  • C&C 2023 More human than human: LLM-generated narratives outperform human-LLM interleaved narratives [paper]
    Z. Zhao, Sophie Song, Bridget Duah, J. Macbeth, Scott A. Carter, Monica P. Van, N. Bravo, M. Klenk, Kate Sick, Alexandre L. S. Filipowicz

    Finds through two roughly 500-participant studies that readers prefer purely LLM-generated stories over human-LLM interleaved stories.

  • INLG 2023 The Next Chapter: A Study of Large Language Models in Storytelling [paper]
    Zhuohan Xie, Trevor Cohn, Jey Han Lau

    Finds that prompted LLMs write stories rivaling human authors and beating prior generators, though they sometimes replicate real stories.

๐Ÿค Co-creation

๐Ÿ› ๏ธ Tools

  • EMNLP Findings 2026 Generating Constructive Feedback on Stories via Reinforcement Learning [paper]
    Maja Stahl, Timon Ziegenbein, Henning Wachsmuth

    Trains LLMs with GRPO and a multi-component constructiveness reward to give story-specific feedback, finding actionable suggestions drive constructiveness most.

  • ArXiv 2026 GraphStory: Collaborative Story Writing through Event-Based Narrative Editing [paper]
    X. Lรช, Minh-Loi Nguyen, Khanh-Duy Le, Minh-Triet Tran, Trung-Nghia Le

    Builds a graph-based writing assistant for organizing plot points and exploring alternative branches, which writers found reduced structuring effort.

  • ArXiv 2026 Fabula: Building a Narrative Storytelling Sidekick with the Writers' Community [paper]
    Piotr Mirowski, Benjamin D. Wedin, Reinald Kim Amplayo, Rich Galt, Duncan Williams, Rida Qadri, Jaume Sanchez-Elias, Erin Drake-Kajioka, Sian Gooding, Lucรญa Lรณpez-Rivilla, et al.

    Designs and evaluates a narratology-based fiction writing app with 42 writers, probing auto-evaluators, plan-exposing interfaces, and cultural fit of story structures.

  • CHI 2026 Exploring Creator-Centric Methods for LLM-Assisted Interactive Storytelling [paper]
    Yue-Lu Li, Siyi Wu, Lu-Jin Zhang, Zhihan Guo, Wenchuan Lu, David Yip

    Designs CoNoder, a creator-centered LLM prototype for interactive narratives with node-graph editing, ripple-effect analysis, and simulated reader feedback, informed by creator interviews.

  • CHI 2026 Orchid-Creator: An Authoring Tool Supporting LLM-Driven Interactive Narrative Creation [paper]
    Zhen Wu, Serkan Kumyol, Zhengyang Ma, T. Braud

    Builds an LLM authoring tool representing interactive narratives as card-based story graphs, found easier for structuring than Twine and AI Dungeon.

  • CHI 2026 Plotania: Exploring Transparency Trade-offs in AI Co-Writing Through Virtual Readers and Transparent Attribution [paper]
    Yu-Feng Hu, Jinyi Zhang, Ze-Hua Wang, Chun Yu

    Builds a co-writing system with virtual reader reactions and AI attribution, finding transparency raises awareness but lowers creative agency and AI usage.

  • CHI 2026 Narrix: Remixing Narrative Strategies from Examples for Story Writing [paper]
    Chao Zhang, Shunan Guo, Abe Davis, Eunyee Koh

    Builds a writing tool that highlights narrative strategies in example stories and lets novices apply them to drafts via strategy-steered generation.

  • CHI 2026 NarrativeLoom: Enhancing Creative Storytelling through Multi-Persona Collaborative Improvisation [paper]
    Yuxi Ma, Yongqian Peng, Fengyuan Yang, Siyu Zha, Chi Zhang, Zi-Xia Jia, Zilong Zheng, Yixin Zhu

    Builds a multi-persona co-creative storytelling system based on blind variation and selective retention; experts rated co-authored stories more creative.

  • CHI 2026 PlayWrite: A Multimodal System for AI Supported Narrative Co-Authoring Through Play in XR [paper]
    Esen K. Tutuncu, Qian Zhou, Frederik Brudy, George W. Fitzmaurice, Fraser Anderson

    Builds a mixed-reality system where users author stories by manipulating virtual characters and props, which multi-agent AI turns into rearrangeable narrative beats.

  • CHI 2026 DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection [paper]
    Yuying Tang, Xinyi Chen, Haotian Li, Xing Xie, Xiaojuan Ma, Huamin Qu

    Builds a screenplay refinement system whose AI agent first simulates character experience, then evaluates it to give feedback that deepens screenwriters' reflection.

  • CHI 2026 Vidmento: Creating Video Stories through Context-Aware Expansion with Generative Video [paper]
    Catherine Yeh, Anh Truong, Mira Dontcheva, Bryan Wang

    Builds a tool that fills narrative gaps in video stories by generating context-aware clips that blend stylistically and narratively with captured footage.

  • CHI 2026 DiaryPlay: AI-Assisted Creation of Interactive Story Vignettes for Everyday Storytelling [paper]
    Jiangnan Xu, Haeseul Cha, Gosu Choi, Gyu-cheol Lee, Y. Yoon, Zucheul Lee, Konstantinos Papangelis, D. Kim, Juho Kim

    Builds an AI authoring system turning text stories into interactive vignettes, using LLM-controlled divergence to keep NPC behavior within the intended story.

  • ArXiv 2025 Elsewise: Authoring Open-ended Interactive Narrative with Possibility Space Visualization [paper]
    Yi Wang, John Joon Young Chung, Melissa Roemmele, Yuqian Sun, Tiffany Wang, S. Almeda, Brett A. Halperin, Yuwen Lu, Max Kreminski

    Builds an authoring tool that visualizes bundled storylines of LLM-driven interactive narratives, helping authors anticipate player-experienced stories in a 12-user study.

  • CHI 2025 Toward Personalizable AI Node Graph Creative Writing Support: Insights on Preferences for Generative AI Features and Information Presentation Across Story Writing Processes [paper]
    Hua-Xuan Qin, Guangzhi Zhu, Mingming Fan, Pan Hui

    Studies a FigJam plugin combining node-graph story structure, LLM audience impersonation, and image/audio generation for personalized story writing and moral reflection.

  • CHI 2025 WhatELSE: Shaping Narrative Spaces at Configurable Level of Abstraction for AI-bridged Interactive Storytelling [paper]
    Zhuoran Lu, Qian Zhou, Yi Wang

    Builds an authoring system deriving narrative possibility spaces from example stories, letting authors bound them and unfold them into game events.

  • CHI 2025 Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols [paper]
    John Joon Young Chung, Melissa Roemmele, Max Kreminski

    Builds a storytelling system where users steer LLM story text by moving character symbols like toys, via a shared motion-text semantic space.

  • CHI 2024 ๐ŸŒŸ CharacterMeet: Supporting Creative Writers' Entire Story Character Construction Processes Through Conversation with LLM-Powered Chatbot Avatars [paper]
    Hua-Xuan Qin, Shan Jin, Ze Gao, Mingming Fan, Pan Hui

    Builds a system letting writers develop characters by conversing with customizable chatbot avatars; a 14-writer study shows it supports iterative character construction.

  • Frontiers in Robotics and AI 2024 Fostering childrenโ€™s creativity through LLM-driven storytelling with a social robot [paper]
    Maha Elgarf, Hanan Salam, Christopher Peters

    Fine-tunes an LLM to drive creative or non-creative social robot storytelling, finding the creative robot boosts children's fluency, flexibility, and elaboration.

  • C&C 2024 Ai.llude: Encouraging Rewriting AI-Generated Text to Support Creative Expression [paper]
    David Zhou, S. Sterman

    Finds from 27 writing sessions that deliberately imperfect intermediate AI suggestions encourage writers to rewrite, supporting creative ownership and reflection.

  • ArXiv 2024 GhostWriter: Augmenting Collaborative Human-AI Writing Experiences Through Personalization and Agency [paper]
    Catherine Yeh, Gonzalo Ramos, Rachel Ng, Andy Huntington, R. Banks

    Builds GhostWriter, a writing design probe that implicitly learns user style while offering explicit controls, studying how it supports agency and personalization.

  • TiiS 2023 ๐ŸŒŸ ID.8: Co-Creating Visual Stories with Generative AI [paper] GitHub stars
    Victor Antony, Chien-Ming Huang

    Introduces ID.8, an open-source system for co-creating visual stories with generative AI, with a user study highlighting enjoyment and remaining gaps.

  • EACL 2023 Fiction-Writing Mode: An Effective Control for Human-Machine Collaborative Writing [paper]
    Wenjie Zhong, Jason Naradowsky, Hiroya Takamura, Ichiro Kobayashi, Yusuke Miyao

    Annotates narrative paragraphs with writing-mode labels and fine-tunes LLMs conditioned on these modes, finding authors prefer mode-controlled suggestions in collaborative fiction writing.

๐Ÿ‘ฅ User Studies

  • COLM 2026 The Garden of Forking Prompts: How Users Explore Narrative Space in Story Generation [paper]
    Advait Deshmukh, N. Benedict, Melanie Walsh, Maria Antoniak

    Analyzes how users iteratively revise story prompts in wild chatbot logs, releasing WildStories and WildEdits and an edit-type framework for benchmarking.

  • ArXiv 2026 AI Fiction in the Wild [paper]
    Neel Gupta, Maria Antoniak, Melanie Walsh

    Analyzes 500,000 ChatGPT conversations, finding over a third involve fiction generation, dominated by power users favoring fanfiction, erotica, and repetition.

  • CHI 2026 Proactive AI as a Catalyst for Creativity? Balancing Human Agency and AI Contribution in Collaborative Story Writing [paper]
    Yiwen Yin, Ming-Ze Wu, R. Huang, X. Tong, Jun Zhou, Chun Yu, Yuanchun Shi

    Wizard-of-Oz study of intrusive versus non-intrusive proactive AI suggestions in story outlining, revealing a creativity-agency trade-off moderated by how inspiring suggestions are.

  • ACL 2025 Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback [paper]
    Hannah Rashkin, Elizabeth Clark, Fantine Huot, Mirella Lapata

    Introduces a task and 1,300 deliberately corrupted stories to evaluate LLM writing feedback, finding models often miss the biggest writing issue.

  • Interacciรณn 2025 Once More with (the Right) Feeling: How Historical Fiction Writing Processes of Character Design, Plot Outline, and Context Checking Are Affected by Co-Writing with ChatGPT [paper]
    Yun Chen, Yiwei Wang, Antoni B. Chan, Jixing Li, LC Ray

    Examines how co-writing with ChatGPT affects historical fiction writers' character design, plot outlining, and context checking processes.

  • CHI 2025 Understanding Screenwriters' Practices, Attitudes, and Future Expectations in Human-AI Co-Creation [paper]
    Yuying Tang, Haotian Li, Minghe Lan, Xiao-Juan Ma, Huamin Qu

    Interviews 23 screenwriters on how they integrate AI across workflow stages and categorizes expected AI roles as actor, audience, expert, and executor.

  • CHI 2024 ๐ŸŒŸ Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models [paper]
    Paramveer S. Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub, Shaochun Zheng, L. P. Robert

    Finds with 131 participants a U-shaped effect of AI scaffolding: paragraph-level suggestions improve writing quality and productivity, while sentence-level ones do not.

  • CSCW 2024 'It was 80% me, 20% AI': Seeking Authenticity in Co-Writing with Large Language Models [paper]
    Angel Hsing-Chi Hwang, Q. Liao, Su Lin Blodgett, Alexandra Olteanu, Adam Trischler

    Interviews 19 professional writers and surveys readers on authenticity in AI co-writing, finding personalization should support writer growth beyond text production.

  • CHI 2024 The Value, Benefits, and Concerns of Generative AI-Powered Assistance in Writing [paper]
    Zhuo-Yan Li, Chen Liang, Jing Peng, Ming Yin

    Finds through an experiment that people will forgo payment for AI writing help, which boosts productivity and confidence but raises ownership concerns.

  • CHI 2023 ๐ŸŒŸ Social Dynamics of AI Support in Creative Writing [paper]
    Katy Ilonka Gero, Tao Long, Lydia B. Chilton

    Interviews 20 creative writers to identify what help they want, how they perceive supporters, and values shaping AI-versus-human support choices.

  • ArXiv 2023 Creativity Support in the Age of Large Language Models: An Empirical Study Involving Emerging Writers [paper]
    Tuhin Chakrabarty, Vishakh Padmakumar, Faeze Brahman, S. Muresan

    Studies 30 writers using an LLM interface based on the cognitive process model, finding LLMs most helpful for translating and reviewing.

  • IDC 2023 Design implications of generative AI systems for visual storytelling for young learners [paper]
    Ariel Han, Zhenyao Cai

    Elicits parent, teacher, and researcher views on generative AI for children's visual storytelling and proposes AIStory, a prototype app supporting literacy.

๐Ÿ“š Surveys

  • ArXiv 2026 Narrative Theory-Driven LLM Methods for Automatic Story Generation and Understanding: A Survey [paper]
    David Y. Liu, A. Joshi, Paul Dawson

    Surveys LLM story generation and understanding through narratology, finding generation lags understanding and recommending theory-based metrics over a single quality benchmark.

  • EMNLP Findings 2025 ๐ŸŒŸ A Survey on LLMs for Story Generation [paper]
    Maria Teleki, Vedangi Bengali, Xiangjue Dong, Sai Janjur, Haoran Liu, Tian Liu, Cong Wang, Ting-Yiu Liu, Yin Zhang, Frank Shipman, et al.

    Surveys LLM story generation, organizing work into autonomous generation versus author assistance and comparing methods, datasets, story types, and evaluations.

  • ArXiv 2024 ๐ŸŒŸ What Makes a Good Story and How Can We Measure It? A Comprehensive Survey of Story Evaluation [paper]
    Dingyi Yang, Qin Jin

    Surveys story evaluation across text-to-text, visual-to-text, and text-to-visual tasks, proposing a taxonomy of human criteria, benchmarks, and automatic metrics.

  • Neurocomputing 2023 Open-world Story Generation with Structured Knowledge Enhancement: A Comprehensive Survey [paper]
    Yuxin Wang, Jieru Lin, Zhiwei Yu, Wei Hu, Bรถrje F. Karlsson

    Surveys structured knowledge-enhanced story generation, offering a taxonomy of how knowledge is injected to improve coherence and grounding, plus future directions.

๐Ÿงฐ Public Resources

๐Ÿ“ฆ Datasets

  • WritingPrompts: about 300K human-written stories paired with Reddit writing prompts; the most widely used dataset for open-ended story generation.
  • VIST: photo sequences paired with human-written stories; the standard dataset for Visual2Story.
  • PG-19: full-length books from Project Gutenberg, a common source of long-form fiction.

๐Ÿ† Leaderboards

๐Ÿ›๏ธ Venues & Workshops

  • ICIDS: International Conference on Interactive Digital Storytelling.
  • AIIDE: AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment.
  • WNU: Workshop on Narrative Understanding, co-located with ACL conferences.
  • In2Writing: Workshop on Intelligent and Interactive Writing Assistants.
  • Wordplay: When Language Meets Games workshop.

๐Ÿ”— Related Lists

๐Ÿค Contributing

We welcome paper recommendations and corrections. Please read CONTRIBUTING.md for what the list includes and how to format an entry, then open a paper recommendation issue or a pull request.

๐Ÿ“ Citation

If you find this list useful, please consider citing it:

@misc{ma2023awesomestorygeneration,
  title        = {Awesome-Story-Generation: A Curated List of Papers on Story Generation in the Era of Large Language Models},
  author       = {Ma, Yingpeng and Ma, Yan},
  year         = {2023},
  howpublished = {\url{https://github.com/yingpengma/Awesome-Story-Generation}}
}

โญ Star History

Star History Chart

language-generation
large-language-model
narrative
natural-language-processing
nlp
novel
open-ended
story
story-generation
storytelling

Significant stargazers

Fredrik Norรฉn

309 followers ยท starred Oct 2024

flashmouse

23 followers ยท starred Apr 2026

Languages

Python

100.0%