emanuelevivoli/awesome-comics-understanding

The official repo of the Comics Survey: "A missing piece in Vision and Language: A Survey on Comics Understanding"

141

16 commits

updated Jan 2, 2025

See the code

README

Awesome Comics Understanding Awesome

This repository contains a curated list of research papers and resources focusing on Comics Understanding.

arXiv

🔥 One missing piece in Vision and Language: A Survey on Comics Understanding 🔥

Authors: Emanuele Vivoli, Andrey Barsky, Mohamed Ali Souibgui, Artemis Llabrés, Marco Bertini, Dimosthenis Karatzas

📣 Latest News 📣

  • 🚧 This repo is a work in progress, please contribute here
  • 14 September 2024 Our survey paper have dropped in arXiv !!

📚 Table of Contents

Overview of Vision-Language Tasks of the Layers of Comics Understanding. The ranking is based on input and output modalities and dimensions, as illustrated in the paper.

Layers of Comics Understanding

Every survey worthy of the name includes illustrative visuals to enhance understanding. We've followed this approach by providing examples for each task in the Layer of Comics Understanding.
Go check every Layer's tasks image ⬇️.

Layer 1: Tagging and Augmentation

  • Tagging

    • Image classification
      YearConference / JournalTitleAuthorsLinks
      2023TIPPanel-Page-Aware Comic Genre UnderstandingXu, Chenshu et al.📜 Paper
      2019ICDAR WorkshopAnalysis Based on Distributed Representations of Various Parts Images in Four-Scene Comics Story DatasetTerauchi, Akira et al.📜 Paper
      2018TPAMILearning Consensus Representation for Weak Style ClassificationJiang, Shuhui et al.📜 Paper
      2018ICDARComic Story Analysis Based on Genre ClassificationDaiku, Yuki et al.📜 Paper
      2017ICDARHistogram of Exclamation Marks and Its Application for Comics AnalysisHiroe, Sotaro et al.📜 Paper
      2014ACM MultimediaLine-Based Drawing Style Description for Manga ClassificationChu, Wei-Ta et al.📜 Paper
    • Emotion classification
      YearConference / JournalTitleAuthorsLinks
      2023MMMManga Text Detection with Manga-Specific Data Augmentation and Its Applications on Emotion AnalysisYang, Yi-Ting et al.📜 Paper
      2021ICDARCompetition on Multimodal Emotion Recognition on Comics ScenesNguyen, Nhu-Van et al.📜 Paper, 👨‍💻 Code
      2016MANPU (ICPR)Manga Content Analysis Using Physiological SignalsSanches, Charles Lima et al.📜 Paper
      2015IIAI-AAIRelation Analysis between Speech Balloon Shapes and Their Serif Descriptions in ComicTanaka, Hideki et al.📜 Paper
    • Action Detection
      YearConference / JournalTitleAuthorsLinks
      2024ArxivMangaUB: A Manga Understanding Benchmark for Large Multimodal ModelsIkuta, Hikaru et al.📜 Paper
      2024MANPU (ICDAR)ComicBERT: A Transformer Model and Pre-training Strategy for Contextual Understanding in ComicsSoykan, Gurkan et al.📜 Paper, 👨‍💻 Code
      2024ICDARMultimodal Transformer for Comics Text-ClozeVivoli, Emanuele et al.📜 Paper
      2020ArxivA Comprehensive Study of Deep Video Action RecognitionZhu, Yi et al.📜 Paper, 👨‍💻 Code
      2017CVPRThe Amazing Mysteries of the Gutter: Drawing Inferences Between Panels in Comic Book NarrativesIyyer, Mohit et al.📜 Paper
    • Page Stream Segmentation
      YearConference / JournalTitleAuthorsLinks
      2022ICPRSemantic Parsing of Interpage RelationsDemirtaş, Mehmet Arif et al.📜 Paper
      2018LRECPage Stream Segmentation with Convolutional Neural Nets Combining Textual and Visual FeaturesWiedemann, Gregor et al.📜 Paper
      2013ICDARDocument Classification and Page Stream Segmentation for Digital Mailroom ApplicationsGordo, Albert et al.📜 Paper
  • Augmentation ( Image-2-Image )

    • Image Super-Resolution
      YearConference / JournalTitleAuthorsLinks
      2023MTAAutomatic Dewarping of Camera-Captured Comic Document ImagesGarai, Arpan et al.📜 Paper
    • Style Transfer
      YearConference / JournalTitleAuthorsLinks
      2023ArxivInkn'hue: Enhancing Manga Colorization from Multiple Priors with Alignment Multi-Encoder VAEJiramahapokee, Tawin📜 Paper, 👨‍💻 Code
      2023IEEE AccessRobust Manga Page Colorization via Coloring Latent SpaceGolyadkin, Maksim et al.📜 Paper
      2023TVCGShading-Guided Manga Screening from ReferenceWu, Huisi et al.📜 Paper
      2022ArxivDASS-Detector: Domain-Adaptive Self-Supervised Pre-Training for Face & Body Detection in DrawingsTopal, Barış Batuhan et al.📜 Paper, 👨‍💻 Code
      2021CVPRGenerating Manga from Illustrations via Mimicking Manga Creation WorkflowZhang, LM et al.*📜 Paper, 👨‍💻 Code
      2021CVPRUnbiased Mean Teacher for Cross-domain Object DetectionDeng, Jinhong et al.📜 Paper, 👨‍💻 Code
      2021CVPREncoding in Style: A StyleGAN Encoder for Image-to-Image TranslationRichardson, Elad et al.📜 Paper, 👨‍💻 Code
      2021AAAIMangaGAN: Unpaired Photo-to-Manga Translation Based on The Methodology of Manga DrawingSu, Hao et al.📜 Paper
      2019ISMSynthesis of Screentone Patterns of Manga CharactersTsubota, K. et al.📜 Paper, 👨‍💻 Code
      2018SciVisColor Interpolation for Non-Euclidean Color SpacesZeyen, Max et al.📜 Paper
      2017ACM-SIGGRAPH AsiaComicolorization: Semi-automatic Manga ColorizationFurusawa, Chie et al.📜 Paper, 👨‍💻 Code
      2017ICDARCGAN-Based Manga Colorization Using a Single Training ImageHensman, Paulina et al.📜 Paper, 👨‍💻 Code
      2017CVPRImage-to-Image Translation with Conditional Adversarial NetworksIsola, Phillip et al.📜 Paper
      2017ACM-TGDeep Extraction of Manga Structural LinesLi, Chengze et al.📜 Paper, 👨‍💻 Code
    • Vectorization
      YearConference / JournalTitleAuthorsLinks
      2023TCSVTMARVEL: Raster Gray-level Manga Vectorization via Primitive-wise Deep Reinforcement LearningH. Su et al.📜 Paper, 👨‍💻 Code
      2022CVPRTowards Layer-wise Image VectorizationMa, Xu et al.📜 Paper, 👨‍💻 Code
      2017ACM-TGDeep Extraction of Manga Structural LinesLi, Chengze et al.📜 Paper, 👨‍💻 Code
      2017TVCGManga Vectorization and Manipulation with Procedural Simple ScreentoneYao, Chih-Yuan et al.📜 Paper
      2011ACM-SIGGRAPHDepixelizing Pixel ArtKopf, Johannes et al.📜 Paper
      2003N/APotrace : A Polygon-Based Tracing AlgorithmSelinger, Peter📜 Paper, 👨‍💻 Code
    • Depth Estimation
      YearConference / JournalTitleAuthorsLinks
      2023CVPR WorkshopDense Multitask Learning to Reconfigure ComicsBhattacharjee, Deblina et al.📜 Paper
      2022WACVEstimating Image Depth in the Comics DomainBhattacharjee, Deblina et al.📜 Paper, 👨‍💻 Code
      2022CVPRMulT: An End-to-End Multitask Learning TransformerBhattacharjee, Deblina et al.📜 Paper, 👨‍💻 Code

Layer 2: Grounding, Analysis and Segmentation

  • Grounding

    • Object detection
      YearConference / JournalTitleAuthorsLinks
      2024MANPU (ICDAR)Comics Datasets Framework: Mix of Comics datasets for detection benchmarkingVivoli, Emanuele et al.📜 Paper
      2024MANPU (ICDAR)A Comprehensive Gold Standard and Benchmark for Comics Text Detection and RecognitionGurkan Soykan et al.📜 Paper, 👨‍💻 Code
      2023MMMManga Text Detection with Manga-Specific Data Augmentation and Its Applications on Emotion AnalysisYang, \relax YT et al.📜 Paper
      2023CSNTCPD: Faster RCNN-based DragonBall Comic Panel DetectionSharma, Rishabh et al.📜 Paper
      2022IJDARBCBId: First Bangla Comic Dataset and Its ApplicationsDutta, Arpita et al.📜 Paper
      2022ECCVCOO/ Comic Onomatopoeia Dataset for Recognizing Arbitrary or Truncated TextsBaek, Jeonghun et al.📜 Paper, 👨‍💻 Code
      2019ICDAR WorkshopWhat Do We Expect from Comic Panel Extraction?Nguyen Nhu, Van et al.📜 Paper
      2019ICDAR WorkshopCNN Based Extraction of Panels/Characters from Bengali Comic Book Page ImagesDutta, Arpita et al.📜 Paper
      2018VCIPText Detection in Manga by Deep Region Proposal, Classification, and RegressionChu, Wei-Ta et al.📜 Paper
      2018IWAITA Study on Object Detection Method from Manga Images Using CNNYanagisawa, Hideaki et al.📜 Paper
      2018IWAITA Study on Object Detection Method from Manga Images Using CNNYanagisawa, Hideaki et al.📜 Paper
      2017ICDARA Faster R-CNN Based Method for Comic Characters Face DetectionQin, Xiaoran et al.📜 Paper
      2016IJCGText-Aware Balloon Extraction from MangaLiu, Xueting et al.📜 Paper
      2016IJCNNLine-Wise Text Identification in Comic Books: A Support Vector Machine-Based ApproachPal, Srikanta et al.📜 Paper
      2016ICIPText Detection in Manga by Combining Connected-Component-Based and Region-Based ClassificationsAramaki, Yuji et al.📜 Paper
      2015ICIAPPanel Tracking for the Extraction and the Classification of Speech BalloonsJomaa, Hadi S. et al.📜 Paper
      2012DASPanel and Speech Balloon Extraction from Comic BooksHo, Anh Khoi Ngo et al.📜 Paper
      2011IJIMethod for Real Time Text Extraction of Digital Manga ComicArai, Kohei et al.📜 Paper
      2011ICDARRecognizing Text Elements for SVG Comic Compression and Its Novel ApplicationsSu, Chung-Yuan et al.📜 Paper
      2010ICITMethod for Automatic E-Comic Scene Frame Extraction for Reading Comic on Mobile DevicesArai, Kohei et al.📜 Paper
      2009IJHCIEnhancing the Accessibility for All of Digital Comic BooksPonsard, Christophe📜 Paper
    • Character Re-Identification
      YearConference / JournalTitleAuthorsLinks
      2024ArxivTails Tell Tales: Chapter-Wide Manga Transcriptions with Character NamesRagav Sachdeva et al.📜 Paper, 👨‍💻 Code
      2024NeurIPSCoMix: A Comprehensive Benchmark for Multi-Task Comic UnderstandingEmanuele Vivoli et al.📜 Paper, 👨‍💻 Code
      2024CVPRThe Manga Whisperer: Automatically Generating Transcriptions for ComicsSachdeva, Ragav et al.📜 Paper, 👨‍💻 Code
      2023IET Image ProcessingToward Cross-Domain Object Detection in Artwork Images Using Improved YoloV5 and XGBoostingAhmad, Tasweer et al.📜 Paper
      2023ArxivIdentity-Aware Semi-Supervised Learning for Comic Character Re-IdentificationSoykan, Gürkan et al.📜 Paper
      2023ACM-MM AsiaOcclusion-Aware Manga Character Re-Identification with Self-Paced Contrastive LearningZhang, Ci-Yin et al.📜 Paper
      2022ArxivUnsupervised Manga Character Re-Identification via Face-Body and Spatial-Temporal Associated ClusteringZhang, Z et al.📜 Paper
      2022ICIRCAST: Character Labeling in Animation Using Self‐supervision by TrackingNir, Oron et al.📜 Paper
      2020ICPRDual Loss for Manga Character Recognition with Imbalanced Training DataLi, Yonggang et al.📜 Paper
      2020ICMLA Simple Framework for Contrastive Learning of Visual RepresentationsChen, Ting et al.📜 Paper
      2015ACPRSimilarity Learning Based on Pool-Based Active Learning for Manga Character RetrievalIwata, Motoi et al.📜 Paper
      2014DASA Study to Achieve Manga Character Retrieval Method for Manga ImagesIwata, M. et al.📜 Paper
      2012CVPRColor Attributes for Object DetectionKhan, Fahad Shahbaz et al.📜 Paper
      2012ECCVPHOG Analysis of Self-Similarity in Aesthetic ImagesRedies, Christoph et al.📜 Paper
      2011ICDARSimilar Manga Retrieval Using Visual Vocabulary Based on Regions of InterestSun, Weihan et al.📜 Paper
    • Sentence-based Grounding
      YearConference / JournalTitleAuthorsLinks
      2024AAAIGroundVLP: Harnessing Zero-shot Visual Grounding from Vision-Language Pre-training and Open-Vocabulary Object DetectionShen, Haozhan et al.📜 Paper, 👨‍💻 Code
      2024ECCVGrounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object DetectionLiu, Shilong et al.📜 Paper, 👨‍💻 Code
      2020CVPR WorkshopExploring Phrase Grounding without Training: Contextualisation and Extension to Text-Based Image RetrievalParcalabescu, Letitia et al.📜 Paper
      2019AAAIZero-Shot Object Detection with Textual DescriptionsLi, Zhihui et al.📜 Paper
  • Analysis

    • Text-Character association
      YearConference / JournalTitleAuthorsLinks
      2024ArxivTails Tell Tales: Chapter-Wide Manga Transcriptions with Character NamesRagav Sachdeva et al.📜 Paper, 👨‍💻 Code
      2024NeurIPSCoMix: A Comprehensive Benchmark for Multi-Task Comic UnderstandingEmanuele Vivoli et al.📜 Paper, 👨‍💻 Code
      2024MANPU (ICDAR)Spatially Augmented Speech Bubble to Character Association via Comic Multi-task LearningSoykan, Gurkan et al.📜 Paper 💻 Code
      2024CVPRThe Manga Whisperer: Automatically Generating Transcriptions for ComicsSachdeva, Ragav et al.📜 Paper, 👨‍💻 Code
      2023arXivManga109Dialog A Large-scale Dialogue Dataset for Comics Speaker DetectionLi, Yingxuan et al.📜 Paper
      2022IIAI-AAIAlgorithms for Estimation of Comic Speakers Considering Reading Order of Frames and TextsOmori, Yuga et al.📜 Paper
      2019IJDARComic MTL: Optimized Multi-Task Learning for Comic Book Image AnalysisNguyen, Nhu-Van et al.📜 Paper
      2015ICDARSpeech Balloon and Speaker Association for Comics and Manga UnderstandingRigaud, Christophe et al.📜 Paper
    • Panel Sorting
      YearConference / JournalTitleAuthorsLinks
      2017ICDARStory Pattern Analysis Based on Scene Order Information in Four-Scene ComicsUeno, Miki et al.📜 Paper
    • Dialog transcription
      YearConference / JournalTitleAuthorsLinks
      2024ArxivTails Tell Tales: Chapter-Wide Manga Transcriptions with Character NamesRagav Sachdeva et al.📜 Paper, 👨‍💻 Code
      2024NeurIPSCoMix: A Comprehensive Benchmark for Multi-Task Comic UnderstandingEmanuele Vivoli et al.📜 Paper, 👨‍💻 Code
      2024CVPRThe Manga Whisperer: Automatically Generating Transcriptions for ComicsSachdeva, Ragav et al.📜 Paper, 👨‍💻 Code
      2023arXivManga109Dialog A Large-scale Dialogue Dataset for Comics Speaker DetectionLi, Yingxuan et al.📜 Paper
    • Translation
      YearConference / JournalTitleAuthorsLinks
      2024ArXivContext-Informed Machine Translation of Manga using Multimodal Large Language ModelsLippmann, Philip et al.📜 Paper, 👨‍💻 Code
      2024ArXivLarge Language Models as Manga Translators: A Case StudyZhishen Yang et al.paper
      2024ArXivGenerating Visual Stories with Grounded and Coreferent CharactersDanyang Liu et al.paper
      2024ICKECSThe Future of Graphic Novel Translation: Fully Automated SystemsSandeep Singh et al.paper
      2024ACM MultimediaZero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal FusionYingxuan Li et al.paper
      2023ArXivMulti-Teacher Knowledge Distillation For Text Image Machine TranslationCong Ma et al.paper
      2020ArXivTowards Fully Automated Manga TranslationRyota Hinami et al.paper
      2014inTRAlineaVisual adaptation in translated comicsFederico, Zanettinpaper
  • Segmentation

    • Instance Segmentation
      YearConference / JournalTitleAuthorsLinks
      2024AI4VA (ECCV)Unlocking Comics: The AI4VA Dataset for Visual UnderstandingGrönquist, Peter et al.📜 Paper,👨‍💻 Code
      2024ICDARInvestigating Neural Networks and Transformer Models for Enhanced Comic DecodingKouletou, Eleanna et al.📜 Paper
      2022DataverseNLThe Visual Language Research Corpus (VLRC) ProjectCohn, Neil📜 Paper

Layer 3: Retrieval and Modification

  • Retrieval

    • Image-Text Retrieval
      YearConference / JournalTitleAuthorsLinks
      2014DASA Study to Achieve Manga Character Retrieval Method for Manga ImagesIwata, M. et al.📜 Paper
      2011ICDARSimilar Manga Retrieval Using Visual Vocabulary Based on Regions of InterestSun, Weihan et al.📜 Paper
      2011CAVWComic Character Animation Using Bayesian EstimationChou, Yun-Feng et al.📜 Paper
      2010ICGCSearching Digital Political CartoonsWu, Yejun📜 Paper
    • Text-Image Retrieval
      YearConference / JournalTitleAuthorsLinks
      2014ICIPSketch2Manga: Sketch-based Manga RetrievalMatsui, Yusuke et al.📜 Paper
    • Composed Image Retrieval
      YearConference / JournalTitleAuthorsLinks
      2023ArxivMaRU: A Manga Retrieval and Understanding System Connecting Vision and LanguageShen, Conghao Tom et al.📜 Paper
      2022DICTAComicLib: A New Large-Scale Comic Dataset for Sketch UnderstandingWei, Xin et al.📜 Paper
      2021ICDARManga-MMTL: Multimodal Multitask Transfer Learning for Manga Character AnalysisNguyen, Nhu-Van et al.📜 Paper
      2017ICDARSketch-Based Manga Retrieval Using Deep FeaturesNarita, Rei et al.📜 Paper
      2017ArxivA Neural Representation of Sketch DrawingsHa, David et al.📜 Paper
      2017ArxivStyle Transfer for Anime Sketches with Enhanced Residual U-net and Auxiliary Classifier GANZhang, Lvmin et al.📜 Paper
      2015MM-TASketch-Based Manga Retrieval Using Manga109 DatasetMatsui, Yusuke et al.📜 Paper
    • Personalized Image Retrieval
      YearConference / JournalTitleAuthorsLinks
      2022BMVCPersonalised CLIP or: How to Find Your Vacation VideosKorbar, Bruno et al.📜 Paper
  • Modification

    • Image Impainting and Editing
      YearConference / JournalTitleAuthorsLinks
      2022ACM-UISTCodeToon: Story Ideation, Auto Comic Generation, and Structure Mapping for Code-Driven StorytellingSuh, Sangho et al.📜 Paper
      2022TVCGInteractive Data ComicsWang, Zezhong et al.📜 Paper

Layer 4: Understanding

  • Understanding

    • Visual Entailment
      YearConference / JournalTitleAuthorsLinks
    • Visual-Question Answer
      YearConference / JournalTitleAuthorsLinks
      2022WACVChallenges in Procedural Multimodal Machine Comprehension: A Novel Way To BenchmarkSahu, Pritish et al.📜 Paper
      2021ArxivTowards Solving Multimodal ComprehensionSahu, Pritish et al.📜 Paper
      2020MDPI-ASA Survey on Machine Reading Comprehension—Tasks, Evaluation Metrics and Benchmark DatasetsZeng, Changchang et al.📜 Paper
      2017IIWASComicQA: Contextual Navigation Aid by Hyper-Comic RepresentationSumi, Yasuyuki et al.📜 Paper
      2016MANPU (ICPR)Designing a Question-Answering System for Comic ContentsMoriyama, Yukihiro et al.📜 Paper
    • Visual-Dialog
      YearConference / JournalTitleAuthorsLinks
    • Visual Reasoning
      YearConference / JournalTitleAuthorsLinks

Layer 5: Generation and Synthesis

  • Generation

    • Comics generation from other media
      YearConference / JournalTitleAuthorsLinks
      2023SIGCSEDeveloping Comic-based Learning Toolkits for Teaching Computing to Elementary School LearnersCastro, Francico et al.📜 Paper
      2022THMSAugmenting Conversations With Comic-Style Word BalloonsZhang, H. et al.📜 Paper
      2022LACLOComics as a Pedagogical Tool for TeachingLima, Antonio Alexandre et al.📜 Paper
      2021TVCGChartStory: Automated Partitioning, Layout, and Captioning of Charts into Comic-Style NarrativesZhao, Jian et al.📜 Paper
      2021SIGCSEUsing Comics to Introduce and Reinforce Programming Concepts in CS1Suh, Sangho et al.📜 Paper
      2021MM-CCAAutomatic Comic Generation with Stylistic Multi-page Layouts and Emotion-driven Text Balloon GenerationYang, Xin et al.📜 Paper
      2018ACMComixify: Transform Video into a ComicsPesko, Maciej et al.📜 Paper
      2015TOMMContent-Aware Video2Comics With Manga-Style LayoutJing, Guangmei et al.📜 Paper
      2012TOMMMovie2Comics: Towards a Lively Video Content PresentationWang, Meng et al.📜 Paper
      2012ACM-TGAutomatic Stylistic Manga LayoutCao, Ying et al.📜 Paper
      2012TOMMScalable Comic-like Video Summaries and Layout DisturbanceHerranz, Luis et al.📜 Paper
      2011ACM-MMAutomatic Preview Generation of Comic Episodes for Digitized Comic SearchHoashi, Keiichiro et al.📜 Paper
      2011ISPACSAutomatic Comic Strip Generation Using Extracted Keyframes from Cartoon AnimationTanapichet, Pakpoom et al.📜 Paper
      2011ICMLCCaricaturation for Human Face PicturesChang, I-Cheng et al.📜 Paper
      2010SICEComic Live Chat Communication Tool Based on Concept of DowngradingMatsuda, Misaki et al.📜 Paper
      2010CAIDCDResearch and Development of the Generation in Japanese Manga Based on Frontal Face ImageXuexiong, Deng et al.📜 Paper
    • Comics to Scene graph
      YearConference / JournalTitleAuthorsLinks
    • Image-2-Text Generation
      YearConference / JournalTitleAuthorsLinks
      2024NeurIPSCracking the Code of Juxtaposition: Can AI Models Understand the Humorous ContradictionsHu, Zhe et al.📜 Paper
      2024AI4VA (ECCV)ComiCap: A VLMs pipeline for dense captioning of Comic PanelsVivoli, Emanuele et al.📜 Paper
      2024ICDARMultimodal Transformer for Comics Text-ClozeVivoli, Emanuele et al.📜 Paper
      2024MANPU (ICDAR)Toward Accessible Comics for Blind and Low Vision ReadersRigaud, Christophe et al.📜 Paper
      2024CVPRThe Manga Whisperer: Automatically Generating Transcriptions for ComicsSachdeva, Ragav et al.📜 Paper, 👨‍💻 Code
      2023ArxivComics for Everyone: Generating Accessible Text Descriptions for Comic StripsRamaprasad, Reshma et al.📜 Paper
      2023ACLMultimodal Persona Based Generation of Comic DialogsAgrawal, Harsh et al.📜 Paper
      2023ArxivM2C: Towards Automatic Multimodal Manga ComplementGuo, Hongcheng et al.📜 Paper
    • Text-2-Image Generation
      YearConference / JournalTitleAuthorsLinks
      2023ICCVDiffusion in StyleEveraert, Martin Nicolas et al.📜 Paper, 👨‍💻 Code
      2023MDPI-ASA Study on Generating Webtoons Using Multilingual Text-to-Image ModelsYu, Kyungho et al.📜 Paper
      2023ArxivGenerating Coherent Comic with Rich Story Using ChatGPT and Stable DiffusionJin, Ze et al.📜 Paper
      2022ISMConditional GAN for Small DatasetsHiruta, Komei et al.📜 Paper
      2021NAACLImproving Generation and Evaluation of Visual Stories via Semantic ConsistencyMaharana, Adyasha et al.📜 Paper, 👨‍💻 Code
      2021CoRRIntegrating Visuospatial, Linguistic and Commonsense Structure intoStory VisualizationMaharana, Adyasha et al.📜 Paper, 👨‍💻 Code
      2021ICCCA Deep Learning Pipeline for the Synthesis of Graphic NovelsMelistas, Thomas et al.📜 Paper
      2021ArxivComicGAN: Text-to-Comic Generative Adversarial NetworkProven-Bessel, Ben et al.📜 Paper, 👨‍💻 Code
      2019CVPRStoryGAN: A Sequential Conditional GAN for Story VisualizationLi, Yitong et al.📜 Paper, 👨‍💻 Code
      2018CVPRCross-Domain Weakly-Supervised Object Detection through Progressive Domain AdaptationInoue, Naoto et al.📜 Paper, 👨‍💻 Code
      2017ArxivTowards the Automatic Anime Characters Creation with Generative Adversarial NetworksJin, Yanghua et al.📜 Paper, 👨‍💻 Code
    • Scene-graph Generation for captioning
      YearConference / JournalTitleAuthorsLinks
    • Sound generation
      YearConference / JournalTitleAuthorsLinks
      2023ACM-TACAccessComics2: Understanding the User Experience of an Accessible Comic Book Reader for Blind People with Textual Sound EffectsLee, Yun Jung et al.📜 Paper
      2019ACM-TGComic-Guided Speech SynthesisWang, Yujia et al.📜 Paper
  • Synthesis

    • 3D Generation from Images
      YearConference / JournalTitleAuthorsLinks
      2023ECCVAnimeCeleb: Large-Scale Animation CelebHeads Dataset for Head ReenactmentKim, Kangyeol et al.📜 Paper, 👨‍💻 Code
      2023IJCAICollaborative Neural Rendering Using Anime Character SheetsLin, Zuzeng et al.📜 Paper, 👨‍💻 Code
      2023CVPRPAniC-3D: Stylized Single-view 3D Reconstruction from Portraits of Anime CharactersChen, Shuhong et al.📜 Paper
      2023ArxivSketch-A-Shape: Zero-Shot Sketch-to-3D Shape GenerationSanghi, Aditya et al.📜 Paper
      2021N/ATalking Head Anime from a Single Image 2: More ExpressiveKhungurn, Pramook et al.👨‍💻 Code
      2020ICLRU-GAT-IT: Unsupervised Generative Attentional Networks with Adaptive Layer-Instance Normalization for Image-to-Image TranslationKim, Junho et al.📜 Paper, 👨‍💻 Code
      20173DV3D Shape Reconstruction from Sketches via Multi-view Convolutional NetworksLun, Zhaoliang et al.📜 Paper
    • Video generation
      YearConference / JournalTitleAuthorsLinks
      2023ArxivDreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text GuidanceWang, Cong et al.📜 Paper, 👨‍💻 Code
      2023ArxivPhotorealistic Video Generation with Diffusion ModelsGupta, Agrim et al.📜 Paper
      2023ArxivMotion-Conditioned Image Animation for Video EditingYan, Wilson et al.📜 Paper, 👨‍💻 Code
      2021ICDARC2VNet: A Deep Learning Framework Towards Comic Strip to Audio-Visual Scene SynthesisGupta, Vaibhavi et al.📜 Paper, 👨‍💻 Code
      2016TOMMDynamic Manga: Animating Still Manga via Camera MovementCao, Ying et al.📜 Paper
    • Narrative-based complex scene generation
      YearConference / JournalTitleAuthorsLinks
      2024WACVSynthesizing Coherent Story with Auto-Regressive Latent Diffusion ModelsPan, Xichen et al.📜 Paper, 👨‍💻 Code
      2023CVPRMake-A-Story: Visual Memory Conditioned Consistent Story GenerationRahman, Tanzila et al.📜 Paper
      2023NeurIPS WorkshopPersonalized Comic Story GenerationPeng, Wenxuan et al.📜 Paper
      2022ECCVStoryDALL-E: Adapting Pretrained Text-to-Image Transformers for Story ContinuationMaharana, Adyasha et al.📜 Paper, 👨‍💻 Code
      2022EMNLPCharacter-Centric Story Visualization via Visual Planning and Token AlignmentChen, Hong et al.📜 Paper, 👨‍💻 Code
      2021NAACLImproving Generation and Evaluation of Visual Stories via Semantic ConsistencyMaharana, Adyasha et al.📜 Paper, 👨‍💻 Code
      2018CoRRStoryGAN: A Sequential Conditional GAN for Story VisualizationYitong Li et al.📜 Paper

Datasets & Benchmarks 📂📎

  • Datasets

    Overview of Comic/Manga Datasets and Tasks

    This table provides an overview of Comic/Manga datasets and tasks, including information on their availability, published year, source, and properties such as languages, number of comic/manga books, and pages. The rows are repeated according to the supported tasks. Accessibility is indicated with ⚠️ for no longer existing datasets, ❌ indicates existing but not accessible, and ✅ means existing and accessible. The link [proj] directs to the project websites, while [data] directs to dataset websites. For CoMix, mix means that it inherits from a mixture of four datasets.

    TaskNameYearAccess.LanguageOrigin# books# pages
    Image ClassificationSequencity [proj]2017⚠️EN, JP--140000
    BAM! [proj]2017⚠️---2500000
    Manga109 [proj][data]2018✅JP1970-201010921142
    EmoRecCom [proj][data]2021✅EN1938-1954--
    Object DetectionFahad18 [proj]2012❌---586
    eBDtheque [proj][data]2013✅EN, FR, JP1905-201225100
    sun70 [proj]2013❌FR-660
    COMICS [proj][data]2017✅EN1938-19543948198657
    BAM! [proj]2017⚠️---2500000
    JC2463 [proj]2017❌JP-142463
    AEC912 [proj]2017❌EN, FR--912
    GCN [proj][data]2017❌EN, JP1978-201325338000
    Sequencity612 [proj]2017⚠️EN, JP--612
    SSGCI [proj][data]2016❌EN, FR, JP1905-2012-500
    Comics3w [proj]2017❌JP, EN-10329845
    comics2k [proj][data]2018⚠️----
    DCM772 [proj][data]2018✅EN1938-195427772
    Manga109 [proj][data]2018✅JP1970-201010921142
    BCBId [proj][data]2022✅BN-643327
    COO [proj][data]2022✅JP1970-201010910602
    COMICS-Text+ [proj][data]2022✅EN1938-19543948198657
    PopManga [proj][data]2024✅EN1990-2020251925
    CoMix [proj][data]2024✅EN, FR1938-20231003800
    Re-IdentificationFahad18 [proj]2012❌---586
    Ho422013❌---42
    Manga109 [proj][data]2018✅JP1970-201010921142
    PopManga [proj][data]2024✅EN1990-2020251925
    CoMix [proj][data]2024✅EN, FR1938-20231003800
    LinkingeBDtheque [proj][data]2013✅EN, FR, JP1905-201225100
    sun702013❌FR-660
    GCN [proj][data]2017❌EN, JP1978-201325338000
    Manga109 [proj][data]2018✅JP1970-201010921142
    PopManga [proj][data]2024✅EN1990-2020251925
    CoMix [proj][data]2024✅EN, FR1938-20231003800
    SegmentationSequencity4k [proj]2020⚠️EN, FR, JP--4479
    Dialog GenerationPopManga [proj][data]2024✅EN1990-2020251925
    CoMix [proj][data]2024✅EN, FR1938-20231003800
    UnknownVLRC [proj][data]2023❌JP, FR, EN, 6+1940-present3767773

Venues

  • Journals
    • TPAMI: IEEE Transactions on Pattern Analysis and Machine Intelligence
    • TIP: IEEE Transactions on Image Processing
    • TOMM: IEEE Transactions on Multimedia
    • TVCG: IEEE Transactions on Visualization and Computer Graphics
    • TCSVT: IEEE Transactions on Circuits and Systems for Video Technology
    • THMS: IEEE Transactions on Human-Machine Systems
    • ACM-TG: Transactions on Graphics
    • ACM-TAC: Transactions on Accessible Computing
    • IJHCI: International Journal on Human-Computer Interaction
    • IJI: The International Journal on the Image
    • IJCG: The Visual Computer: International Journal of Computer Graphics
    • IJDAR: International Journal on Document Analysis and Recognition
    • MM-CCA: Transaction on Multimedia Computing, Communication and Applications
  • Conferences
    • NeurIPS: Neural Information Processing Systems
    • ICML: International Conference on Machine Learning
    • CVPR: IEEE/CVF Conference on Computer Vision and Pattern Recognition
    • ICCV: IEEE/CVF International Conference of Computer Vision
    • ECCV: IEEE/CVF European Conference of Computer Vision
    • WACV: IEEE/CVF Winter Conference on Applications of Computer Vision
    • SciVis: IEEE Scientific Visualization Conference
    • ICIP: IEEE International Conference on Image Processing
    • VCIP: IEEE International Conference Visual Communication Image Process
    • CSNT: IEEE International Conference on Communication Systems and Network Technologies
    • CAIDCD: IEEE International Conference on Computer-Aided Industrial Design and Conceptual Design
    • ACM: Association for Computing Machinery
    • ICDAR: IAPR International Conference on Document Analysis and Recognition
    • ICPR: International Conference on Pattern Recognition
    • ICIR: International Conference on Intelligent Reality
    • IIAI-AAI: International Congress on Advanced Applied Informatics
    • MMM: Multimedia Modeling
    • LREC: International Conference on Language Resources and Evaluation
    • MTA: Multimedia Tools and Applications
    • ICIT: International Conference on Information Technology
    • ICIAP: International Conference on Image Analysis and Processing
    • IJCNN: International Joint Conference on Neural Networks
    • ACPR: IAPR Asian Conference on Pattern Recognition
    • ICGC: IEEE International Conference on Granular Computing
    • CAVW: Computer Animation and Virtual Worlds
    • MM-TA: Multimedia Tools and Applications
    • DICTA: International Conference on Digital Image Computing: Techniques and Applications
    • UIST: ACM Symposium on User Interface Software and Technology
    • EMNLP: ACM Conference on Empirical Methods in Natural Language Processing
    • IIWAS: International Conference on Information Integration and Web-based Applications and Services
    • MDPI-AS: MDPI Applied Science
    • ICMLC: International Conference on Machine Learning and Cybernetics
    • LACLO: Latin American Conference on Learning Technologies
    • ACL: Association for Computational Linguistics
    • ICCC: International Conference on Computational Creativity
    • 3DV: International Conference on 3D Vision
  • Workshops
    • MANPU: IAPR International Workshop on Comics Analysis, Processing and Understanding
    • DAS: IAPR International Workshop on Document Analysis Systems
    • IWAIT: International Workshop on Advanced Image Technology
    • ISPACS: Symposium on Intelligent Signal Processing and Communication Systems
    • SIGCSE: ACM Technical Symposium on Computer Science Education
    • ISM: IEEE International Symposium in Multimedia

Links

🔧 Tools & Repositories

How to Contribute 🚀

You can contribute in two ways:

  1. The easiest is to open an Issue (see an example in issue #1) and we can discuss if there are missing papers, wrong associations or links, or misspelled venues.
  2. The second one is making a pull request with the implemented changes, following the steps:
    1. Fork this repository and clone it locally.
    2. Create a new branch for your changes: git checkout -b feature-name.
    3. Make your changes and commit them: git commit -m 'Description of the changes'.
    4. Push to your fork: git push origin feature-name.
    5. Open a pull request on the original repository by providing a description of your changes.

This project is in constant development, and we welcome contributions to include the latest research papers in the field or report issues 💥💥.

Star History ⭐

Star History Chart

Acknowledge

Many thanks to my co-authors for taking the time to help me with the various refactoring of the survey. Thanks to Beppe Folder for its Awesome Human Visual Attention repo that inspired the ✨style✨ of this repository.

Significant stargazers

Marcello Seri

192 followers · starred Sep 2024

emanuelevivoli/awesome-comics-understanding

The official repo of the Comics Survey: "A missing piece in Vision and Language: A Survey on Comics Understanding"

141

16 commits

updated Jan 2, 2025

See the code

README

Awesome Comics Understanding Awesome

This repository contains a curated list of research papers and resources focusing on Comics Understanding.

arXiv

🔥 One missing piece in Vision and Language: A Survey on Comics Understanding 🔥

Authors: Emanuele Vivoli, Andrey Barsky, Mohamed Ali Souibgui, Artemis Llabrés, Marco Bertini, Dimosthenis Karatzas

📣 Latest News 📣

  • 🚧 This repo is a work in progress, please contribute here
  • 14 September 2024 Our survey paper have dropped in arXiv !!

📚 Table of Contents

Overview of Vision-Language Tasks of the Layers of Comics Understanding. The ranking is based on input and output modalities and dimensions, as illustrated in the paper.

Layers of Comics Understanding

Every survey worthy of the name includes illustrative visuals to enhance understanding. We've followed this approach by providing examples for each task in the Layer of Comics Understanding.
Go check every Layer's tasks image ⬇️.

Layer 1: Tagging and Augmentation

  • Tagging

    • Image classification
      YearConference / JournalTitleAuthorsLinks
      2023TIPPanel-Page-Aware Comic Genre UnderstandingXu, Chenshu et al.📜 Paper
      2019ICDAR WorkshopAnalysis Based on Distributed Representations of Various Parts Images in Four-Scene Comics Story DatasetTerauchi, Akira et al.📜 Paper
      2018TPAMILearning Consensus Representation for Weak Style ClassificationJiang, Shuhui et al.📜 Paper
      2018ICDARComic Story Analysis Based on Genre ClassificationDaiku, Yuki et al.📜 Paper
      2017ICDARHistogram of Exclamation Marks and Its Application for Comics AnalysisHiroe, Sotaro et al.📜 Paper
      2014ACM MultimediaLine-Based Drawing Style Description for Manga ClassificationChu, Wei-Ta et al.📜 Paper
    • Emotion classification
      YearConference / JournalTitleAuthorsLinks
      2023MMMManga Text Detection with Manga-Specific Data Augmentation and Its Applications on Emotion AnalysisYang, Yi-Ting et al.📜 Paper
      2021ICDARCompetition on Multimodal Emotion Recognition on Comics ScenesNguyen, Nhu-Van et al.📜 Paper, 👨‍💻 Code
      2016MANPU (ICPR)Manga Content Analysis Using Physiological SignalsSanches, Charles Lima et al.📜 Paper
      2015IIAI-AAIRelation Analysis between Speech Balloon Shapes and Their Serif Descriptions in ComicTanaka, Hideki et al.📜 Paper
    • Action Detection
      YearConference / JournalTitleAuthorsLinks
      2024ArxivMangaUB: A Manga Understanding Benchmark for Large Multimodal ModelsIkuta, Hikaru et al.📜 Paper
      2024MANPU (ICDAR)ComicBERT: A Transformer Model and Pre-training Strategy for Contextual Understanding in ComicsSoykan, Gurkan et al.📜 Paper, 👨‍💻 Code
      2024ICDARMultimodal Transformer for Comics Text-ClozeVivoli, Emanuele et al.📜 Paper
      2020ArxivA Comprehensive Study of Deep Video Action RecognitionZhu, Yi et al.📜 Paper, 👨‍💻 Code
      2017CVPRThe Amazing Mysteries of the Gutter: Drawing Inferences Between Panels in Comic Book NarrativesIyyer, Mohit et al.📜 Paper
    • Page Stream Segmentation
      YearConference / JournalTitleAuthorsLinks
      2022ICPRSemantic Parsing of Interpage RelationsDemirtaş, Mehmet Arif et al.📜 Paper
      2018LRECPage Stream Segmentation with Convolutional Neural Nets Combining Textual and Visual FeaturesWiedemann, Gregor et al.📜 Paper
      2013ICDARDocument Classification and Page Stream Segmentation for Digital Mailroom ApplicationsGordo, Albert et al.📜 Paper
  • Augmentation ( Image-2-Image )

    • Image Super-Resolution
      YearConference / JournalTitleAuthorsLinks
      2023MTAAutomatic Dewarping of Camera-Captured Comic Document ImagesGarai, Arpan et al.📜 Paper
    • Style Transfer
      YearConference / JournalTitleAuthorsLinks
      2023ArxivInkn'hue: Enhancing Manga Colorization from Multiple Priors with Alignment Multi-Encoder VAEJiramahapokee, Tawin📜 Paper, 👨‍💻 Code
      2023IEEE AccessRobust Manga Page Colorization via Coloring Latent SpaceGolyadkin, Maksim et al.📜 Paper
      2023TVCGShading-Guided Manga Screening from ReferenceWu, Huisi et al.📜 Paper
      2022ArxivDASS-Detector: Domain-Adaptive Self-Supervised Pre-Training for Face & Body Detection in DrawingsTopal, Barış Batuhan et al.📜 Paper, 👨‍💻 Code
      2021CVPRGenerating Manga from Illustrations via Mimicking Manga Creation WorkflowZhang, LM et al.*📜 Paper, 👨‍💻 Code
      2021CVPRUnbiased Mean Teacher for Cross-domain Object DetectionDeng, Jinhong et al.📜 Paper, 👨‍💻 Code
      2021CVPREncoding in Style: A StyleGAN Encoder for Image-to-Image TranslationRichardson, Elad et al.📜 Paper, 👨‍💻 Code
      2021AAAIMangaGAN: Unpaired Photo-to-Manga Translation Based on The Methodology of Manga DrawingSu, Hao et al.📜 Paper
      2019ISMSynthesis of Screentone Patterns of Manga CharactersTsubota, K. et al.📜 Paper, 👨‍💻 Code
      2018SciVisColor Interpolation for Non-Euclidean Color SpacesZeyen, Max et al.📜 Paper
      2017ACM-SIGGRAPH AsiaComicolorization: Semi-automatic Manga ColorizationFurusawa, Chie et al.📜 Paper, 👨‍💻 Code
      2017ICDARCGAN-Based Manga Colorization Using a Single Training ImageHensman, Paulina et al.📜 Paper, 👨‍💻 Code
      2017CVPRImage-to-Image Translation with Conditional Adversarial NetworksIsola, Phillip et al.📜 Paper
      2017ACM-TGDeep Extraction of Manga Structural LinesLi, Chengze et al.📜 Paper, 👨‍💻 Code
    • Vectorization
      YearConference / JournalTitleAuthorsLinks
      2023TCSVTMARVEL: Raster Gray-level Manga Vectorization via Primitive-wise Deep Reinforcement LearningH. Su et al.📜 Paper, 👨‍💻 Code
      2022CVPRTowards Layer-wise Image VectorizationMa, Xu et al.📜 Paper, 👨‍💻 Code
      2017ACM-TGDeep Extraction of Manga Structural LinesLi, Chengze et al.📜 Paper, 👨‍💻 Code
      2017TVCGManga Vectorization and Manipulation with Procedural Simple ScreentoneYao, Chih-Yuan et al.📜 Paper
      2011ACM-SIGGRAPHDepixelizing Pixel ArtKopf, Johannes et al.📜 Paper
      2003N/APotrace : A Polygon-Based Tracing AlgorithmSelinger, Peter📜 Paper, 👨‍💻 Code
    • Depth Estimation
      YearConference / JournalTitleAuthorsLinks
      2023CVPR WorkshopDense Multitask Learning to Reconfigure ComicsBhattacharjee, Deblina et al.📜 Paper
      2022WACVEstimating Image Depth in the Comics DomainBhattacharjee, Deblina et al.📜 Paper, 👨‍💻 Code
      2022CVPRMulT: An End-to-End Multitask Learning TransformerBhattacharjee, Deblina et al.📜 Paper, 👨‍💻 Code

Layer 2: Grounding, Analysis and Segmentation

  • Grounding

    • Object detection
      YearConference / JournalTitleAuthorsLinks
      2024MANPU (ICDAR)Comics Datasets Framework: Mix of Comics datasets for detection benchmarkingVivoli, Emanuele et al.📜 Paper
      2024MANPU (ICDAR)A Comprehensive Gold Standard and Benchmark for Comics Text Detection and RecognitionGurkan Soykan et al.📜 Paper, 👨‍💻 Code
      2023MMMManga Text Detection with Manga-Specific Data Augmentation and Its Applications on Emotion AnalysisYang, \relax YT et al.📜 Paper
      2023CSNTCPD: Faster RCNN-based DragonBall Comic Panel DetectionSharma, Rishabh et al.📜 Paper
      2022IJDARBCBId: First Bangla Comic Dataset and Its ApplicationsDutta, Arpita et al.📜 Paper
      2022ECCVCOO/ Comic Onomatopoeia Dataset for Recognizing Arbitrary or Truncated TextsBaek, Jeonghun et al.📜 Paper, 👨‍💻 Code
      2019ICDAR WorkshopWhat Do We Expect from Comic Panel Extraction?Nguyen Nhu, Van et al.📜 Paper
      2019ICDAR WorkshopCNN Based Extraction of Panels/Characters from Bengali Comic Book Page ImagesDutta, Arpita et al.📜 Paper
      2018VCIPText Detection in Manga by Deep Region Proposal, Classification, and RegressionChu, Wei-Ta et al.📜 Paper
      2018IWAITA Study on Object Detection Method from Manga Images Using CNNYanagisawa, Hideaki et al.📜 Paper
      2018IWAITA Study on Object Detection Method from Manga Images Using CNNYanagisawa, Hideaki et al.📜 Paper
      2017ICDARA Faster R-CNN Based Method for Comic Characters Face DetectionQin, Xiaoran et al.📜 Paper
      2016IJCGText-Aware Balloon Extraction from MangaLiu, Xueting et al.📜 Paper
      2016IJCNNLine-Wise Text Identification in Comic Books: A Support Vector Machine-Based ApproachPal, Srikanta et al.📜 Paper
      2016ICIPText Detection in Manga by Combining Connected-Component-Based and Region-Based ClassificationsAramaki, Yuji et al.📜 Paper
      2015ICIAPPanel Tracking for the Extraction and the Classification of Speech BalloonsJomaa, Hadi S. et al.📜 Paper
      2012DASPanel and Speech Balloon Extraction from Comic BooksHo, Anh Khoi Ngo et al.📜 Paper
      2011IJIMethod for Real Time Text Extraction of Digital Manga ComicArai, Kohei et al.📜 Paper
      2011ICDARRecognizing Text Elements for SVG Comic Compression and Its Novel ApplicationsSu, Chung-Yuan et al.📜 Paper
      2010ICITMethod for Automatic E-Comic Scene Frame Extraction for Reading Comic on Mobile DevicesArai, Kohei et al.📜 Paper
      2009IJHCIEnhancing the Accessibility for All of Digital Comic BooksPonsard, Christophe📜 Paper
    • Character Re-Identification
      YearConference / JournalTitleAuthorsLinks
      2024ArxivTails Tell Tales: Chapter-Wide Manga Transcriptions with Character NamesRagav Sachdeva et al.📜 Paper, 👨‍💻 Code
      2024NeurIPSCoMix: A Comprehensive Benchmark for Multi-Task Comic UnderstandingEmanuele Vivoli et al.📜 Paper, 👨‍💻 Code
      2024CVPRThe Manga Whisperer: Automatically Generating Transcriptions for ComicsSachdeva, Ragav et al.📜 Paper, 👨‍💻 Code
      2023IET Image ProcessingToward Cross-Domain Object Detection in Artwork Images Using Improved YoloV5 and XGBoostingAhmad, Tasweer et al.📜 Paper
      2023ArxivIdentity-Aware Semi-Supervised Learning for Comic Character Re-IdentificationSoykan, Gürkan et al.📜 Paper
      2023ACM-MM AsiaOcclusion-Aware Manga Character Re-Identification with Self-Paced Contrastive LearningZhang, Ci-Yin et al.📜 Paper
      2022ArxivUnsupervised Manga Character Re-Identification via Face-Body and Spatial-Temporal Associated ClusteringZhang, Z et al.📜 Paper
      2022ICIRCAST: Character Labeling in Animation Using Self‐supervision by TrackingNir, Oron et al.📜 Paper
      2020ICPRDual Loss for Manga Character Recognition with Imbalanced Training DataLi, Yonggang et al.📜 Paper
      2020ICMLA Simple Framework for Contrastive Learning of Visual RepresentationsChen, Ting et al.📜 Paper
      2015ACPRSimilarity Learning Based on Pool-Based Active Learning for Manga Character RetrievalIwata, Motoi et al.📜 Paper
      2014DASA Study to Achieve Manga Character Retrieval Method for Manga ImagesIwata, M. et al.📜 Paper
      2012CVPRColor Attributes for Object DetectionKhan, Fahad Shahbaz et al.📜 Paper
      2012ECCVPHOG Analysis of Self-Similarity in Aesthetic ImagesRedies, Christoph et al.📜 Paper
      2011ICDARSimilar Manga Retrieval Using Visual Vocabulary Based on Regions of InterestSun, Weihan et al.📜 Paper
    • Sentence-based Grounding
      YearConference / JournalTitleAuthorsLinks
      2024AAAIGroundVLP: Harnessing Zero-shot Visual Grounding from Vision-Language Pre-training and Open-Vocabulary Object DetectionShen, Haozhan et al.📜 Paper, 👨‍💻 Code
      2024ECCVGrounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object DetectionLiu, Shilong et al.📜 Paper, 👨‍💻 Code
      2020CVPR WorkshopExploring Phrase Grounding without Training: Contextualisation and Extension to Text-Based Image RetrievalParcalabescu, Letitia et al.📜 Paper
      2019AAAIZero-Shot Object Detection with Textual DescriptionsLi, Zhihui et al.📜 Paper
  • Analysis

    • Text-Character association
      YearConference / JournalTitleAuthorsLinks
      2024ArxivTails Tell Tales: Chapter-Wide Manga Transcriptions with Character NamesRagav Sachdeva et al.📜 Paper, 👨‍💻 Code
      2024NeurIPSCoMix: A Comprehensive Benchmark for Multi-Task Comic UnderstandingEmanuele Vivoli et al.📜 Paper, 👨‍💻 Code
      2024MANPU (ICDAR)Spatially Augmented Speech Bubble to Character Association via Comic Multi-task LearningSoykan, Gurkan et al.📜 Paper 💻 Code
      2024CVPRThe Manga Whisperer: Automatically Generating Transcriptions for ComicsSachdeva, Ragav et al.📜 Paper, 👨‍💻 Code
      2023arXivManga109Dialog A Large-scale Dialogue Dataset for Comics Speaker DetectionLi, Yingxuan et al.📜 Paper
      2022IIAI-AAIAlgorithms for Estimation of Comic Speakers Considering Reading Order of Frames and TextsOmori, Yuga et al.📜 Paper
      2019IJDARComic MTL: Optimized Multi-Task Learning for Comic Book Image AnalysisNguyen, Nhu-Van et al.📜 Paper
      2015ICDARSpeech Balloon and Speaker Association for Comics and Manga UnderstandingRigaud, Christophe et al.📜 Paper
    • Panel Sorting
      YearConference / JournalTitleAuthorsLinks
      2017ICDARStory Pattern Analysis Based on Scene Order Information in Four-Scene ComicsUeno, Miki et al.📜 Paper
    • Dialog transcription
      YearConference / JournalTitleAuthorsLinks
      2024ArxivTails Tell Tales: Chapter-Wide Manga Transcriptions with Character NamesRagav Sachdeva et al.📜 Paper, 👨‍💻 Code
      2024NeurIPSCoMix: A Comprehensive Benchmark for Multi-Task Comic UnderstandingEmanuele Vivoli et al.📜 Paper, 👨‍💻 Code
      2024CVPRThe Manga Whisperer: Automatically Generating Transcriptions for ComicsSachdeva, Ragav et al.📜 Paper, 👨‍💻 Code
      2023arXivManga109Dialog A Large-scale Dialogue Dataset for Comics Speaker DetectionLi, Yingxuan et al.📜 Paper
    • Translation
      YearConference / JournalTitleAuthorsLinks
      2024ArXivContext-Informed Machine Translation of Manga using Multimodal Large Language ModelsLippmann, Philip et al.📜 Paper, 👨‍💻 Code
      2024ArXivLarge Language Models as Manga Translators: A Case StudyZhishen Yang et al.paper
      2024ArXivGenerating Visual Stories with Grounded and Coreferent CharactersDanyang Liu et al.paper
      2024ICKECSThe Future of Graphic Novel Translation: Fully Automated SystemsSandeep Singh et al.paper
      2024ACM MultimediaZero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal FusionYingxuan Li et al.paper
      2023ArXivMulti-Teacher Knowledge Distillation For Text Image Machine TranslationCong Ma et al.paper
      2020ArXivTowards Fully Automated Manga TranslationRyota Hinami et al.paper
      2014inTRAlineaVisual adaptation in translated comicsFederico, Zanettinpaper
  • Segmentation

    • Instance Segmentation
      YearConference / JournalTitleAuthorsLinks
      2024AI4VA (ECCV)Unlocking Comics: The AI4VA Dataset for Visual UnderstandingGrönquist, Peter et al.📜 Paper,👨‍💻 Code
      2024ICDARInvestigating Neural Networks and Transformer Models for Enhanced Comic DecodingKouletou, Eleanna et al.📜 Paper
      2022DataverseNLThe Visual Language Research Corpus (VLRC) ProjectCohn, Neil📜 Paper

Layer 3: Retrieval and Modification

  • Retrieval

    • Image-Text Retrieval
      YearConference / JournalTitleAuthorsLinks
      2014DASA Study to Achieve Manga Character Retrieval Method for Manga ImagesIwata, M. et al.📜 Paper
      2011ICDARSimilar Manga Retrieval Using Visual Vocabulary Based on Regions of InterestSun, Weihan et al.📜 Paper
      2011CAVWComic Character Animation Using Bayesian EstimationChou, Yun-Feng et al.📜 Paper
      2010ICGCSearching Digital Political CartoonsWu, Yejun📜 Paper
    • Text-Image Retrieval
      YearConference / JournalTitleAuthorsLinks
      2014ICIPSketch2Manga: Sketch-based Manga RetrievalMatsui, Yusuke et al.📜 Paper
    • Composed Image Retrieval
      YearConference / JournalTitleAuthorsLinks
      2023ArxivMaRU: A Manga Retrieval and Understanding System Connecting Vision and LanguageShen, Conghao Tom et al.📜 Paper
      2022DICTAComicLib: A New Large-Scale Comic Dataset for Sketch UnderstandingWei, Xin et al.📜 Paper
      2021ICDARManga-MMTL: Multimodal Multitask Transfer Learning for Manga Character AnalysisNguyen, Nhu-Van et al.📜 Paper
      2017ICDARSketch-Based Manga Retrieval Using Deep FeaturesNarita, Rei et al.📜 Paper
      2017ArxivA Neural Representation of Sketch DrawingsHa, David et al.📜 Paper
      2017ArxivStyle Transfer for Anime Sketches with Enhanced Residual U-net and Auxiliary Classifier GANZhang, Lvmin et al.📜 Paper
      2015MM-TASketch-Based Manga Retrieval Using Manga109 DatasetMatsui, Yusuke et al.📜 Paper
    • Personalized Image Retrieval
      YearConference / JournalTitleAuthorsLinks
      2022BMVCPersonalised CLIP or: How to Find Your Vacation VideosKorbar, Bruno et al.📜 Paper
  • Modification

    • Image Impainting and Editing
      YearConference / JournalTitleAuthorsLinks
      2022ACM-UISTCodeToon: Story Ideation, Auto Comic Generation, and Structure Mapping for Code-Driven StorytellingSuh, Sangho et al.📜 Paper
      2022TVCGInteractive Data ComicsWang, Zezhong et al.📜 Paper

Layer 4: Understanding

  • Understanding

    • Visual Entailment
      YearConference / JournalTitleAuthorsLinks
    • Visual-Question Answer
      YearConference / JournalTitleAuthorsLinks
      2022WACVChallenges in Procedural Multimodal Machine Comprehension: A Novel Way To BenchmarkSahu, Pritish et al.📜 Paper
      2021ArxivTowards Solving Multimodal ComprehensionSahu, Pritish et al.📜 Paper
      2020MDPI-ASA Survey on Machine Reading Comprehension—Tasks, Evaluation Metrics and Benchmark DatasetsZeng, Changchang et al.📜 Paper
      2017IIWASComicQA: Contextual Navigation Aid by Hyper-Comic RepresentationSumi, Yasuyuki et al.📜 Paper
      2016MANPU (ICPR)Designing a Question-Answering System for Comic ContentsMoriyama, Yukihiro et al.📜 Paper
    • Visual-Dialog
      YearConference / JournalTitleAuthorsLinks
    • Visual Reasoning
      YearConference / JournalTitleAuthorsLinks

Layer 5: Generation and Synthesis

  • Generation

    • Comics generation from other media
      YearConference / JournalTitleAuthorsLinks
      2023SIGCSEDeveloping Comic-based Learning Toolkits for Teaching Computing to Elementary School LearnersCastro, Francico et al.📜 Paper
      2022THMSAugmenting Conversations With Comic-Style Word BalloonsZhang, H. et al.📜 Paper
      2022LACLOComics as a Pedagogical Tool for TeachingLima, Antonio Alexandre et al.📜 Paper
      2021TVCGChartStory: Automated Partitioning, Layout, and Captioning of Charts into Comic-Style NarrativesZhao, Jian et al.📜 Paper
      2021SIGCSEUsing Comics to Introduce and Reinforce Programming Concepts in CS1Suh, Sangho et al.📜 Paper
      2021MM-CCAAutomatic Comic Generation with Stylistic Multi-page Layouts and Emotion-driven Text Balloon GenerationYang, Xin et al.📜 Paper
      2018ACMComixify: Transform Video into a ComicsPesko, Maciej et al.📜 Paper
      2015TOMMContent-Aware Video2Comics With Manga-Style LayoutJing, Guangmei et al.📜 Paper
      2012TOMMMovie2Comics: Towards a Lively Video Content PresentationWang, Meng et al.📜 Paper
      2012ACM-TGAutomatic Stylistic Manga LayoutCao, Ying et al.📜 Paper
      2012TOMMScalable Comic-like Video Summaries and Layout DisturbanceHerranz, Luis et al.📜 Paper
      2011ACM-MMAutomatic Preview Generation of Comic Episodes for Digitized Comic SearchHoashi, Keiichiro et al.📜 Paper
      2011ISPACSAutomatic Comic Strip Generation Using Extracted Keyframes from Cartoon AnimationTanapichet, Pakpoom et al.📜 Paper
      2011ICMLCCaricaturation for Human Face PicturesChang, I-Cheng et al.📜 Paper
      2010SICEComic Live Chat Communication Tool Based on Concept of DowngradingMatsuda, Misaki et al.📜 Paper
      2010CAIDCDResearch and Development of the Generation in Japanese Manga Based on Frontal Face ImageXuexiong, Deng et al.📜 Paper
    • Comics to Scene graph
      YearConference / JournalTitleAuthorsLinks
    • Image-2-Text Generation
      YearConference / JournalTitleAuthorsLinks
      2024NeurIPSCracking the Code of Juxtaposition: Can AI Models Understand the Humorous ContradictionsHu, Zhe et al.📜 Paper
      2024AI4VA (ECCV)ComiCap: A VLMs pipeline for dense captioning of Comic PanelsVivoli, Emanuele et al.📜 Paper
      2024ICDARMultimodal Transformer for Comics Text-ClozeVivoli, Emanuele et al.📜 Paper
      2024MANPU (ICDAR)Toward Accessible Comics for Blind and Low Vision ReadersRigaud, Christophe et al.📜 Paper
      2024CVPRThe Manga Whisperer: Automatically Generating Transcriptions for ComicsSachdeva, Ragav et al.📜 Paper, 👨‍💻 Code
      2023ArxivComics for Everyone: Generating Accessible Text Descriptions for Comic StripsRamaprasad, Reshma et al.📜 Paper
      2023ACLMultimodal Persona Based Generation of Comic DialogsAgrawal, Harsh et al.📜 Paper
      2023ArxivM2C: Towards Automatic Multimodal Manga ComplementGuo, Hongcheng et al.📜 Paper
    • Text-2-Image Generation
      YearConference / JournalTitleAuthorsLinks
      2023ICCVDiffusion in StyleEveraert, Martin Nicolas et al.📜 Paper, 👨‍💻 Code
      2023MDPI-ASA Study on Generating Webtoons Using Multilingual Text-to-Image ModelsYu, Kyungho et al.📜 Paper
      2023ArxivGenerating Coherent Comic with Rich Story Using ChatGPT and Stable DiffusionJin, Ze et al.📜 Paper
      2022ISMConditional GAN for Small DatasetsHiruta, Komei et al.📜 Paper
      2021NAACLImproving Generation and Evaluation of Visual Stories via Semantic ConsistencyMaharana, Adyasha et al.📜 Paper, 👨‍💻 Code
      2021CoRRIntegrating Visuospatial, Linguistic and Commonsense Structure intoStory VisualizationMaharana, Adyasha et al.📜 Paper, 👨‍💻 Code
      2021ICCCA Deep Learning Pipeline for the Synthesis of Graphic NovelsMelistas, Thomas et al.📜 Paper
      2021ArxivComicGAN: Text-to-Comic Generative Adversarial NetworkProven-Bessel, Ben et al.📜 Paper, 👨‍💻 Code
      2019CVPRStoryGAN: A Sequential Conditional GAN for Story VisualizationLi, Yitong et al.📜 Paper, 👨‍💻 Code
      2018CVPRCross-Domain Weakly-Supervised Object Detection through Progressive Domain AdaptationInoue, Naoto et al.📜 Paper, 👨‍💻 Code
      2017ArxivTowards the Automatic Anime Characters Creation with Generative Adversarial NetworksJin, Yanghua et al.📜 Paper, 👨‍💻 Code
    • Scene-graph Generation for captioning
      YearConference / JournalTitleAuthorsLinks
    • Sound generation
      YearConference / JournalTitleAuthorsLinks
      2023ACM-TACAccessComics2: Understanding the User Experience of an Accessible Comic Book Reader for Blind People with Textual Sound EffectsLee, Yun Jung et al.📜 Paper
      2019ACM-TGComic-Guided Speech SynthesisWang, Yujia et al.📜 Paper
  • Synthesis

    • 3D Generation from Images
      YearConference / JournalTitleAuthorsLinks
      2023ECCVAnimeCeleb: Large-Scale Animation CelebHeads Dataset for Head ReenactmentKim, Kangyeol et al.📜 Paper, 👨‍💻 Code
      2023IJCAICollaborative Neural Rendering Using Anime Character SheetsLin, Zuzeng et al.📜 Paper, 👨‍💻 Code
      2023CVPRPAniC-3D: Stylized Single-view 3D Reconstruction from Portraits of Anime CharactersChen, Shuhong et al.📜 Paper
      2023ArxivSketch-A-Shape: Zero-Shot Sketch-to-3D Shape GenerationSanghi, Aditya et al.📜 Paper
      2021N/ATalking Head Anime from a Single Image 2: More ExpressiveKhungurn, Pramook et al.👨‍💻 Code
      2020ICLRU-GAT-IT: Unsupervised Generative Attentional Networks with Adaptive Layer-Instance Normalization for Image-to-Image TranslationKim, Junho et al.📜 Paper, 👨‍💻 Code
      20173DV3D Shape Reconstruction from Sketches via Multi-view Convolutional NetworksLun, Zhaoliang et al.📜 Paper
    • Video generation
      YearConference / JournalTitleAuthorsLinks
      2023ArxivDreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text GuidanceWang, Cong et al.📜 Paper, 👨‍💻 Code
      2023ArxivPhotorealistic Video Generation with Diffusion ModelsGupta, Agrim et al.📜 Paper
      2023ArxivMotion-Conditioned Image Animation for Video EditingYan, Wilson et al.📜 Paper, 👨‍💻 Code
      2021ICDARC2VNet: A Deep Learning Framework Towards Comic Strip to Audio-Visual Scene SynthesisGupta, Vaibhavi et al.📜 Paper, 👨‍💻 Code
      2016TOMMDynamic Manga: Animating Still Manga via Camera MovementCao, Ying et al.📜 Paper
    • Narrative-based complex scene generation
      YearConference / JournalTitleAuthorsLinks
      2024WACVSynthesizing Coherent Story with Auto-Regressive Latent Diffusion ModelsPan, Xichen et al.📜 Paper, 👨‍💻 Code
      2023CVPRMake-A-Story: Visual Memory Conditioned Consistent Story GenerationRahman, Tanzila et al.📜 Paper
      2023NeurIPS WorkshopPersonalized Comic Story GenerationPeng, Wenxuan et al.📜 Paper
      2022ECCVStoryDALL-E: Adapting Pretrained Text-to-Image Transformers for Story ContinuationMaharana, Adyasha et al.📜 Paper, 👨‍💻 Code
      2022EMNLPCharacter-Centric Story Visualization via Visual Planning and Token AlignmentChen, Hong et al.📜 Paper, 👨‍💻 Code
      2021NAACLImproving Generation and Evaluation of Visual Stories via Semantic ConsistencyMaharana, Adyasha et al.📜 Paper, 👨‍💻 Code
      2018CoRRStoryGAN: A Sequential Conditional GAN for Story VisualizationYitong Li et al.📜 Paper

Datasets & Benchmarks 📂📎

  • Datasets

    Overview of Comic/Manga Datasets and Tasks

    This table provides an overview of Comic/Manga datasets and tasks, including information on their availability, published year, source, and properties such as languages, number of comic/manga books, and pages. The rows are repeated according to the supported tasks. Accessibility is indicated with ⚠️ for no longer existing datasets, ❌ indicates existing but not accessible, and ✅ means existing and accessible. The link [proj] directs to the project websites, while [data] directs to dataset websites. For CoMix, mix means that it inherits from a mixture of four datasets.

    TaskNameYearAccess.LanguageOrigin# books# pages
    Image ClassificationSequencity [proj]2017⚠️EN, JP--140000
    BAM! [proj]2017⚠️---2500000
    Manga109 [proj][data]2018✅JP1970-201010921142
    EmoRecCom [proj][data]2021✅EN1938-1954--
    Object DetectionFahad18 [proj]2012❌---586
    eBDtheque [proj][data]2013✅EN, FR, JP1905-201225100
    sun70 [proj]2013❌FR-660
    COMICS [proj][data]2017✅EN1938-19543948198657
    BAM! [proj]2017⚠️---2500000
    JC2463 [proj]2017❌JP-142463
    AEC912 [proj]2017❌EN, FR--912
    GCN [proj][data]2017❌EN, JP1978-201325338000
    Sequencity612 [proj]2017⚠️EN, JP--612
    SSGCI [proj][data]2016❌EN, FR, JP1905-2012-500
    Comics3w [proj]2017❌JP, EN-10329845
    comics2k [proj][data]2018⚠️----
    DCM772 [proj][data]2018✅EN1938-195427772
    Manga109 [proj][data]2018✅JP1970-201010921142
    BCBId [proj][data]2022✅BN-643327
    COO [proj][data]2022✅JP1970-201010910602
    COMICS-Text+ [proj][data]2022✅EN1938-19543948198657
    PopManga [proj][data]2024✅EN1990-2020251925
    CoMix [proj][data]2024✅EN, FR1938-20231003800
    Re-IdentificationFahad18 [proj]2012❌---586
    Ho422013❌---42
    Manga109 [proj][data]2018✅JP1970-201010921142
    PopManga [proj][data]2024✅EN1990-2020251925
    CoMix [proj][data]2024✅EN, FR1938-20231003800
    LinkingeBDtheque [proj][data]2013✅EN, FR, JP1905-201225100
    sun702013❌FR-660
    GCN [proj][data]2017❌EN, JP1978-201325338000
    Manga109 [proj][data]2018✅JP1970-201010921142
    PopManga [proj][data]2024✅EN1990-2020251925
    CoMix [proj][data]2024✅EN, FR1938-20231003800
    SegmentationSequencity4k [proj]2020⚠️EN, FR, JP--4479
    Dialog GenerationPopManga [proj][data]2024✅EN1990-2020251925
    CoMix [proj][data]2024✅EN, FR1938-20231003800
    UnknownVLRC [proj][data]2023❌JP, FR, EN, 6+1940-present3767773

Venues

  • Journals
    • TPAMI: IEEE Transactions on Pattern Analysis and Machine Intelligence
    • TIP: IEEE Transactions on Image Processing
    • TOMM: IEEE Transactions on Multimedia
    • TVCG: IEEE Transactions on Visualization and Computer Graphics
    • TCSVT: IEEE Transactions on Circuits and Systems for Video Technology
    • THMS: IEEE Transactions on Human-Machine Systems
    • ACM-TG: Transactions on Graphics
    • ACM-TAC: Transactions on Accessible Computing
    • IJHCI: International Journal on Human-Computer Interaction
    • IJI: The International Journal on the Image
    • IJCG: The Visual Computer: International Journal of Computer Graphics
    • IJDAR: International Journal on Document Analysis and Recognition
    • MM-CCA: Transaction on Multimedia Computing, Communication and Applications
  • Conferences
    • NeurIPS: Neural Information Processing Systems
    • ICML: International Conference on Machine Learning
    • CVPR: IEEE/CVF Conference on Computer Vision and Pattern Recognition
    • ICCV: IEEE/CVF International Conference of Computer Vision
    • ECCV: IEEE/CVF European Conference of Computer Vision
    • WACV: IEEE/CVF Winter Conference on Applications of Computer Vision
    • SciVis: IEEE Scientific Visualization Conference
    • ICIP: IEEE International Conference on Image Processing
    • VCIP: IEEE International Conference Visual Communication Image Process
    • CSNT: IEEE International Conference on Communication Systems and Network Technologies
    • CAIDCD: IEEE International Conference on Computer-Aided Industrial Design and Conceptual Design
    • ACM: Association for Computing Machinery
    • ICDAR: IAPR International Conference on Document Analysis and Recognition
    • ICPR: International Conference on Pattern Recognition
    • ICIR: International Conference on Intelligent Reality
    • IIAI-AAI: International Congress on Advanced Applied Informatics
    • MMM: Multimedia Modeling
    • LREC: International Conference on Language Resources and Evaluation
    • MTA: Multimedia Tools and Applications
    • ICIT: International Conference on Information Technology
    • ICIAP: International Conference on Image Analysis and Processing
    • IJCNN: International Joint Conference on Neural Networks
    • ACPR: IAPR Asian Conference on Pattern Recognition
    • ICGC: IEEE International Conference on Granular Computing
    • CAVW: Computer Animation and Virtual Worlds
    • MM-TA: Multimedia Tools and Applications
    • DICTA: International Conference on Digital Image Computing: Techniques and Applications
    • UIST: ACM Symposium on User Interface Software and Technology
    • EMNLP: ACM Conference on Empirical Methods in Natural Language Processing
    • IIWAS: International Conference on Information Integration and Web-based Applications and Services
    • MDPI-AS: MDPI Applied Science
    • ICMLC: International Conference on Machine Learning and Cybernetics
    • LACLO: Latin American Conference on Learning Technologies
    • ACL: Association for Computational Linguistics
    • ICCC: International Conference on Computational Creativity
    • 3DV: International Conference on 3D Vision
  • Workshops
    • MANPU: IAPR International Workshop on Comics Analysis, Processing and Understanding
    • DAS: IAPR International Workshop on Document Analysis Systems
    • IWAIT: International Workshop on Advanced Image Technology
    • ISPACS: Symposium on Intelligent Signal Processing and Communication Systems
    • SIGCSE: ACM Technical Symposium on Computer Science Education
    • ISM: IEEE International Symposium in Multimedia

Links

🔧 Tools & Repositories

How to Contribute 🚀

You can contribute in two ways:

  1. The easiest is to open an Issue (see an example in issue #1) and we can discuss if there are missing papers, wrong associations or links, or misspelled venues.
  2. The second one is making a pull request with the implemented changes, following the steps:
    1. Fork this repository and clone it locally.
    2. Create a new branch for your changes: git checkout -b feature-name.
    3. Make your changes and commit them: git commit -m 'Description of the changes'.
    4. Push to your fork: git push origin feature-name.
    5. Open a pull request on the original repository by providing a description of your changes.

This project is in constant development, and we welcome contributions to include the latest research papers in the field or report issues 💥💥.

Star History ⭐

Star History Chart

Acknowledge

Many thanks to my co-authors for taking the time to help me with the various refactoring of the survey. Thanks to Beppe Folder for its Awesome Human Visual Attention repo that inspired the ✨style✨ of this repository.

Significant stargazers

Marcello Seri

192 followers · starred Sep 2024