๐ญ Daily Papers (2026-09-22)
Document Understanding
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | Chart-RVR | University of Virginia | Monitorable Chart Reasoning Agents via Verifiable Process Rewards |  | Sep. 2026 |
Chart-RVR โ A reinforcement learning framework that trains chart reasoning agents with verifiable process rewards, decomposing reasoning into auditable structure, evidence-table and derivation blocks instead of answer-only outputs.
Visual Text Generation
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | DuetGen | USTC | Planning and Rendering in Concert: DeepFusion of Autoregressive Layouts and Diffusion for Visual Text Generation | - | Sep. 2026 |
DuetGen โ An autonomous visual text generator that jointly learns autoregressive layout planning and diffusion rendering so that rendering supervision shapes the planner's representations instead of optimizing the two stages separately.
๐ Contents
Overview
A curated, continuously updated reading list of OCR in the era of large language models, covering document parsing and understanding, visual text generation, benchmarks, challenges, and future perspectives, with a focus on research around the past five years (2021โnow).
Scope. This list tracks OCR in the LLM era: work that applies large vision-language or multimodal models to text-rich images and documents (parsing, understanding, benchmarks, and specialized text tasks). It is not a general document-AI list, a generic MLLM list, or a classical OCR-1.0 list; such work appears only when it directly bears on text-rich visual understanding.
A note on evaluation. Most recent systems are released as technical reports with self-reported numbers, private test sets, and inconsistent protocols, so cross-paper scores are rarely comparable in a rigorous sense. We list results as reported and, where known, indicate the evaluation basis. The field still lacks a unified, contamination-resistant, reproducible benchmark, and we see building one as a prerequisite for trustworthy leaderboard claims.
๐ News
- [2026-2-11] ๐ฅ We release an open-source resource to help the community easily track recent OCR research!
Contributing. PRs welcome. One row per model, newest first; please include venue/date, affiliation, and a code or model link.
๐ Emerging Trends
Emerging Trends (2023โ2026)
- End-to-end VLM-based parsing replaces modular OCR pipelines.
- Reinforcement learning for layout and reading order modeling.
- OCR-free document understanding models.
- Scaling down: compact document VLMs under 1B parameters.
- Long-doc OCR is a scaling problem of tokens and consistency, not just context length.
- Structure (layout + logic) is the new accuracy.
- Generative OCR shifts the core risk from โmisrecognitionโ to โhallucinationโ.
- Benchmarks are moving toward executable evaluation.
- Document agents and autonomous reasoning over PDFs.
- Document Agents need recoverability, not one-shot perfection.
๐ Document Parsing
Document parsing focuses on converting visually complex documents into structured, machine-readable representations. In the LLM era, parsing is no longer a pipeline of isolated modules, but increasingly unified within end-to-end VLM architectures.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | PrismAlign | Huawei | PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR | - | Sep. 2026 |
 | WeVisDoc | Tencent | WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing |  | Sep. 2026 |
 | Jina-OCR-v1 | Jina AI | Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards |  | Sep. 2026 |
 | OCR-EDR | Tencent | OCR-EDR: Rendering-Aware Diagnosis and Repair for Closed-Loop OCR Improvement | - | Sep. 2026 |
 | Wayu-Paxa-OCR-Zero | Wayu Research | How Far Can Synthetic Data Take Thai OCR? | - | Sep. 2026 |
 | SCVER | Fudan University | State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models | - | Sep. 2026 |
 | FinixDoc | Ant Group | FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks |  | Aug. 2026 |
 | SmolDocling-KV-Link | IBM | Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images | - | Aug. 2026 |
 | ArmorOCR | Ant Group | ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation |  | Aug. 2026 |
 | NaviDC-OCR | China Telecom AI | NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents | - | Aug. 2026 |
 | TongGuOCR | SCUT | TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents |  | Aug. 2026 |
 | PaDoc | Tsinghua University | PaDoc: Layout-Grounded Parallel Decoding for Document Parsing |  | Aug. 2026 |
 | Logographic Pretraining | QMUL | Logographic Character Visual Pretraining via Semantic-based Contrastive Learning | - | Aug. 2026 |
 | SPIRAL | SEU | Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression |  | Aug. 2026 |
 | DocPO | Tencent | DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards | - | Aug. 2026 |
 | DrawAI | BUPT | DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable |  | Aug. 2026 |
 | LayoutLite | Yuanli Technology | LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR |  | Jul. 2026 |
 | HPD-Parsing | Baidu | HPD-Parsing: Hierarchical Parallel Document Parsing |  | Jul. 2026 |
 | OvisOCR2 | Alibaba | OvisOCR2 Technical Report |  | Jul. 2026 |
 | DocOCR-Eval | University of Melbourne | DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth | - | Jul. 2026 |
 | MonkeyOCRv2 | HUST&Kingsoft | MonkeyOCRv2: A Visual-Text Foundation Model for Document AI |  | Jul. 2026 |
 | Infinity-Parser2 | INF Team | Infinity-Parser2 Technical Report |   | Jul. 2026 |
 | HunyuanOCR-1.5 | Tencent | HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better |  | Jul. 2026 |
 | SAYRE | Alibaba | Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis | - | Jul. 2026 |
 | P-MTP | Baidu | P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling | - | Jun. 2026 |
 | RT-DocLayout | Baidu | RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild | - | Jun. 2026 |
 | Unlimited-OCR | Baidu | Unlimited OCR Works |  | Jun. 2026 |
 | Beaver | Microsoft Research | Building Agent Harnesses for Scientific Curation from Multimodal Sources | - | Jun. 2026 |
 | Agents-K1 | Shanghai AI Laboratory | Agents-K1: Towards Agent-native Knowledge Orchestration |  | Jun. 2026 |
 | PaddleOCR-VL-1.6 | Baidu | PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training |  | Jun. 2026 |
 | PP-OCRv6 | Baidu | PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks |  | Jun. 2026 |
 | StrucTab | IIE, CAS | StrucTab: A Structured Optimization Framework for Table Parsing |  | Jun. 2026 |
 | ExChart | Zhejiang University | Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework | - | Jun. 2026 |
 | MinerU-Popo | Shanghai AI Laboratory & OpenDataLab | MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing |  | May. 2026 |
 | RTPrune | DeepSeek | RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference |  | May. 2026 |
 | ABot-OCR | Alibaba | ABot-OCR Technical Report |  | May. 2026 |
 | BabelDOC | funstory.ai | BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation |  | May. 2026 |
 | Consensus Entropy | Fudan University | Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR |  | May. 2026 |
 | FastOCR | Tsinghua University & JD | FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing | - | May. 2026 |
 | MinerU2.5-Pro | Shanghai AI Laboratory | MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale |  | Apr. 2026 |
 | PixelPrune | OPPO | PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding |  | Apr. 2026 |
 | TexOCR | Yale University & Zhejiang University | TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction |  | Apr. 2026 |
 | Falcon OCR | Falcon Vision Team, TII | Falcon Perception |  | Mar. 2026 |
 | MinerU-Diffusion | Shanghai AI Laboratory | MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding |  | Mar. 2026 |
 | Qianfan-OCR | Baidu | Qianfan-OCR: A Unified End-to-End Model for Document Intelligence |  | Mar. 2026 |
 | dots.mocr | HUST | Multimodal OCR: Parse Anything from Documents |  | Mar. 2026 |
 | FireRed-OCR | Xiaohongshu Inc | FireRed-OCR Technical Report |  | Mar. 2026 |
 | AgenticOCR | Shanghai AI Laboratory & OpenDataLab | AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation |  | Mar. 2026 |
 | PTP | Tencent | Efficient Document Parsing via Parallel Token Prediction | - | Mar. 2026 |
 | Logics-Parsing-Omni | Alibaba | Logics-Parsing-Omni Technical Report |  | Mar. 2026 |
๐ See full list at Document-Parsing.md
๐ Document Understanding
Document understanding extends beyond structural parsing to semantic comprehension and reasoning over visually rich documents.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | Chart-RVR | University of Virginia | Monitorable Chart Reasoning Agents via Verifiable Process Rewards |  | Sep. 2026 |
 | ViSAR | INSA Lyon | ViSAR: Training-Free Adaptive-k Retrieval for Visual Document Question Answering | - | Sep. 2026 |
 | DocIntent | SCUT | DocIntent: Answerability-Guided Agentic Restoration for Real-World Document Visual Question Answering | - | Sep. 2026 |
 | Doc-REFRAG | Zhejiang University | Doc-REFRAG: Rethinking Multimodal Document Retrieval-Augmented Generation |  | Sep. 2026 |
 | AWM | University of Oslo | AWM: Answerable Working Memory for Long-Document VQA Agents |  | Aug. 2026 |
 | SAGE | Fudan University | SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding | - | Aug. 2026 |
 | Q-Guide | Amazon | Question-Guided Evidence Acquisition for Multimodal Visual Question Answering | - | Aug. 2026 |
 | DocClaw | NTU | DocClaw: A Unified Agentic System for Intelligent Document Processing |  | Aug. 2026 |
 | ConceptFormer | Northeastern University | ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval |  | Aug. 2026 |
 | Hyper-M2RAG | Hangzhou Dianzi University | Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement |  | Aug. 2026 |
 | Trident | Emory University | What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering | - | Aug. 2026 |
 | D2-ScaleAgent | Zhejiang University | D2-ScaleAgent: Dual-Dimensional Scaling for Long Document Understanding | - | Aug. 2026 |
 | SEER | UT Austin | SEER: Long-Context Reasoning via Selective Visual-Text Compression |  | Aug. 2026 |
 | HAM-RAG | HKUST | HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation |  | Aug. 2026 |
 | DRUF | Shenzhen MSU-BIT University | Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs |  | Aug. 2026 |
 | DistilVDR | Aalto University | DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation |  | Aug. 2026 |
 | InSight-doc | HKUST | InSight-doc: Agentic Visual Perception for Long-Document Understanding |  | Aug. 2026 |
 | DocAtlas | Wuhan University / Microsoft | DocAtlas: Long-Document Understanding as Mutable-State Interaction | - | Aug. 2026 |
 | DocMemo | HIT, Shenzhen | DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding |  | Aug. 2026 |
 | ECF | Beihang University | Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models? |  | Aug. 2026 |
 | ADOPD 2026 | Georgia Tech | Thinking with Anchors: Grounded and Efficient Document Reasoning |  | Aug. 2026 |
 | VTS | MBZUAI | When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning | - | Aug. 2026 |
 | Q-CueGraph | The University of Tokyo | Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning | - | Aug. 2026 |
 | DocTrace | Baidu | DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning | - | Aug. 2026 |
 | CURV | William & Mary | CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning | - | Aug. 2026 |
 | RAGOCR | Peking University | RAGOCR: Optical Compression of Retrieval-Augmented Text via Visual Representation | - | Aug. 2026 |
 | VaRS-Doc | SJTU | VaRS-Doc: Interpretation-Aware Variant Representations via Latent Self-Probing for Visual Document Retrieval |  | Aug. 2026 |
 | ET-Prune | SJTU | ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs |  | Aug. 2026 |
 | HierDoc | USTC | HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering | - | Aug. 2026 |
 | DualG-MRAG | BUAA | DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation | - | Jul. 2026 |
 | MMLDSum-LLM | OPPO | MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware | - | Jul. 2026 |
 | TAP-RAG | Tianjin University | TAP-RAG: Task-Aware Policy Control for Long-Document Multimodal Question Answering |  | Jul. 2026 |
 | Perception-RFT | Quantiphi | Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment | - | Jul. 2026 |
 | CMDR | NTT | CMDR: Contextual Multimodal Document Retrieval |  | Jul. 2026 |
 | HiEvi-RAG | USTC | Hierarchical Evidence-Driven Reasoning for Long Document Understanding | - | Jul. 2026 |
 | MultAttnAttrib | โ | MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering | - | Jul. 2026 |
 | OracleAnalyser | NUDT | OracleAnalyser: Analysing Implicit Semantics of Oracle Bone Scripts through MLLMs with Post-training | - | Jun. 2026 |
 | DocArena | Adobe Research | DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents | - | Jun. 2026 |
 | ViTexQA | Meituan | ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering |  | Jun. 2026 |
 | PreciseDoc | Tsinghua University | An LMM for Precisely Grounding Elements in Documents | - | Jun. 2026 |
 | LightSTAR | SJTU | LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement |  | Jun. 2026 |
 | SciLens | HKUST | SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding | - | Jun. 2026 |
 | SAFE-Cascade | Walmart | SAFE-Cascade: Cost-Adaptive Vision-Language Routing for Chart Question Answering | - | Jun. 2026 |
 | UMG-RAG | Purdue University | Uncertainty-Aware Hybrid Retrieval for Long-Document RAG | - | Jun. 2026 |
 | MAGE-RAG | BIT | MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA |  | Jun. 2026 |
 | MINARD | University of Maryland | Helping Figures Tell their Story! Paper-Grounded Video Generation Explaining Complex Scientific Figures | - | Jun. 2026 |
 | MM-BizRAG | JPMorgan Chase & Co. | MM-BizRAG: Rethinking Multimodal Retrieval-Augmented Generation for General Purpose Enterprise Q&A | - | Jun. 2026 |
 | KG4VD | National Taiwan University | Multimodal Graph RAG for Long-range Visually Rich Document Understanding |  | Jun. 2026 |
 | ZipRerank | Magellan Technology Research Institute | Very Efficient Listwise Multimodal Reranking for Long Documents |  | May. 2026 |
 | GeoSym127K | SenseTime Research & CUHK Shenzhen | GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning |  | May. 2026 |
๐ See full list at Document-Understanding.md
๐ Visual Text Generation
Visual Text Generation focuses on generating or editing legible, visually harmonious, and semantically consistent text within images, serving as the creative inverse of OCR. In the LLM era, it is no longer a task reliant on specialized modules, but is emerging as a foundational skill for general-purpose generative models.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | DuetGen | USTC | Planning and Rendering in Concert: DeepFusion of Autoregressive Layouts and Diffusion for Visual Text Generation | - | Sep. 2026 |
 | LoGAN | Netflix | LoGAN: Multilingual Font Localization with Generative Agents | - | Sep. 2026 |
 | GlyphAnchor | Fudan University | GlyphAnchor: Enhancing Visual Text Rendering via Position-Anchored Glyph Priors | - | Sep. 2026 |
 | TextRefine | Kuaishou | TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters | - | Aug. 2026 |
 | PosterText | Wuhan University | PosterText: Towards Unified Visual Text Generation and Editing for E-commerce Poster | - | Aug. 2026 |
 | TransAnyText | Wuhan University | TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation | - | Aug. 2026 |
 | onoff | GIST | Bridging Online and Offline Handwriting via Differentiable Physical Rendering |  | Aug. 2026 |
 | PosterMELD | Tsinghua University | PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs |  | Aug. 2026 |
 | InnoText | SYSU | InnoText: A Unified Model for Visual Text Generation and Editing | - | Jul. 2026 |
 | Boogu-Image-0.1 | Boogu Team,Huawei | Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation |  | Jul. 2026 |
 | SciForma | Microsoft & Peking University | SciForma: Structure-Faithful Generation of Scientific Diagrams |  | Jul. 2026 |
 | VecFontLLM | Fuzhou University & Peking University | VecFontLLM: Anchor-Guided Direct Synthesis of Chinese Vector Fonts | - | Jul. 2026 |
 | ArtChart | Ant Group | ArtChart: A Benchmark for Faithful Artistic Chart Generation with Integrated Text Rendering | - | Jul. 2026 |
 | DataEvolver | CSU | DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation |  | Jul. 2026 |
 | Qwen-Image-2.0-RL | Alibaba | Qwen-Image-2.0-RL Technical Report | - | Jun. 2026 |
 | UniTranslator | IIE CAS | UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation |  | Jun. 2026 |
 | SteerVTE | ByteDance & Peking University | SteerVTE: Seamless Video Text Editing with Style and Glyph Control | - | Jun. 2026 |
 | DiffMath | SCUT & Huawei | DiffMath: Symbol- and Graph-Aware Latent Diffusion Transformer for Handwritten Mathematical Expression Generation |  | Jun. 2026 |
 | PhyDrawGen | University of Dhaka | PhyDrawGen: Physically Grounded Diagram Generation from Natural Language | - | Jun. 2026 |
 | NIV | Reichman University | NIV: Neural Axis Variations for Variable Font Generation |  | Jun. 2026 |
 | Qwen-Image-2.0 | Alibaba Group | Qwen-Image-2.0 Technical Report |  | May. 2026 |
 | MangaFlow | The University of Tokyo | MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation | - | May. 2026 |
 | Wan-Image | Alibaba Group | Wan-Image: Pushing the Boundaries of Generative Visual Intelligence | - | Apr. 2026 |
 | PosterIQ | PolyU | PosterIQ: A Design Perspective Benchmark for Poster Understanding and Generation |  | Mar. 2026 |
 | EfficientPosterGen | Tsinghua University | EfficientPosterGen: Semantic-aware Efficient Poster Generation via Token Compression and Accurate Violation Detection |  | Mar. 2026 |
 | TextFlow | NJUST | Towards Training-Free Scene Text Editing |  | Mar. 2026 |
 | GlyphPrinter | Fudan University | GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering |  | Mar. 2026 |
 | CTRL-S | SJTU | Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning | - | Mar. 2026 |
 | LaDe | Adobe Research | LaDe: Unified Multi-Layered Graphic Media Generation and Decomposition | - | Mar. 2026 |
 | EchoGen | USTC | EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding | - | Mar. 2026 |
 | WebVR | StepFun | WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics | - | Mar. 2026 |
 | AutoFigure-Edit | Westlake University | AutoFigure-Edit: Generating Editable Scientific Illustration |  | Mar. 2026 |
 | GlyphBanana | SJTU & Xiaohongshu Inc. | GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows |  | Mar. 2026 |
 | InnoAds-Composer | JD.com | InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster Generation | - | Mar. 2026 |
 | Seeing is Improving | USTC | Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement |  | Mar. 2026 |
 | PosterOmni | HKUST(GZ) | PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback |  | Feb. 2026 |
 | TextPecker | HUST | TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering |  | Feb. 2026 |
 | SSPT | Tongji University | Space Syntax-guided Post-training for Residential Floor Plan Generation | - | Feb. 2026 |
 | ChatUMM | Tsinghua University & Tencent Hunyuan | ChatUMM: Robust Context Tracking for Conversational Interleaved Generation | - | Feb. 2026 |
 | PosterVerse | SCUT | PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable Typography |  | Jan. 2026 |
| - | GPT-Image-1.5 | OpenAI | GPT-Image-1.5 | --- | Dec. 2025 |
| - | Gemini 3 Pro Image | Google DeepMind | Gemini 3 Pro Image (Nano Banana Pro) | --- | Nov. 2025 |
| - | Gemini 2.5 Flash Image | Google DeepMind | Gemini 2.5 Flash Image (Nano Banana) | --- | Oct. 2025 |
 | Qwen-Image | Qwen Team | Qwen-Image Technical Report |  | Sep. 2025 |
 | Seedream 4.0 | ByteDance Seed | Seedream 4.0: Toward Next-generation Multimodal Image Generation | --- | Aug. 2025 |
 | Postergen | Stony Brook University | Postergen: Aesthetic-aware paper-to-poster generation via multi-agent llms |  | Aug. 2025 |
 | UniGlyph | Tsinghua University | UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis | --- | Jul. 2025 |
 | X-Omni | Tencent Hunyuan X | X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again |  | Jul. 2025 |
 | DreamPoster | Intelligent Creation Lab, ByteDance | DreamPoster: A Unified Framework for Image-Conditioned Generative Poster Design | --- | Jul. 2025 |
 | PosterCraft | HKUST(GZ) | PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework |  | Jun. 2025 |
๐ See full list at Visual-Text-Generation.md
๐ Specialized Model
Beyond general document parsing and understanding, traditional OCR research remains essential for specialized visual-text structures, restoration, temporal analysis, and forensic reliability.
๐ Document Dewarping
Document dewarping restores photographed or scanned pages to a geometrically rectified and readable form.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | BookNet | HFUT | BookNet: Dual-Page Book Image Rectification via Cross-Page Attention | - | Jan. 2026 |
 | TADoc | CAS | TADoc: Robust Time-Aware Document Image Dewarping | - | Aug. 2025 |
 | DocDewarpHV | HIT | Dual Dimensions Geometric Representation Learning Based Document Dewarping |  | Jul. 2025 |
 | DocMatcher | FZI & KIT | DocMatcher: Document Image Dewarping via Structural and Textual Line Matching |  | Mar. 2025 |
 | - | Shuya Branch of the Ivanovo State University | Efficient Document Image Dewarping via Hybrid Deep Learning and Cubic Polynomial Geometry Restoration |  | Jan. 2025 |
 | DocScanner | USTC | DocScanner: Robust Document Image Rectification with Progressive Learning |  | Jan. 2025 |
 | DocRes | SCUT | DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks |  | Jun. 2024 |
 | DocNLC | SCUT | DocNLC: A Document Image Enhancement Framework with Normalized and Latent Contrastive Representation for Multiple Degradations |  | Feb. 2024 |
 | DocTr++ | USTC | Deep Unrestricted Document Image Rectification | - | 2024 |
 | LA-DocFlatten | CAS | Layout-aware Single-image Document Flattening |  | 2024 |
 | UVDoc | ETH Zurich | UVDoc: Neural Grid-based Document Unwarping |  | Oct. 2023 |
 | Foreground and Text-lines Aware Model | HIT Shenzhen, China | Foreground and text-lines aware document image rectification |  | Jun. 2023 |
 | DocMAE | USTC & iFLYTEK | DocMAE: Document Image Rectification via Self-supervised Representation Learning | - | 2023 |
 | Marior | SCUT & IntSig | Marior: Margin Removal and Iterative Content Rectification for Document Dewarping in the Wild |  | Oct. 2022 |
๐ Physical Structure Analysis
Physical structure analysis identifies document regions, layouts, and spatial relationships before semantic reasoning.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | PorTEXTO | NOVA School of Science and Technology | PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction | - | Jun. 2026 |
 | - | Insiders Technologies GmbH | Bounding Box Label Propagation for Re-Annotation of Document Layout Analysis Datasets | - | Jun. 2026 |
 | IndustryBench-MIPU | Multimodal and Industrial AI Team, Alibaba | IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products |  | Jun. 2026 |
 | ERN-Net | Tamkang University | ERN-Net : Evolving Reason Node-Net for Document Binarization | - | Jun. 2026 |
 | HiLEx | IEM Kolkata | HiLEx: Image-Based Hierarchical Layout Extraction from Question Papers |  | Mar. 2025 |
 | DCEM-ViT | Sharda University, Greater Noida, India | Devanagari character encoded mix-merge vision transformer for robust document layout analysis | --- | Mar. 2025 |
 | Efficient Additive Attention DLA | Technical University of Kaiserslautern, Germany | Efficient Additive Attention for Transformer-based Semi-supervised Document Layout Analysis | --- | Feb. 2025 |
 | DocSemi | Technical University of Kaiserslautern, Germany | DocSemi: Efficient Document Layout Analysis with Guided Queries | --- | Feb. 2025 |
 | FS-QCSNet | China University of Mining and Technology | Few-Shot Quaternion-valued Correlation Squeeze Network for Document Image Layout Segmentation | --- | Jan. 2025 |
 | LayoutDETR | Salesforce Research | LayoutDETR: Detection Transformer Is a Good Multimodal Layout Designer | - | Sep. 2024 |
 | LayoutLLM | Alibaba Group | LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding |  | Jun. 2024 |
 | RoDLA | KIT & Univ. of Oxford | RoDLA: Benchmarking the Robustness of Document Layout Analysis Models |  | Jun. 2024 |
 | DocLLM | JPMorgan Chase & Co. | DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding |  | May 2024 |
 | SemiDocSeg | Computer Vision Center, Barcelona, Spain | SemiDocSeg: Harnessing Semi-Supervised Learning for Document Layout Analysis | --- | Mar. 2024 |
 | VGT | Alibaba Group | Vision Grid Transformer for Document Layout Analysis |  | Oct. 2023 |
 | GeoLayoutLM | Alibaba Group | GeoLayoutLM: Geometric Pre-training for Visual Information Extraction |  | Jun. 2023 |
 | M6Doc | SCUT | M6Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout Analysis |  | Jun. 2023 |
 | HRDoc | USTC & iFLYTEK | HRDoc: Dataset and Baseline Method Toward Hierarchical Reconstruction of Document Structures |  | Feb. 2023 |
 | LayoutLMv3 | Microsoft | LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking |  | Oct. 2022 |
๐ Reading Order Prediction
Reading order prediction recovers the intended sequence of text and visual elements in complex page layouts.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | Orli | University of Konstanz | End-to-End Text Line Detection and Ordering |  | Jun. 2026 |
 | FocalOrder | Unisound AI Technology Co. Ltd. | FocalOrder: Focal Preference Optimization for Reading Order Detection | - | Jan. 2026 |
 | ROAP | UESTC | ROAP: A Reading-Order and Attention-Prior Pipeline for Optimizing Layout Transformers in Key Information Extraction |  | Jan. 2026 |
 | XY-cut++ | Tianjin University | XY-Cut++: Advanced Layout Ordering via Hierarchical Mask Mechanism on a Novel Benchmark | - | Apr. 2025 |
 | UniHDSA | USTC & Microsoft Research Asia | UniHDSA: A Unified Relation Prediction Approach for Hierarchical Document Structure Analysis |  | 2025 |
 | HanDoc-OrderOCR | NCCU & Academia Sinica | Reading between the Lines: Image-Based Order Detection in OCR for Chinese Historical Documents | - | Feb. 2024 |
 | Detect-Order-Construct | USTC & Microsoft Research Asia | Detect-Order-Construct: A Tree Construction Based Approach for Hierarchical Document Structure Analysis |  | 2024 |
 | Reading Order Matters | Fudan University | Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path Prediction |  | Dec. 2023 |
 | LayoutLMv3 | SYSU | LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking |  | Oct. 2022 |
๐ Mathematical Expression Recognition
Mathematical expression recognition converts printed or handwritten formulas into structured symbolic representations.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | - | University of Alicante | Direct content-based retrieval from music scores images | - | May. 2026 |
 | MusicSynth | - | MusicSynth: An Automated Pipeline for Generating Violin Fingerboard Animations from Sheet Music Using Optical Music Recognition | - | May. 2026 |
 | Transcoda | Heinrich Heine University Dรผsseldorf | Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training |  | May. 2026 |
 | From Image to Music Language | FindLab | From Image to Music Language: A Two-Stage Structure Decoding Approach for Complex Polyphonic OMR | - | Apr. 2026 |
 | UniMERNet | OpenDataLab & Shanghai AI Lab | UniMERNet: A Universal Network for Real-World Mathematical Expression Recognition |  | Jun. 2025 |
 | CDM | OpenDataLab | Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching |  | Jun. 2025 |
 | SSAN | Inner Mongolia University | SSAN: A Symbol Spatial-Aware Network for Handwritten Mathematical Expression Recognition |  | Apr. 2025 |
 | TAMER | Peking University | TAMER: Tree-Aware Transformer for Handwritten Mathematical Expression Recognition |  | Apr. 2025 |
 | VLPG | UCAS | Visionโlanguage pre-training for graph-based handwritten mathematical expression recognition |  | Jan. 2025 |
 | BAT | USTC & iFLYTEK | Bidirectional Trained Tree-Structured Decoder for Handwritten Mathematical Expression Recognition |  | 2025 |
 | LattE | Purdue University | LattE: Improving LaTeX Recognition with Iterative Refinement |  | Sep. 2024 |
 | - | TUAT & VGU & RIT & Nantes Univ. | A Survey on Handwritten Mathematical Expression Recognition: The Rise of Encoder-Decoder and GNN Models | - | Sep. 2024 |
 | NAMER | USTC & iFLYTEK Research | NAMER: Non-Autoregressive Modeling for Handwritten Mathematical Expression Recognition | - | Sep. 2024 |
 | HMEG | Hangzhou Dianzi Univ. | Generating Handwritten Mathematical Expressions from Symbol Graphs: An End-to-End Pipeline | - | Jun. 2024 |
 | BPD | SCUT | A Tree-Based Model with Branch Parallel Decoding for Handwritten Mathematical Expression Recognition | - | 2024 |
 | SAN (Syntax-Aware Net) | Tomorrow Advancing Life | Syntax-Aware Network for Handwritten Mathematical Expression Recognition |  | Jun. 2022 |
๐ Table & Chart Understanding
Table and chart understanding covers structure recognition, data extraction, grounding, retrieval, and visual reasoning over structured graphics.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | - | Teamreboott Inc. | Rethinking the Pointer Loss in Table Structure Recognition: Geometry-Aware Pointer Loss for Spatial Locality |  | Jun. 2026 |
 | - | Preferred Networks, Inc. | Revisiting Structural Dependency in Autoregressive Multi-Task Table Recognition via Order-Independent Cell-Level Representations | - | Jun. 2026 |
 | ChartLens | Shandong University | ChartLens: A Dual-Branch Framework for Chart Data Correction and Factual Summary Refinement |  | Jun. 2026 |
 | POTATR | Kensho Technologies (S&P Global) | POTATR: A Lightweight Image-to-Graph Model for Page-Level Table Extraction | - | Jun. 2026 |
 | ChartReLA | University of Science, Ho Chi Minh city, Vietnam | ChartReLA: A compact vision-language model for comprehensivechart reasoning via relationship modeling |  | Jan. 2026 |
 | Table-R1 | XJTU-Liverpool University | Can GRPO Boost Complex Multimodal Table Understanding? | - | Dec. 2025 |
 | TinyChart | Alibaba & Tsinghua | TinyChart: Efficient Chart Understanding with Program-of-Thoughts Learning and Visual Token Merging | - | Nov. 2024 |
 | ReachQA | Fudan University | Distill Visual Chart Reasoning Ability from LLMs to MLLMs |  | Oct. 2024 |
 | OmniParser | Alibaba Group | OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition |  | Jun. 2024 |
 | DePlot | Google | DePlot: One-shot visual language reasoning by plot-to-table translation |  | Jul. 2023 |
 | TableVLM | Fudan University | TableVLM: Multi-modal Pre-training for Table Structure Recognition | - | Jul. 2023 |
 | VAST | Huawei | Improving Table Structure Recognition with Visual-Alignment Sequential Coordinate Modeling | - | Jun. 2023 |
 | LORE | Alibaba Group | LORE: Logical Location Regression Network for Table Structure Recognition |  | Feb. 2023 |
 | ChartQA | York University | ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning |  | Jul. 2022 |
 | LGPMA | Hikvision | LGPMA: Complicated Table Structure Recognition with Local and Global Pyramid Mask Alignment |  | Sep. 2021 |
๐ Scene Text Understanding
Scene text understanding unifies text detection, recognition, and spotting in natural images and related real-world settings.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | Next-Generation Parallel Decoder for LPDR | NUTECH | Next-Generation Parallel Decoder for LPDR: Architectural Optimization and Class-Balanced GAN-Augmentation | - | Jun. 2026 |
 | - | Polytechnic University of Turin | Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting | - | May. 2026 |
 | MNSP | Nankai University | Masked Next-Scale Prediction for Self-supervised Scene Text Recognition |  | May. 2026 |
 | - | Afeka Academic College of Engineering | Mapping License Plate Recoverability Under Extreme Viewing Angles for Oppor-tunistic Urban Sensing | - | Apr. 2026 |
 | DPText-DETR | Wuhan Univ. | DPText-DETR: Towards Better Scene Text Detection with Dynamic Points in Transformer |  | Feb. 2023 |
 | DeepSolo | Wuhan University | DeepSolo: Let Transformer Decoder with Explicit Points Solo for Text Spotting |  | Nov. 2022 |
 | ABINet++ | USTC | ABINet++: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Spotting |  | Nov. 2022 |
 | SwinTextSpotter | SCUT | SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition |  | Mar. 2022 |
 | ABCNet v2 | SCUT | ABCNet v2: Adaptive Bezier-Curve Network for Real-time End-to-end Text Spotting |  | May. 2021 |
 | PAN++ | Nanjing University | PAN++: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped Text |  | May. 2021 |
 | MANGO | Hikvision Research Institute | MANGO: A Mask Attention Guided One-Stage Scene Text Spotter |  | Dec. 2020 |
 | Mask TextSpotter v3 | HUST | Mask TextSpotter v3: Segmentation Proposal Network for Robust Scene Text Spotting |  | Jul. 2020 |
 | Text perceptron | Hikvision Research Institute | Text perceptron: Towards end-to-end arbitrary-shaped text spotting |  | Apr. 2020 |
 | ABCNet | SCUT | ABCNet: Real-time Scene Text Spotting with Adaptive Bezier-Curve Network |  | Feb. 2020 |
๐ Text Removal & Editing
Text removal and editing modifies textual content in images while preserving visual consistency and surrounding appearance.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | ProductConsistency | Fractal Analytics | ProductConsistency: Improving Product Identity Preservation in Instruction-Based Image Editing via SFT and RL |  | Jun. 2026 |
 | - | Fudan University | Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification | - | Jun. 2026 |
 | UniDDT | Nanjing University | UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer | - | Jun. 2026 |
 | TextWand | Peking University | TextWand: A Unified Framework for Scene Text Editing | - | Jun. 2026 |
 | TMIM | USTC | Leveraging Text Localization for Scene Text Removal via Text-Aware Masked Image Modeling |  | Sep. 2024 |
 | PowerPaint | Tsinghua & Shanghai AI Lab | PowerPaint: A Task is Worth One Word โ Learning with Task Prompts for High-Quality Versatile Image Inpainting |  | Sep. 2024 |
 | MagicEraser | Huaweiย & SIAT, CAS | MagicEraser: Erasing Any Objects via Semantics-Aware Control | - | Sep. 2024 |
 | TurboEdit | Adobe | TurboEdit: Instant Text-based Image Editing |  | Sep. 2024 |
 | SPDInv | HKUST | Source Prompt Disentangled Inversion for Boosting Image Editability |  | Sep. 2024 |
 | DARLING | USTC | Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing | - | Jun. 2024 |
 | SceneTextGen | Meta & Rutgers | Layout-Agnostic Scene Text Image Synthesis with Diffusion Models | - | Jun. 2024 |
 | - | KAIST | Prompt Augmentation for Self-supervised Text-guided Image Manipulation | - | Jun. 2024 |
 | ViTEraser | SCUT | ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining |  | Feb. 2024 |
 | FETNet | Kyushu Univ. | FETNet: Feature Erasing and Transferring Network for Scene Text Removal |  | 2023 |
๐ Text Image Super-Resolution
Text image super-resolution restores low-resolution text regions while preserving character identity and legibility.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | FDF | Anhui University | Frequency Decoupled Framework for Screen Content Image Super-Resolution | - | Jun. 2026 |
 | DTG-Restore | Virginia Tech | DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution | - | Jun. 2026 |
 | Everything at Every Scale | MIT | Everything at Every Scale: Scale-Invariant Diffusion with Continuous Super-Resolution | - | May. 2026 |
 | PRISM | SJTU | PRISM: Prior Rectification and Uncertainty-Aware Structure Modeling for Diffusion-Based Text Image Super-Resolution |  | May. 2026 |
 | TextDiff | BUPT | TextDiff: Mask-Guided Residual Diffusion Models for Scene Text Image Super-Resolution | - | 2025 |
 | PEAN | Southeast Univ. | PEAN: A Diffusion-Based Prior-Enhanced Attention Network for Scene Text Image Super-Resolution |  | Oct. 2024 |
 | DCDM | IIT Roorkee | DCDM: Diffusion-Conditioned-Diffusion Model for Scene Text Image Super-Resolution |  | Sep. 2024 |
 | DiffTSR | BIT & SenseTime | Diffusion-based Blind Text Image Super-Resolution |  | Jun. 2024 |
 | SGENet | Fudan Univ. & Videt Tech. | SGENet: Efficient Scene Text Image Super-Resolution with Semantic Guidance | - | Apr. 2024 |
 | STIRER | Tongji Univ. & Fudan Univ. | STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and Recognition |  | Oct. 2023 |
 | DocDiff | BUPT | DocDiff: Document Enhancement via Residual Diffusion Models |  | Oct. 2023 |
 | - | HIT & NTU | Learning Generative Structure Prior for Blind Text Image Super-Resolution |  | Jun. 2023 |
 | DPMN | Southeast Univ. | DPMN: Improving Scene Text Image Super-Resolution via Dual Prior Modulation Network |  | Feb. 2023 |
 | TPGSR | Hong Kong PolyU | Text Prior Guided Scene Text Image Super-Resolution |  | 2023 |
๐ Handwritten Document Analysis
Handwritten document analysis covers handwriting recognition, writer-related analysis, signatures, and document-level handwritten content.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | - | LTU | Performance Gap Analysis between Latin and Arabic Scripts HTR | - | Jun. 2026 |
 | Prototypical Signature Approach | ETS / LIVIA Lab, Universite du Quebec | A Prototypical Signature Approach for Writer-Independent Offline Signature Verification |  | Jun. 2026 |
 | Stringalign | - | Stringalign: Moving beyond summary statistics with a transparent Unicode-aware tool for evaluating automatic transcription models | - | Jun. 2026 |
 | - | Offenburg UAS | Intelligent Character Recognition of Handwritten Forms with Deep Neural Networks | - | Jun. 2026 |
 | MetaWriter | Concordia Univ. & Mila | MetaWriter: Personalized Handwritten Text Recognition Using Meta-Learned Prompt Tuning | - | Jun. 2025 |
 | - | Univ. of Alicante | On the Generalization of Handwritten Text Recognition Models | - | Jun. 2025 |
 | DetailSemNet | NYCU, Taiwan | DetailSemNet: Elevating Signature Verification through Detail-Semantic Integration | - | Sep. 2024 |
 | TransOSV | Xi'an Jiaotong Univ. | TransOSV: Offline Signature Verification with Transformers | - | 2024 |
๐ Video Text Analysis
Video text analysis studies temporal text detection, recognition, spotting, tracking, and reasoning across video frames.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | TraRA | RIKEN | TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance |  | Jun. 2026 |
 | Beyond Detection | Nankai University | Beyond Detection: A Structure-Aware Framework for Scene Text Tracking | - | May. 2026 |
 | VTAgent | Wuhan University | VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA | - | May. 2026 |
 | VimTS | HUST | VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-Domain Generalization |  | Apr. 2025 |
 | GoMatching | Wuhan Univ. & NTU | GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching |  | Dec. 2024 |
 | TransDeTR | Zhejiang Univ. & BUPT | TransDeTR: End-to-End Video Text Spotting with Transformer |  | 2024 |
 | BiRViT-1K | CASIA | Video Text Detection With Robust Feature Representation | - | 2024 |
๐ Historical Document Analysis
Historical document analysis addresses OCR, restoration, structure recovery, and interpretation for manuscripts and archival materials.
| Venue | Name | Primary affiliation | Title | GitHub | Date |
|---|
 | AlphaOracle | HUST | AlphaOracle: Oracle bone script decipherment via human-workflow-inspired deep learning |  | Jul. 2026 |
 | Urdu Katib Handwritten Dataset | University of Gujrat, Pakistan | Urdu Katib Handwritten Dataset: A Historical Document Dataset for Offline Urdu Handwritten Text Recognition with CRNN-Based Baseline Evaluation | - | Jun. 2026 |
 | - | UMRE, Italy | A Text Recognition Dataset from Sahidic Coptic Ancient Manuscripts | - | Jun. 2026 |
 | - | Charles University | Optical Music Recognition for Real-World Manuscripts with Synthetic Data | - | Jun. 2026 |
 | - | LIGM | Leveraging Morphology for Historical Script Metrological Analysis | - | Jun. 2026 |
 | KaiRacters | TU Wien | KaiRacters: Character-Level-Based Writer Retrieval for Greek Papyri | - | Dec. 2024 |
 | - | Univ. Modena e Reggio Emilia | Binarizing Documents by Leveraging both Space and Frequency | - | Aug. 2024 |
 | CATMuS Medieval | PSL Univ. & ENC | CATMuS Medieval: A Multilingual Large-Scale Cross-Century Dataset in Latin Script for Handwritten Text Recognition | - | Aug. 2024 |
 | DELINE8K | Brigham Young Univ. | DELINE8K: A Synthetic Data Pipeline for Semantic Segmentation of Historical Documents |  | Aug. 2024 |
 | - | Univ. Rennes, IRISA | Training Transformer Architectures on Few Annotated Data: Application to Historical Handwritten Text Recognition | - | 2024 |
 | ColDBin | DFKI | ColDBin: Cold Diffusion for Document Image Binarization | - | Aug. 2023 |
 | - | Univ. of Sousse | Historical Document Image Segmentation Combining Deep Learning and Gabor Features | - | Aug. 2023 |
 | - | TU Wien | Feature Mixing for Writer Retrieval and Identification on Papyri Fragments |  | Aug. 2023 |
๐ Tampered Text Detection & Forensics
Tampered text detection and forensics studies document authenticity, manipulation localization, forgery detection, and evidence-grounded verification.
| Venue | Name | Primary affiliation | Title | GitHub