A Survey on Medical Report Generation: From Deep Neural Networks to Large Language Models
32
5 commits
updated Feb 23, 2024
This repository contains a list of papers, codes, datasets in medical report generation (MRG) field. If you found any error, please don't hesitate to open an issue or pull request.
Given a radiology image, the main objective of medical report generation (MRG) is to generate a descriptive medical report as shown in Figure 1. Current studies leverage medical reports written by professional radiologists as the reference, whose output is expected to be as close to as possible. Formally, they adhere to standard optimization procedures by employing the cross-entropy loss to compare the generated report against the gold standard report.

Figure 1: Two representative cases in MRG, which consist of chest radiology images with their corresponding radiology reports, respectively. Formally, the goal of MRG is to generate the 'Findings' content from one or multiple radiology images.
The most widely used datasets are Indiana University Chest X-ray (IU X-Ray) and MIMIC Chest X-ray (MIMIC-CXR).
IU X-Ray is a widely-used benchmark in MRG systems, which was introduced by Indiana University. It comprises 7,470 chest X-ray images associated with 3,955 radiology reports. Existing MRG studies mostly follow the data splits proposed by Chen et al., which first exclude the samples without the ``Finding'' section, and partition IU X-Ray into the train, validation, and test sets with the ratio of 70%, 10%, and 20%, respectively.
MIMIC-CXR is a recently released large-scale benchmark dataset, which is provided by the Beth Israel Deaconess, USA. It includes 377,110 chest X-ray images and 227,835 reports. Formally, the official data splits are widely adopted. Thus, there are 368,960/2,991/5,159 cases for train/validation/test.
To calculate the performance of MRG, the natural language generation (NLG) metrics, i.e., BLEU-n, METEOR, ROUGE-n, and CIDEr, are widely used. These metrics measure the match between the generated reports and reference reports annotated by professional radiologists. In detail, NLG Metrics are utilized to measure the descriptive accuracy of predicted reports.
bilingual evaluation understudy (BLEU-n) is initially introduced for machine translation, which measures the n-gram precision of generated tokens. BLEU-n is usually employed in the evaluation of MRG approaches, with n ranging from 1 up to 4. This metric assesses the accuracy and coherence of the generated reports to a certain extent.
metric for evaluation of translation with explicit ordering (METEOR) is initially proposed for machine translation, which computes the recall of matching uni-grams from tokens in produced and gold standard reports according to their exact stemmed form and meaning.
recall-oriented understudy for gisting evaluation (ROUGE-L and METEOR) is initially designed for summarization, which measures the similarity between the generated and gold standard report based on their longest common subsequence (LCS) tokens.
Consensus-based image description evaluation (CIDEr) is initially designed to evaluate the quality of generated descriptions for natural images. In MRG systems, CIDEr evaluates models by rewarding topic-specific terms (terminologies in MRG) and penalizing overly frequent terms.
However, existing NLG evaluation metrics are not tailored to evaluate the accurate reporting of abnormalities in the image, which is the core value and urgent problem of MRG. Thus, additional clinical efficacy (CE) metrics are proposed to specifically measure the correctness of descriptions of clinical abnormalities. CE metrics are widely employed to capture and evaluate clinical correctness of predicted reports.
To calculate CE metrics, medical labelers, i.e., CheXpert are utilized to annotate tokens in both generated report and the gold standard one across 14 categories of diseases and support devices, producing the Precision, Recall, and F1 scores.
In addition to above automatic metrics, the human evaluation is also conducted in Liu et al, Liu et al etc. In these works, they invite professional radiologists to rate the quality of generated reports from faithfulness and comprehensiveness perspectives. However, the manual evaluation is both time-consuming and costly given large amount of reports and different radiologists many have conflicted opinion in labeling.

Figure 2: Three categories for MRG, including (a) Data-driven Encoder-Decoder based; (b) Medical Knowledge Enhanced, and (c) Large Language Model based frameworks.
On the Automatic Generation of Medical Imaging Reports [paper], [code]
Hybrid Retrieval-Generation Reinforced Agent for Medical Image Report Generation [paper]
Show, Describe and Conclude: On Exploiting the Structure Information of Chest X-ray Reports [paper]
Competence-based Multimodal Curriculum Learning for Medical Report Generation [paper]
Automatic Radiology Report Generation by Learning with Increasingly Hard Negatives [paper]
Contrastive Attention for Automatic Chest X-ray Report Generation [paper]
AlignTransformer: Hierarchical Alignment of Visual Regions and Disease Tags for Medical Report Generation [paper]
DeltaNet: Conditional Medical Report Generation for COVID-19 Diagnosis [paper]
Auto-Encoding Knowledge Graph for Unsupervised Medical Report Generation [paper]
MedCLIP: Contrastive Learning from Unpaired Medical Images and Text [paper]
Generating Radiology Reports via Memory-driven Transformer [paper]
Radiology report generation with a learned knowledge base and multi-modal alignment [paper]
Automatic Radiology Report Generation Based on Multi-view Image Fusion and Medical Concept Enrichment [paper]
When Radiology Report Generation Meets Knowledge Graph [paper]
Exploring and Distilling Posterior and Prior Knowledge for Radiology Report Generation [paper]
Auxiliary signal-guided knowledge encoder-decoder for medical report generation [paper]
Knowledge matters: Chest radiology report generation with general and specific knowledge [paper]
KiUT: Knowledge-injected U-Transformer for Radiology Report Generation [paper]
Dynamic Graph Enhanced Contrastive Learning for Chest X-ray Report Generation [paper]
ChatGPT makes medicine easy to swallow: an exploratory case study on simplified radiology reports [paper]
Translating radiology reports into plain language using ChatGPT and GPT-4 with prompt learning: results, limitations, and potential [paper]
Evaluating the performance of Generative Pre-trained Transformer-4 (GPT-4) in standardizing radiology reports [paper]
Evaluating GPT-4 on Impressions Generation in Radiology Reports [paper]
ImpressionGPT: An Iterative Optimizing Framework for Radiology Report Summarization with ChatGPT [paper]
ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models [paper], [code]
ChatCAD+: Towards a Universal and Reliable Interactive CAD using LLMs [paper], [code]
Style-Aware Radiology Report Generation with RadGraph and Few-Shot Prompting [paper]
Exploring the Boundaries of GPT-4 in Radiology [paper]
RadLLM: A Comprehensive Healthcare Benchmark of Large Language Models for Radiology [paper]
Radiology-GPT: A Large Language Model for Radiology [paper]
Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data [paper]
MAIRA-1: A specialised large multimodal model for radiology report generation [paper], [code]
A Survey on Medical Report Generation: From Deep Neural Networks to Large Language Models
32
5 commits
updated Feb 23, 2024
This repository contains a list of papers, codes, datasets in medical report generation (MRG) field. If you found any error, please don't hesitate to open an issue or pull request.
Given a radiology image, the main objective of medical report generation (MRG) is to generate a descriptive medical report as shown in Figure 1. Current studies leverage medical reports written by professional radiologists as the reference, whose output is expected to be as close to as possible. Formally, they adhere to standard optimization procedures by employing the cross-entropy loss to compare the generated report against the gold standard report.

Figure 1: Two representative cases in MRG, which consist of chest radiology images with their corresponding radiology reports, respectively. Formally, the goal of MRG is to generate the 'Findings' content from one or multiple radiology images.
The most widely used datasets are Indiana University Chest X-ray (IU X-Ray) and MIMIC Chest X-ray (MIMIC-CXR).
IU X-Ray is a widely-used benchmark in MRG systems, which was introduced by Indiana University. It comprises 7,470 chest X-ray images associated with 3,955 radiology reports. Existing MRG studies mostly follow the data splits proposed by Chen et al., which first exclude the samples without the ``Finding'' section, and partition IU X-Ray into the train, validation, and test sets with the ratio of 70%, 10%, and 20%, respectively.
MIMIC-CXR is a recently released large-scale benchmark dataset, which is provided by the Beth Israel Deaconess, USA. It includes 377,110 chest X-ray images and 227,835 reports. Formally, the official data splits are widely adopted. Thus, there are 368,960/2,991/5,159 cases for train/validation/test.
To calculate the performance of MRG, the natural language generation (NLG) metrics, i.e., BLEU-n, METEOR, ROUGE-n, and CIDEr, are widely used. These metrics measure the match between the generated reports and reference reports annotated by professional radiologists. In detail, NLG Metrics are utilized to measure the descriptive accuracy of predicted reports.
bilingual evaluation understudy (BLEU-n) is initially introduced for machine translation, which measures the n-gram precision of generated tokens. BLEU-n is usually employed in the evaluation of MRG approaches, with n ranging from 1 up to 4. This metric assesses the accuracy and coherence of the generated reports to a certain extent.
metric for evaluation of translation with explicit ordering (METEOR) is initially proposed for machine translation, which computes the recall of matching uni-grams from tokens in produced and gold standard reports according to their exact stemmed form and meaning.
recall-oriented understudy for gisting evaluation (ROUGE-L and METEOR) is initially designed for summarization, which measures the similarity between the generated and gold standard report based on their longest common subsequence (LCS) tokens.
Consensus-based image description evaluation (CIDEr) is initially designed to evaluate the quality of generated descriptions for natural images. In MRG systems, CIDEr evaluates models by rewarding topic-specific terms (terminologies in MRG) and penalizing overly frequent terms.
However, existing NLG evaluation metrics are not tailored to evaluate the accurate reporting of abnormalities in the image, which is the core value and urgent problem of MRG. Thus, additional clinical efficacy (CE) metrics are proposed to specifically measure the correctness of descriptions of clinical abnormalities. CE metrics are widely employed to capture and evaluate clinical correctness of predicted reports.
To calculate CE metrics, medical labelers, i.e., CheXpert are utilized to annotate tokens in both generated report and the gold standard one across 14 categories of diseases and support devices, producing the Precision, Recall, and F1 scores.
In addition to above automatic metrics, the human evaluation is also conducted in Liu et al, Liu et al etc. In these works, they invite professional radiologists to rate the quality of generated reports from faithfulness and comprehensiveness perspectives. However, the manual evaluation is both time-consuming and costly given large amount of reports and different radiologists many have conflicted opinion in labeling.

Figure 2: Three categories for MRG, including (a) Data-driven Encoder-Decoder based; (b) Medical Knowledge Enhanced, and (c) Large Language Model based frameworks.
On the Automatic Generation of Medical Imaging Reports [paper], [code]
Hybrid Retrieval-Generation Reinforced Agent for Medical Image Report Generation [paper]
Show, Describe and Conclude: On Exploiting the Structure Information of Chest X-ray Reports [paper]
Competence-based Multimodal Curriculum Learning for Medical Report Generation [paper]
Automatic Radiology Report Generation by Learning with Increasingly Hard Negatives [paper]
Contrastive Attention for Automatic Chest X-ray Report Generation [paper]
AlignTransformer: Hierarchical Alignment of Visual Regions and Disease Tags for Medical Report Generation [paper]
DeltaNet: Conditional Medical Report Generation for COVID-19 Diagnosis [paper]
Auto-Encoding Knowledge Graph for Unsupervised Medical Report Generation [paper]
MedCLIP: Contrastive Learning from Unpaired Medical Images and Text [paper]
Generating Radiology Reports via Memory-driven Transformer [paper]
Radiology report generation with a learned knowledge base and multi-modal alignment [paper]
Automatic Radiology Report Generation Based on Multi-view Image Fusion and Medical Concept Enrichment [paper]
When Radiology Report Generation Meets Knowledge Graph [paper]
Exploring and Distilling Posterior and Prior Knowledge for Radiology Report Generation [paper]
Auxiliary signal-guided knowledge encoder-decoder for medical report generation [paper]
Knowledge matters: Chest radiology report generation with general and specific knowledge [paper]
KiUT: Knowledge-injected U-Transformer for Radiology Report Generation [paper]
Dynamic Graph Enhanced Contrastive Learning for Chest X-ray Report Generation [paper]
ChatGPT makes medicine easy to swallow: an exploratory case study on simplified radiology reports [paper]
Translating radiology reports into plain language using ChatGPT and GPT-4 with prompt learning: results, limitations, and potential [paper]
Evaluating the performance of Generative Pre-trained Transformer-4 (GPT-4) in standardizing radiology reports [paper]
Evaluating GPT-4 on Impressions Generation in Radiology Reports [paper]
ImpressionGPT: An Iterative Optimizing Framework for Radiology Report Summarization with ChatGPT [paper]
ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models [paper], [code]
ChatCAD+: Towards a Universal and Reliable Interactive CAD using LLMs [paper], [code]
Style-Aware Radiology Report Generation with RadGraph and Few-Shot Prompting [paper]
Exploring the Boundaries of GPT-4 in Radiology [paper]
RadLLM: A Comprehensive Healthcare Benchmark of Large Language Models for Radiology [paper]
Radiology-GPT: A Large Language Model for Radiology [paper]
Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data [paper]
MAIRA-1: A specialised large multimodal model for radiology report generation [paper], [code]