ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation
11
12 commits
4 linked in READMEs
updated Jul 7, 2025
Github | Dataset(ModelScope) | Model | Paper
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in chart understanding tasks. However, interpreting charts with textual descriptions often leads to information loss, as it fails to fully capture the dense information embedded in charts. In contrast, parsing charts into code provides lossless representations that can effectively contain all critical details. Although existing open-source MLLMs have achieved success in chart understanding tasks, they still face two major challenges when applied to chart-to-code tasks: (1) Low executability and poor restoration of chart details in the generated code and (2) Lack of large-scale and diverse training data. To address these challenges, we propose ChartCoder, the first dedicated chart-to-code MLLM, which leverages Code LLMs as the language backbone to enhance the executability of the generated code. Furthermore, we introduce Chart2Code-160k, the first large-scale and diverse dataset for chart-to-code generation, and propose the Snippet-of-Thought (SoT) method, which transforms direct chart-to-code generation data into step-by-step generation. Experiments demonstrate that ChartCoder, with only 7B parameters, surpasses existing open-source MLLMs on chart-to-code benchmarks, achieving superior chart restoration and code excitability. Our code is available at this https URL .
This repository contains the code to train and infer ChartCoder.

git clone https://github.com/thunlp/ChartCoder.git
cd ChartCoder
conda create -n chartcoder python=3.10 -y
conda activate chartcoder
pip install --upgrade pip # enable PEP 660 support
pip install -e .
pip install -e ".[train]"
pip install flash-attn --no-build-isolation
| Model | Download Link |
|---|---|
| MLP Connector | projector |
| ChartCoder | ChartCoder |
The MLP Connector is our pre-trained MLP weights, which you could directly use for SFT.
| Dataset | Download Link |
|---|---|
| Chart2Code-160k | HuggingFace |
| Chart2Code-160k | ModelScope |
The whole training process consists of two stages. To train the ChartCoder, siglip-so400m-patch14-384 and deepseek-coder-6.7b-instruct should be downloaded first.
For Pre-training, run
bash scripts/train/pretrain_siglip.sh
For SFT, run
bash scripts/train/finetune_siglip_a4.sh
Please change the model path to your local path. See the corresponding .sh file for details.
We also provide other training scripts, such as using CLIP _clip and multiple machines _m. See scripts/train for further information.
Please see inference.py for details.
Please refer to our paper for detailed performance on ChartMimic, Plot2Code and ChartX benchmarks. Thanks for these contributions to the chart-to-code field.

For any questions, you can contact 2429527z@gmail.com.
If you find this work useful, consider giving this repository a star ⭐️ and citing 📝 our paper as follows:
@misc{zhao2025chartcoderadvancingmultimodallarge,
title={ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation},
author={Xuanle Zhao and Xianzhen Luo and Qi Shi and Chi Chen and Shuo Wang and Wanxiang Che and Zhiyuan Liu and Maosong Sun},
year={2025},
eprint={2501.06598},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2501.06598},
}
The code is based on the LLaVA-NeXT. Thanks for these great works and open sourcing!
ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation
11
12 commits
4 linked in READMEs
updated Jul 7, 2025
Github | Dataset(ModelScope) | Model | Paper
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in chart understanding tasks. However, interpreting charts with textual descriptions often leads to information loss, as it fails to fully capture the dense information embedded in charts. In contrast, parsing charts into code provides lossless representations that can effectively contain all critical details. Although existing open-source MLLMs have achieved success in chart understanding tasks, they still face two major challenges when applied to chart-to-code tasks: (1) Low executability and poor restoration of chart details in the generated code and (2) Lack of large-scale and diverse training data. To address these challenges, we propose ChartCoder, the first dedicated chart-to-code MLLM, which leverages Code LLMs as the language backbone to enhance the executability of the generated code. Furthermore, we introduce Chart2Code-160k, the first large-scale and diverse dataset for chart-to-code generation, and propose the Snippet-of-Thought (SoT) method, which transforms direct chart-to-code generation data into step-by-step generation. Experiments demonstrate that ChartCoder, with only 7B parameters, surpasses existing open-source MLLMs on chart-to-code benchmarks, achieving superior chart restoration and code excitability. Our code is available at this https URL .
This repository contains the code to train and infer ChartCoder.

git clone https://github.com/thunlp/ChartCoder.git
cd ChartCoder
conda create -n chartcoder python=3.10 -y
conda activate chartcoder
pip install --upgrade pip # enable PEP 660 support
pip install -e .
pip install -e ".[train]"
pip install flash-attn --no-build-isolation
| Model | Download Link |
|---|---|
| MLP Connector | projector |
| ChartCoder | ChartCoder |
The MLP Connector is our pre-trained MLP weights, which you could directly use for SFT.
| Dataset | Download Link |
|---|---|
| Chart2Code-160k | HuggingFace |
| Chart2Code-160k | ModelScope |
The whole training process consists of two stages. To train the ChartCoder, siglip-so400m-patch14-384 and deepseek-coder-6.7b-instruct should be downloaded first.
For Pre-training, run
bash scripts/train/pretrain_siglip.sh
For SFT, run
bash scripts/train/finetune_siglip_a4.sh
Please change the model path to your local path. See the corresponding .sh file for details.
We also provide other training scripts, such as using CLIP _clip and multiple machines _m. See scripts/train for further information.
Please see inference.py for details.
Please refer to our paper for detailed performance on ChartMimic, Plot2Code and ChartX benchmarks. Thanks for these contributions to the chart-to-code field.

For any questions, you can contact 2429527z@gmail.com.
If you find this work useful, consider giving this repository a star ⭐️ and citing 📝 our paper as follows:
@misc{zhao2025chartcoderadvancingmultimodallarge,
title={ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation},
author={Xuanle Zhao and Xianzhen Luo and Qi Shi and Chi Chen and Shuo Wang and Wanxiang Che and Zhiyuan Liu and Maosong Sun},
year={2025},
eprint={2501.06598},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2501.06598},
}
The code is based on the LLaVA-NeXT. Thanks for these great works and open sourcing!