ncclab-sustech/BrainFLORA

[ACM MM 2025 Oral] Reproducible multimodal neural embeddings for EEG, MEG, and fMRI visual retrieval, reconstruction, and captioning.

Python

39

36 commits

updated Jun 19, 2026

See the code

README

BrainFLORA: Uncovering Brain Concept Representation
via Multimodal Neural Embeddings

Paper Hugging Face License

Dongyang Li, Haoyang Qin, Mingyang Wu, Chen Wei, Quanying Liu
Southern University of Science and Technology

BrainFLORA overview

BrainFLORA framework

BrainFLORA is a reproducible multimodal neural embedding framework for EEG, MEG, and fMRI visual retrieval, image reconstruction, and image captioning. The repository provides unified command-line entry points for training, inference, and evaluation with the released checkpoints.

News

  • 2026-06-19: The codebase was reorganized and Hugging Face checkpoint weights were updated.
  • 2025-12-21: Preprocessed datasets and pretrained checkpoints were released on Hugging Face.
  • 2025-07-15: The arXiv paper was released.
  • 2025-07-12: The codebase was released.
  • 2025-07-05: BrainFLORA was accepted by ACM MM 2025.

Installation

git clone https://github.com/ncclab-sustech/BrainFLORA.git
cd BrainFLORA

conda env create -f environment.yml
conda activate BrainFLORA
pip install -e .

If you prefer the setup script from the original repository:

bash setup.sh
conda activate BrainFLORA
pip install -e .

Caption evaluation may require optional metric packages such as clip and pycocoevalcap. Caption generation uses Shikra; download the model assets from the original Shikra project and place them locally:

external_models/shikra-7b
external_models/mm_projector.bin

Shikra resources:

Data and Checkpoints

Raw datasets used by the original project:

DatasetLinkDatasetLink
THINGS-EEG1OpenNeuro ds003825THINGS-EEG2OSF 3jk45
THINGS-MEGOpenNeuro ds004212THINGS-fMRIOpenNeuro ds004192
THINGS-ImagesOSF rdxy2

Use data_preparing/ if you need to preprocess raw data yourself.

Quick Run Experiments

Run pretrained checkpoint inference and evaluation for visual retrieval, reconstruction, and captioning.

1. Visual Retrieval

Evaluate retrieval checkpoints:

python eval/reproduce_retrieval.py \
  --device cuda:0 \
  --models single unified \
  --modalities eeg meg fmri \
  --batch-size 128 \
  --output-dir outputs/reproduction

Outputs:

outputs/reproduction/retrieval_reproduction_metrics.csv
outputs/reproduction/retrieval_reproduction_metrics.json
outputs/reproduction/retrieval_reproduction_vs_paper.json

2. Visual Reconstruction

Generate reconstructed images with cached SDXL/IP-Adapter assets:

python eval/FLORA_inference_reconst.py \
  --device cuda:0 \
  --modalities eeg meg fmri \
  --prior-steps 10 \
  --image-steps 4 \
  --local-files-only \
  --skip-existing \
  --output-dir outputs/reconstruction_png_full \
  --summary-json outputs/reconstruction_png_full/reconstruction_png_full_summary.json

The full run contains 3100 generated images:

EEG: 10 subjects x 200 images
MEG: 4 subjects x 200 images
fMRI: 3 subjects x 100 images

Evaluate reconstruction metrics:

python eval/evaluate_reconstruction_metrics.py \
  --recon-root outputs/reconstruction_png_full \
  --modalities eeg meg fmri \
  --gt-root /path/to/dataset_root \
  --output-dir outputs/reconstruction_metrics

3. Visual Captioning

Generate woPrior caption token embeddings and apply the caption diffusion prior:

python eval/FLORA_inference_caption_embeddings.py \
  --device cuda:0 \
  --stage all \
  --modalities eeg meg fmri \
  --features-root features/FLORA \
  --metrics-json outputs/caption_embedding/caption_woPrior_metrics.json \
  --summary-json outputs/caption_embedding/caption_prior_summary.json

Generate caption JSON/TXT files with Shikra:

python eval/shikra_caption.py \
  --device cuda:0 \
  --modalities eeg meg fmri \
  --conditions Prior woPrior \
  --embeddings-root features/FLORA \
  --output-root Caption \
  --shikra-path external_models/shikra-7b \
  --mm-projector-path external_models/mm_projector.bin \
  --local-files-only

Dry-run Shikra caption generation without loading the model:

python eval/shikra_caption.py \
  --modalities fmri \
  --conditions Prior \
  --subjects sub-01 \
  --embeddings-root features/FLORA \
  --dry-run

Evaluate caption metrics:

python Caption/evaluate_caption_metrics.py \
  --modalities eeg fmri \
  --conditions Prior woPrior \
  --candidate-pattern 'shikra_{tag}_sub_{subject}_caption.json' \
  --output-dir outputs/caption_metrics

The default references are:

Caption/EEG_caption/caption_EEG_GT.json
Caption/fMRI_caption/caption_fMRI_GT.json

Quick Training

Unified training scripts are provided for retraining and ablations. Update dataset paths in the configs or pass CLI arguments before launching.

Train retrieval encoders:

# EEG
python Retrieval/train_retrieval.py \
  --modality eeg \
  --gpu cuda:0 \
  --output-dir outputs/contrast

# MEG
python Retrieval/train_retrieval.py \
  --modality meg \
  --gpu cuda:0 \
  --output-dir outputs/contrast

# fMRI
python Retrieval/train_retrieval.py \
  --modality fmri \
  --gpu cuda:0 \
  --output-dir outputs/contrast

Train the unified encoder for retrieval:

python train/train_unified_encoder.py \
  --task retrieval \
  --modalities eeg meg fmri \
  --gpu cuda:0 \
  --output_dir outputs/contrast

Train the unified encoder for reconstruction:

python train/train_unified_encoder.py \
  --task reconstruction \
  --modalities eeg meg fmri \
  --gpu cuda:0 \
  --output_dir outputs/contrast

Train the caption-aligned unified encoder:

python train/train_unified_encoder.py \
  --task caption \
  --use-caption \
  --modalities eeg meg fmri \
  --gpu cuda:0 \
  --output_dir outputs/contrast

Distributed reconstruction training:

accelerate launch train/train_unified_encoder.py \
  --task reconstruction \
  --distributed \
  --modalities eeg meg fmri \
  --output_dir outputs/contrast

Citation

@inproceedings{li2025brainflora,
  author = {Li, Dongyang and Qin, Haoyang and Wu, Mingyang and Wei, Chen and Liu, Quanying},
  title = {BrainFLORA: Uncovering Brain Concept Representation via Multimodal Neural Embeddings},
  year = {2025},
  isbn = {9798400720352},
  publisher = {Association for Computing Machinery},
  address = {New York, NY, USA},
  url = {https://doi.org/10.1145/3746027.3754996},
  doi = {10.1145/3746027.3754996},
  booktitle = {Proceedings of the 33rd ACM International Conference on Multimedia},
  pages = {5577--5586}
}

@article{li2024visual,
  title = {Visual Decoding and Reconstruction via EEG Embeddings with Guided Diffusion},
  author = {Li, Dongyang and Wei, Chen and Li, Shiying and Zou, Jiachen and Liu, Quanying},
  journal = {Advances in Neural Information Processing Systems},
  volume = {37},
  pages = {102822--102864},
  year = {2024}
}

@inproceedings{wei2024cocog,
  title = {CoCoG: controllable visual stimuli generation based on human concept representations},
  author = {Wei, Chen and Zou, Jiachen and Heinke, Dietmar and Liu, Quanying},
  booktitle = {Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence},
  pages = {3178--3186},
  year = {2024}
}

😺Acknowledge

1.Thanks to Y Song et al. for their contribution in data set preprocessing and neural network structure, we refer to their work:"Decoding Natural Images from EEG for Object Recognition". Yonghao Song, Bingchuan Liu, Xiang Li, Nanlin Shi, Yijun Wang, and Xiaorong Gao.

2.We also thank the authors of SDRecon for providing the codes and the results. Some parts of the training script are based on MindEye and MindEye2. Thanks for the awesome research works.

3.Here we provide the THING-EEG2 dataset cited in the paper: "A large and rich EEG dataset for modeling human visual object recognition". Alessandro T. Gifford, Kshitij Dwivedi, Gemma Roig, Radoslaw M. Cichy.

4.Another used THINGS-MEG and THINGS-fMRI data set provides a reference:"THINGS-data, a multimodal collection of large-scale datasets for investigating object representations in human brain and behavior". Hebart, Martin N., Oliver Contier, Lina Teichmann, Adam H. Rockter, Charles Y. Zheng, Alexis Kidder, Anna Corriveau, Maryam Vaziri-Pashkam, and Chris I. Baker.

5.We use the "BrainHub" for visual caption evaluation from "UMBRAE: Unified Multimodal Brain Decoding (ECCV 2024)" Xia, Weihao and de Charette, Raoul and Oztireli, Cengiz and Xue, Jing-Hao.

Contact Dongyang Li if you have any questions or suggestions.

License

This repository is released under the MIT license. See LICENSE for details.

ncclab-sustech/BrainFLORA

[ACM MM 2025 Oral] Reproducible multimodal neural embeddings for EEG, MEG, and fMRI visual retrieval, reconstruction, and captioning.

Python

39

36 commits

updated Jun 19, 2026

See the code

README

BrainFLORA: Uncovering Brain Concept Representation
via Multimodal Neural Embeddings

Paper Hugging Face License

Dongyang Li, Haoyang Qin, Mingyang Wu, Chen Wei, Quanying Liu
Southern University of Science and Technology

BrainFLORA overview

BrainFLORA framework

BrainFLORA is a reproducible multimodal neural embedding framework for EEG, MEG, and fMRI visual retrieval, image reconstruction, and image captioning. The repository provides unified command-line entry points for training, inference, and evaluation with the released checkpoints.

News

  • 2026-06-19: The codebase was reorganized and Hugging Face checkpoint weights were updated.
  • 2025-12-21: Preprocessed datasets and pretrained checkpoints were released on Hugging Face.
  • 2025-07-15: The arXiv paper was released.
  • 2025-07-12: The codebase was released.
  • 2025-07-05: BrainFLORA was accepted by ACM MM 2025.

Installation

git clone https://github.com/ncclab-sustech/BrainFLORA.git
cd BrainFLORA

conda env create -f environment.yml
conda activate BrainFLORA
pip install -e .

If you prefer the setup script from the original repository:

bash setup.sh
conda activate BrainFLORA
pip install -e .

Caption evaluation may require optional metric packages such as clip and pycocoevalcap. Caption generation uses Shikra; download the model assets from the original Shikra project and place them locally:

external_models/shikra-7b
external_models/mm_projector.bin

Shikra resources:

Data and Checkpoints

Raw datasets used by the original project:

DatasetLinkDatasetLink
THINGS-EEG1OpenNeuro ds003825THINGS-EEG2OSF 3jk45
THINGS-MEGOpenNeuro ds004212THINGS-fMRIOpenNeuro ds004192
THINGS-ImagesOSF rdxy2

Use data_preparing/ if you need to preprocess raw data yourself.

Quick Run Experiments

Run pretrained checkpoint inference and evaluation for visual retrieval, reconstruction, and captioning.

1. Visual Retrieval

Evaluate retrieval checkpoints:

python eval/reproduce_retrieval.py \
  --device cuda:0 \
  --models single unified \
  --modalities eeg meg fmri \
  --batch-size 128 \
  --output-dir outputs/reproduction

Outputs:

outputs/reproduction/retrieval_reproduction_metrics.csv
outputs/reproduction/retrieval_reproduction_metrics.json
outputs/reproduction/retrieval_reproduction_vs_paper.json

2. Visual Reconstruction

Generate reconstructed images with cached SDXL/IP-Adapter assets:

python eval/FLORA_inference_reconst.py \
  --device cuda:0 \
  --modalities eeg meg fmri \
  --prior-steps 10 \
  --image-steps 4 \
  --local-files-only \
  --skip-existing \
  --output-dir outputs/reconstruction_png_full \
  --summary-json outputs/reconstruction_png_full/reconstruction_png_full_summary.json

The full run contains 3100 generated images:

EEG: 10 subjects x 200 images
MEG: 4 subjects x 200 images
fMRI: 3 subjects x 100 images

Evaluate reconstruction metrics:

python eval/evaluate_reconstruction_metrics.py \
  --recon-root outputs/reconstruction_png_full \
  --modalities eeg meg fmri \
  --gt-root /path/to/dataset_root \
  --output-dir outputs/reconstruction_metrics

3. Visual Captioning

Generate woPrior caption token embeddings and apply the caption diffusion prior:

python eval/FLORA_inference_caption_embeddings.py \
  --device cuda:0 \
  --stage all \
  --modalities eeg meg fmri \
  --features-root features/FLORA \
  --metrics-json outputs/caption_embedding/caption_woPrior_metrics.json \
  --summary-json outputs/caption_embedding/caption_prior_summary.json

Generate caption JSON/TXT files with Shikra:

python eval/shikra_caption.py \
  --device cuda:0 \
  --modalities eeg meg fmri \
  --conditions Prior woPrior \
  --embeddings-root features/FLORA \
  --output-root Caption \
  --shikra-path external_models/shikra-7b \
  --mm-projector-path external_models/mm_projector.bin \
  --local-files-only

Dry-run Shikra caption generation without loading the model:

python eval/shikra_caption.py \
  --modalities fmri \
  --conditions Prior \
  --subjects sub-01 \
  --embeddings-root features/FLORA \
  --dry-run

Evaluate caption metrics:

python Caption/evaluate_caption_metrics.py \
  --modalities eeg fmri \
  --conditions Prior woPrior \
  --candidate-pattern 'shikra_{tag}_sub_{subject}_caption.json' \
  --output-dir outputs/caption_metrics

The default references are:

Caption/EEG_caption/caption_EEG_GT.json
Caption/fMRI_caption/caption_fMRI_GT.json

Quick Training

Unified training scripts are provided for retraining and ablations. Update dataset paths in the configs or pass CLI arguments before launching.

Train retrieval encoders:

# EEG
python Retrieval/train_retrieval.py \
  --modality eeg \
  --gpu cuda:0 \
  --output-dir outputs/contrast

# MEG
python Retrieval/train_retrieval.py \
  --modality meg \
  --gpu cuda:0 \
  --output-dir outputs/contrast

# fMRI
python Retrieval/train_retrieval.py \
  --modality fmri \
  --gpu cuda:0 \
  --output-dir outputs/contrast

Train the unified encoder for retrieval:

python train/train_unified_encoder.py \
  --task retrieval \
  --modalities eeg meg fmri \
  --gpu cuda:0 \
  --output_dir outputs/contrast

Train the unified encoder for reconstruction:

python train/train_unified_encoder.py \
  --task reconstruction \
  --modalities eeg meg fmri \
  --gpu cuda:0 \
  --output_dir outputs/contrast

Train the caption-aligned unified encoder:

python train/train_unified_encoder.py \
  --task caption \
  --use-caption \
  --modalities eeg meg fmri \
  --gpu cuda:0 \
  --output_dir outputs/contrast

Distributed reconstruction training:

accelerate launch train/train_unified_encoder.py \
  --task reconstruction \
  --distributed \
  --modalities eeg meg fmri \
  --output_dir outputs/contrast

Citation

@inproceedings{li2025brainflora,
  author = {Li, Dongyang and Qin, Haoyang and Wu, Mingyang and Wei, Chen and Liu, Quanying},
  title = {BrainFLORA: Uncovering Brain Concept Representation via Multimodal Neural Embeddings},
  year = {2025},
  isbn = {9798400720352},
  publisher = {Association for Computing Machinery},
  address = {New York, NY, USA},
  url = {https://doi.org/10.1145/3746027.3754996},
  doi = {10.1145/3746027.3754996},
  booktitle = {Proceedings of the 33rd ACM International Conference on Multimedia},
  pages = {5577--5586}
}

@article{li2024visual,
  title = {Visual Decoding and Reconstruction via EEG Embeddings with Guided Diffusion},
  author = {Li, Dongyang and Wei, Chen and Li, Shiying and Zou, Jiachen and Liu, Quanying},
  journal = {Advances in Neural Information Processing Systems},
  volume = {37},
  pages = {102822--102864},
  year = {2024}
}

@inproceedings{wei2024cocog,
  title = {CoCoG: controllable visual stimuli generation based on human concept representations},
  author = {Wei, Chen and Zou, Jiachen and Heinke, Dietmar and Liu, Quanying},
  booktitle = {Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence},
  pages = {3178--3186},
  year = {2024}
}

😺Acknowledge

1.Thanks to Y Song et al. for their contribution in data set preprocessing and neural network structure, we refer to their work:"Decoding Natural Images from EEG for Object Recognition". Yonghao Song, Bingchuan Liu, Xiang Li, Nanlin Shi, Yijun Wang, and Xiaorong Gao.

2.We also thank the authors of SDRecon for providing the codes and the results. Some parts of the training script are based on MindEye and MindEye2. Thanks for the awesome research works.

3.Here we provide the THING-EEG2 dataset cited in the paper: "A large and rich EEG dataset for modeling human visual object recognition". Alessandro T. Gifford, Kshitij Dwivedi, Gemma Roig, Radoslaw M. Cichy.

4.Another used THINGS-MEG and THINGS-fMRI data set provides a reference:"THINGS-data, a multimodal collection of large-scale datasets for investigating object representations in human brain and behavior". Hebart, Martin N., Oliver Contier, Lina Teichmann, Adam H. Rockter, Charles Y. Zheng, Alexis Kidder, Anna Corriveau, Maryam Vaziri-Pashkam, and Chris I. Baker.

5.We use the "BrainHub" for visual caption evaluation from "UMBRAE: Unified Multimodal Brain Decoding (ECCV 2024)" Xia, Weihao and de Charette, Raoul and Oztireli, Cengiz and Xue, Jing-Hao.

Contact Dongyang Li if you have any questions or suggestions.

License

This repository is released under the MIT license. See LICENSE for details.

Languages

Python

99.7%