OPPOer/X2I

Model

8

stars

13

commits

1

linked in READMEs

Apr 3, 2025

updated

any-to-image
audio-to-image
diffusers
flux.1
internvl
minicpm-o
multi-image-to-image
multilingual
qwenvl
speech-to-image
text_image-to-image
text-to-image
video-to-image
Browse cluster: Multimodal Vision-Language Models

README

X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation

Citation

🌟 If you find our work helpful, please consider citing our paper and leaving valuable stars

@misc{ma2025x2i,
    title={X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation},
    author={Jian Ma and Qirong Peng and Xu Guo and Chen Chen and Haonan Lu and Zhenyu Yang},
    year={2025},
    eprint={2503.06134},
    archivePrefix={arXiv},
    primaryClass={cs.CV}
}

License

This model is released under the Apache 2.0 License.

Contributors

majian0318

12 commits

nielsr

1 commits

OPPOer/X2I

Model

8

stars

13

commits

1

linked in READMEs

Apr 3, 2025

updated

any-to-image
audio-to-image
diffusers
flux.1
internvl
minicpm-o
multi-image-to-image
multilingual
qwenvl
speech-to-image
text_image-to-image
text-to-image
video-to-image
Browse cluster: Multimodal Vision-Language Models

README

X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation

Citation

🌟 If you find our work helpful, please consider citing our paper and leaving valuable stars

@misc{ma2025x2i,
    title={X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation},
    author={Jian Ma and Qirong Peng and Xu Guo and Chen Chen and Haonan Lu and Zhenyu Yang},
    year={2025},
    eprint={2503.06134},
    archivePrefix={arXiv},
    primaryClass={cs.CV}
}

License

This model is released under the Apache 2.0 License.

Contributors

majian0318

12 commits

nielsr

1 commits