For the Chinese version of the README, please refer to 中文文档.
In the VLM (Visual Language Model), the visual component utilizes the CLIP or SIGLIP models, which have already achieved preliminary semantic alignment. A two-layer MLP is used for feature mapping. By overriding the forward method of the QWenModel, the corresponding image tokens are replaced with visual features.
The code for running the model can be found at Basic-Visual-Language-Model.
Special thanks to the following projects for their great work 🙌:
If you have any questions or ideas, feel free to reach out to me 😊:
I will respond as soon as I see your email!
8 commits
For the Chinese version of the README, please refer to 中文文档.
In the VLM (Visual Language Model), the visual component utilizes the CLIP or SIGLIP models, which have already achieved preliminary semantic alignment. A two-layer MLP is used for feature mapping. By overriding the forward method of the QWenModel, the corresponding image tokens are replaced with visual features.
The code for running the model can be found at Basic-Visual-Language-Model.
Special thanks to the following projects for their great work 🙌:
If you have any questions or ideas, feel free to reach out to me 😊:
I will respond as soon as I see your email!
8 commits