9 repos
X-iZhang/Libra
[ACL 2025] ⚖️ Temporally-aware MLLM for Biomedical Radiology Analysis and Report Generation.…
32
158 commits
ATH-MaaS/Ovis
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual…
1,522
57 commits
AIDC-AI/Ovis
X-iZhang/RRG-BioNLP-ACL2024
[ACL 2024] 🔬 Med-CXRGen, developed by Glasgow AI4BioMed Lab, brings vision-language adaptation to…
1
29 commits
KAIST-Visual-AI-Group/APC-VLM
[ICCV 2025] Official code for Perspective-Aware Reasoning in Vision-Language Models via Mental…
66
6 commits
ByteDance-Seed/Seed1.5-VL
Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal…
1,588
cambrian-mllm/cambrian-s
Cambrian-S: Towards Spatial Supersensing in Video
569
16 commits
jolibrain/colette
Multimodal RAG to search and interact locally with technical documents of any kind
301
101 commits
X-iZhang/CCD
[ACL 2026] 📷 CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding —…
15
55 commits