Live demo for GRACE-VLM: INT4 Quantization-Aware Distillation for Vision-Language Models, accepted at ICML 2026. Read the paper at arXiv:2601.22709.
Upload an image and ask a question to run the GRACE 2B BF16 checkpoint on free
ZeroGPU hardware. For genuine packed INT4 inference, use
ForeverBlue/Qwen3-VL-2B-GRACE-W4G128-AWQ
with the copy-ready GRACE loader.
9 commits
Live demo for GRACE-VLM: INT4 Quantization-Aware Distillation for Vision-Language Models, accepted at ICML 2026. Read the paper at arXiv:2601.22709.
Upload an image and ask a question to run the GRACE 2B BF16 checkpoint on free
ZeroGPU hardware. For genuine packed INT4 inference, use
ForeverBlue/Qwen3-VL-2B-GRACE-W4G128-AWQ
with the copy-ready GRACE loader.
9 commits