ramendik/granite-4.0-h-small-stonebnb

Model

0

stars

5

commits

1

linked in READMEs

Feb 9, 2026

updated

granitemoehybrid
pytorch

README

IBM Granite 4-h Small with its MLP expert layers quantized with BitsandBytes, saved in the custom StoneBnB format to enable VRAM-efficient training.

(While this quant can be technically used for Transformers inference, it is not supported by any commmon server and the GGUF quants are probably much better. StoneBnB is intended for fine-tuning, producing an adapter that you can merge into the unquantized model)

Contributors

ramendik

5 commits

ramendik/granite-4.0-h-small-stonebnb

Model

0

stars

5

commits

1

linked in READMEs

Feb 9, 2026

updated

granitemoehybrid
pytorch

README

IBM Granite 4-h Small with its MLP expert layers quantized with BitsandBytes, saved in the custom StoneBnB format to enable VRAM-efficient training.

(While this quant can be technically used for Transformers inference, it is not supported by any commmon server and the GGUF quants are probably much better. StoneBnB is intended for fine-tuning, producing an adapter that you can merge into the unquantized model)

Contributors

ramendik

5 commits