Huihui Gemma 4 26B A4B IT Abliterated — GGUF Quantizations
36
81 commits
1 linked in READMEs
updated Aug 22, 2026
Huihui-gemma-4-26B-A4B-it-abliterated-GGUF is a GGUF release for llama.cpp-compatible runtimes and local inference, published by groxaxo.
It is intended for open-source evaluation, reproducible experimentation, and compatible local or
hosted inference workflows. The wording below is deliberately limited to what can be verified
from this repository's metadata and artifacts.
The repository name identifies a behavior-modified or reduced-filtering lineage. That label describes the source or conversion history; it is not a guarantee of unrestricted behavior in every prompt or runtime. Test outputs carefully before sharing or deploying them.
| Field | Details |
|---|---|
| Format | GGUF |
| Source / base | huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated |
| Intended task | image-text-to-text |
| License | the license declared in the repository files |
*.gguf (22 files)config.jsongeneration_config.jsontokenizer.jsontokenizer_config.jsonprocessor_config.jsonchat_template.jinjaDownload a .gguf file that fits your available memory, then run it with a current llama.cpp
build:
llama-cli \
-m /path/to/model.gguf \
-p "Write a concise technical summary."
For vision or any-to-any models, download the matching multimodal projection file when one is provided and follow the source model's modality-specific instructions.
Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.
Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.
This repository contains GGUF / llama.cpp quantized builds of:
huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated
These are UD quantizations prepared for efficient local inference with llama.cpp, including support for multimodal image-text-to-text workflows when used with the corresponding mmproj file.
This release is designed for users who want to run the Huihui Gemma 4 26B A4B abliterated model locally with reduced VRAM and RAM requirements while preserving as much output quality as possible.
The quantization variants use an optimized tensor distribution strategy inspired by Unsloth-style mixed-quality quantization recipes, balancing model fidelity, speed, and memory efficiency across different hardware targets.
.gguf model file from this repository.mmproj file.Example:
./llama-cli \
-m Huihui-Gemma-4-26B-A4B-it-abliterated-UD-Q4_K_XL.gguf \
--mmproj mmproj-model.gguf \
-p "Describe this image in detail."
Adjust the model filename and mmproj filename to match the files you downloaded.
Choose based on your available memory and quality target:
For best results, use the largest quantization your hardware can comfortably run.
This model supports image-text-to-text inference when used with the appropriate multimodal projection file.
Make sure the mmproj file matches this model family. Using an incorrect projection file may result in broken or degraded vision-language behavior.
This repository only provides quantized GGUF builds. Model behavior, alignment characteristics, and training details are inherited from the original base model and fine-tune.
Huihui Gemma 4 26B A4B IT Abliterated — GGUF Quantizations
36
81 commits
1 linked in READMEs
updated Aug 22, 2026
Huihui-gemma-4-26B-A4B-it-abliterated-GGUF is a GGUF release for llama.cpp-compatible runtimes and local inference, published by groxaxo.
It is intended for open-source evaluation, reproducible experimentation, and compatible local or
hosted inference workflows. The wording below is deliberately limited to what can be verified
from this repository's metadata and artifacts.
The repository name identifies a behavior-modified or reduced-filtering lineage. That label describes the source or conversion history; it is not a guarantee of unrestricted behavior in every prompt or runtime. Test outputs carefully before sharing or deploying them.
| Field | Details |
|---|---|
| Format | GGUF |
| Source / base | huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated |
| Intended task | image-text-to-text |
| License | the license declared in the repository files |
*.gguf (22 files)config.jsongeneration_config.jsontokenizer.jsontokenizer_config.jsonprocessor_config.jsonchat_template.jinjaDownload a .gguf file that fits your available memory, then run it with a current llama.cpp
build:
llama-cli \
-m /path/to/model.gguf \
-p "Write a concise technical summary."
For vision or any-to-any models, download the matching multimodal projection file when one is provided and follow the source model's modality-specific instructions.
Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.
Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.
This repository contains GGUF / llama.cpp quantized builds of:
huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated
These are UD quantizations prepared for efficient local inference with llama.cpp, including support for multimodal image-text-to-text workflows when used with the corresponding mmproj file.
This release is designed for users who want to run the Huihui Gemma 4 26B A4B abliterated model locally with reduced VRAM and RAM requirements while preserving as much output quality as possible.
The quantization variants use an optimized tensor distribution strategy inspired by Unsloth-style mixed-quality quantization recipes, balancing model fidelity, speed, and memory efficiency across different hardware targets.
.gguf model file from this repository.mmproj file.Example:
./llama-cli \
-m Huihui-Gemma-4-26B-A4B-it-abliterated-UD-Q4_K_XL.gguf \
--mmproj mmproj-model.gguf \
-p "Describe this image in detail."
Adjust the model filename and mmproj filename to match the files you downloaded.
Choose based on your available memory and quality target:
For best results, use the largest quantization your hardware can comfortably run.
This model supports image-text-to-text inference when used with the appropriate multimodal projection file.
Make sure the mmproj file matches this model family. Using an incorrect projection file may result in broken or degraded vision-language behavior.
This repository only provides quantized GGUF builds. Model behavior, alignment characteristics, and training details are inherited from the original base model and fine-tune.