8 repos
Louym/awq4nvomni
A temporary repo
0
116 commits
mit-han-lab/llm-awq
[MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and…
3,633
vimarsh244/llm-awq
casper-hansen/AutoAWQ
AutoAWQ implements the AWQ algorithm for 4-bit quantization with a 2x speedup during inference.…
2,348
546 commits
phunggiahuy159/awq-embed
No description
9 commits
mit-han-lab/TinyChatEngine
TinyChatEngine: On-Device LLM Inference Library
962
14 commits
SqueezeBits/QUICK
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
123
13 commits
Efficient-Large-Model/NVILA-AWQ
3