yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF Text Generation • 12B • Updated Jun 19 • 295k • 2.76k
view reply I tested this https://huggingface.co/unsloth/gemma-4-12B-it-qat-GGUF on an RTX 4060 (8GB VRAM) using https://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant, and it worked perfectly. I even used the assistant for MTP https://huggingface.co/Janvitos/gemma-4-12B-it-qat-assistant-MTP-Q8_0-GGUF/tree/main and everything loaded into VRAM.