Inference Providers
Active filters: gptq
StillDeadcode/qwen3.8-next-flash-fp8-iq4r-moe
Image-Text-to-Text
• Updated • 4
EliovpAI/Qwen3.8-27B-W3Rot-INT3-Paiton-RDNA4
Text Generation
• Updated • 127
• 3
syvai/Qwen3.8-27B-DFlash2-W4A16
2B • Updated • 9.03k
• 18
canada-quant/GLM-5.3-Flash-W4A16-MTP
Image-Text-to-Text
• 50B • Updated • 11.2k
• 17
systalyze/gemma-4-26B-A4B-it-ExpertsINT4-FP8
Image-Text-to-Text
• 26B • Updated • 1.14k
• 2
azampatti/Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound
Text Generation
• 124B • Updated • 1.88k
• 9
lvkaokao/DeepSeek-V4.1-Flash-W4A16-Engram-AutoRound
Text Generation
• 301B • Updated • 387
• 3
INCModel3/DeepSeek-V4.1-Flash-W4A16-Engram-AutoRound
Text Generation
• 301B • Updated • 422
• 2
bjonor/Swift-1.5-Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16
Image-Text-to-Text
• 28B • Updated • 614
• 2
valoomba/Qwen3.8-Flash-Next-Uncensored-W4A16-Attn8-FP8PLE
Image-Text-to-Text
• 71B • Updated • 192
• 2
0xBakeer/TandemLLM-Qwen3.8-27B-NVFP4
TheBloke/dolphin-2.2.1-mistral-7B-GPTQ
Text Generation
• 7B • Updated • 115
• 33
TheBloke/deepseek-coder-33B-base-GPTQ
Text Generation
• 33B • Updated • 79
• 3
Qwen/Qwen2.5-0.5B-Instruct-GPTQ-Int8
Text Generation
• 0.5B • Updated • 608
• 11
Qwen/Qwen2.5-Coder-7B-Instruct-GPTQ-Int4
Text Generation
• 8B • Updated • 441k
• 17
Qwen/Qwen2.5-Coder-14B-Instruct-GPTQ-Int4
Text Generation
• 15B • Updated • 19.6k
• 9
Qwen/Qwen2.5-Coder-32B-Instruct-GPTQ-Int4
Text Generation
• 33B • Updated • 1.62k
• 25
AtlaAI/Selene-1-Mini-Llama-3.1-8B-GPTQ-W4A16
Text Generation
• 8B • Updated • 88
• 2
hfl/Qwen2.5-VL-7B-Instruct-GPTQ-Int4
Image-Text-to-Text
• 8B • Updated • 2.75k
• 11
OPEA/DeepSeek-R1-Distill-Llama-70B-int4-gptq-sym-inc
71B • Updated • 21
• 4
Qwen/Qwen3-30B-A3B-GPTQ-Int4
Text Generation
• 31B • Updated • 60.7k
• 58
tencent/HY-MT1.5-1.8B-GPTQ-Int4
Translation
• 2B • Updated • 447
• 16
Qwen/Qwen3.5-35B-A3B-GPTQ-Int4
Image-Text-to-Text
• 36B • Updated • 239k
• 93
ciocan/gemma-4-E4B-it-W4A16
Image-Text-to-Text
• 4B • Updated • 2.93k
• 3
marcusnogueira/granite-guardian-4.1-8b-gptq-4bit
8B • Updated • 19
• 1
Sebesky/MiniMax-M3-W4A16-GPTQ
Image-Text-to-Text
• 430B • Updated • 168
• 4
Text Generation
• 299B • Updated • 259
• 10
pipecat-ai/NVIDIA-NemotronLabs-VoiceChat-11B-Spark
Audio-to-Audio
• Updated • 25
Vishva007/Qwen3.8-27B-W4A16-AutoRound-GPTQ
Image-Text-to-Text
• 3B • Updated • 61k
• 11
dbirks/Qwen3.8-27B-W4A16-AutoRound
Image-Text-to-Text
• 6B • Updated • 46.1k
• 38