Instructions to use guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M # Run inference directly in the terminal: llama cli -hf guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M # Run inference directly in the terminal: llama cli -hf guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M
Use Docker
docker model run hf.co/guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M with Ollama:
ollama run hf.co/guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M with Docker Model Runner:
docker model run hf.co/guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M
- Lemonade
How to use guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M
Run and chat with the model
lemonade run user.DeepSeek-R1-0528-Qwen3-8B_Q4_K_M-Q4_K_M
List all available models
lemonade list
- Atomic Chat
DeepSeek-R1-0528-Qwen3-8B - GGUF Quantized (Q4_K_M)
1. Introduction
This repository contains the
DeepSeek-R1-0528 8B
model quantized to GGUF format using llama.cpp. It is in Q4_K_M format,
suitable for fast inference on CPU or GPU.
2. Model Info
- Base model:
deepseek-ai/DeepSeek-R1-0528-Qwen3-8B - Format: GGUF
- Tool:
llama.cppbuilt with-DGGML_ENABLE_ARM=ON. - Conversion: bfloat16 safetensors => bfloat16
- Quantization: bfloat16 => Q4_K_M
- File:
deepseek-r1-0528-qwen3-8b-q4_k_m.gguf
3. How to Run Locally
Tested on RockChip RK3588/RK3588S running ARM CPU cores.
llama.cpp
After cloning and building llama.cpp with -DGGML_ENABLE_ARM=ON use
llama-run tool.
./llama.cpp/build/bin/llama-run ./deepseek-r1-0528-qwen3-8b-q4_k_m.gguf --prompt "What is the capital of France?"
Ollama
ollama run hf.co/guynich/DeepSeek-R1-0528-Qwen3-8B_Q4_K_M:Q4_K_M
I did see some looping during thinking discussed in this post.
4. License
Please refer to the original license for terms of use.
- Downloads last month
- 20
4-bit