Qwen · Qwen 3
Qwen 3 8B GGUF
Verified by Badgr · ServeQwen 3 8B GGUF is a 8B chat and text-generation model available through supported Badgr execution routes.
Last validated or meaningfully updated: 2026-10-10
✓ Verified by Badgr · llama.cpp · Serve
Last verified 2026-10-10
badgr serve --runtime llama.cpp --hf-repo Qwen/Qwen3-8B-GGUF --hf-file Qwen3-8B-Q4_K_M.gguf --gpu RTX_4090 --max-cost 0.8 --startup-timeout 20 --env LLAMA_ARG_N_GPU_LAYERS=99 --env LLAMA_ARG_N_PARALLEL=8 --env LLAMA_ARG_KV_UNIFIED=1 --env LLAMA_ARG_CTX_SIZE=20480 --env LLAMA_ARG_CACHE_TYPE_K=q8_0 --env LLAMA_ARG_CACHE_TYPE_V=q8_0 --env LLAMA_ARG_FLASH_ATTN=on- GPU
- NVIDIA GeForce RTX 4090Reported by the provider, not confirmed inside the container
- Time to verified
- 154s
- Deployment
- dep-d95b6fbdc9
- ✓ llama-server loaded Qwen3-8B Q4_K_M with 8 parallel slots and a unified KV cache
- ✓ final_state = VERIFIED
- ✓ 7 streams were generating while new requests were timed: the worst new request waited 4.0 s and 2.55 s in two samples (median 0.73 s)
- ✓ Teardown completed
Default configuration for llama.cpp#30252: with idle-slot prompt caching on, a new request can stall for seconds while other streams are generating. Measured on an RTX 4090 over the public endpoint, two samples of 8 requests each.
GitHub issue this run was for: ggml-org/llama.cpp#30252
Raw evidence
final_state=VERIFIED failure_class=None providers_tried=1/1 teardown_confirmed=False spend_usd=0.0186
Run from Badgr's local development environment. Documents that this deployment path works end-to-end for this model.
✓ Verified by Badgr · llama.cpp · Serve
Last verified 2026-10-10
badgr serve --runtime llama.cpp --hf-repo Qwen/Qwen3-8B-GGUF --hf-file Qwen3-8B-Q4_K_M.gguf --gpu RTX_4090 --max-cost 0.8 --startup-timeout 20 --env LLAMA_ARG_N_GPU_LAYERS=99 --env LLAMA_ARG_N_PARALLEL=8 --env LLAMA_ARG_KV_UNIFIED=1 --env LLAMA_ARG_CTX_SIZE=20480 --env LLAMA_ARG_CACHE_TYPE_K=q8_0 --env LLAMA_ARG_CACHE_TYPE_V=q8_0 --env LLAMA_ARG_FLASH_ATTN=on --env LLAMA_ARG_CACHE_IDLE_SLOTS=0- GPU
- NVIDIA GeForce RTX 4090Reported by the provider, not confirmed inside the container
- Time to verified
- 128s
- Deployment
- dep-f7c7827944
- ✓ Same model and settings with idle-slot prompt caching switched off (LLAMA_ARG_CACHE_IDLE_SLOTS=0)
- ✓ final_state = VERIFIED
- ✓ With 7 streams generating, the worst new request waited 0.94 s and 0.87 s in two samples (median 0.85 s and 0.48 s)
- ✓ Teardown completed
Workaround for llama.cpp#30252. Same hardware and load as the default run: the worst wait fell from 4.0 s / 2.55 s to 0.94 s / 0.87 s. Two samples each, so this shows the stall disappears, not an exact speed-up.
GitHub issue this run was for: ggml-org/llama.cpp#30252
Raw evidence
final_state=VERIFIED failure_class=None providers_tried=1/1 teardown_confirmed=False spend_usd=0.0153
Run from Badgr's local development environment. Documents that this deployment path works end-to-end for this model.
Availability
Not yet Badgr AI API
Not currently validated
✓ Dedicated endpoint
Available through Badgr
✓ Custom GPU deployment
Available through Badgr
Deployment profile
- Architecture
- qwen3 (GGUF)
- Licence
- Apache-2.0
- Minimum VRAM
- ~20GB
- Context
- 40K tokens
Recommended: RTX 4090 24GB. VRAM is an estimate and increases with context, cache, concurrency, and runtime overhead.
Model identity
- Base model
- Qwen/Qwen3-8B
- Updated
- 1 years ago
- HuggingFace revision
- 7c41481f57
Popularity and trust
Downloads (last month)
551.3K
Likes
315
Spaces using this model
17
Files and formats
Weight formats
GGUF
Repository files
9
Chat template
Not detected
GPU deployment scenarios
Estimated from parameter count and quantisation. Will switch to “Verified by Badgr” once a real run backs a tier.
Testing
RTX 4090 24GB
Short context, low concurrency
Small production
L40S 48GB
8K context, 1-4 concurrent requests
Higher throughput
A100 80GB
32K context, continuous batching
Large production
RTX 4090 24GB
Higher concurrency
Estimated — multimodal inputs, long context, and concurrency all increase real VRAM use beyond this estimate.
Also runs well on L40S 48GB, A100 40GB. Available in United States, Europe.
Serving compatibility
| Runtime | Status | Notes |
|---|---|---|
| vLLM | Estimated | Used for Badgr dedicated endpoints |
| Transformers | Unknown | Basic fallback |
| llama.cpp | Verified by Badgr | Confirmed by a real badgr serve run (deployment dep-d95b6fbdc9) |
| SGLang | Unknown | Not evaluated |
| Ollama | Unknown | Not evaluated |
| Diffusers | Unknown | Not evaluated |
| PyTorch | Unknown | Not evaluated |
| TensorRT-LLM | Unknown | Conversion may be required |
| TGI | Unknown | Not evaluated |
Ready-to-run Badgr configurations
Quick test
badgr serve Qwen/Qwen3-8B-GGUF --max-cost 2Validated dedicated endpoint command
badgr serve Qwen/Qwen3-8B-GGUF --gpu RTX 4090 --max-cost 10Advanced configuration (defaults to automatic)
- Runtime
- Quantisation
- GPU and GPU count
- Context length
- Maximum concurrency
- Region
- Maximum hourly spend
- Persistent or capped runtime
Estimated pricing
Provider estimate: cold start
2–5 min
Provider estimate: endpoint cost
$0.19 – $0.45/hr
Idle cost
$0 when stopped
Per-token throughput cost is not shown here: Badgr has not benchmarked this model yet, and this page does not display figures it cannot back with real data.
Recommended Badgr routes
RTX 4090 24GB · United States
Availablefrom $0.19/hr
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: just now
RTX 4090 24GB · United States
Availablefrom $0.44/hr
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: just now
RTX 4090 24GB · United States
Availablefrom $0.45/hr
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: just now
RTX 4090 24GB · Europe
Estimatedstarting from $0.45/hr
Badgr estimated starting price
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: Live marketplace price not available yet
RTX 4090 24GB · Europe
Estimatedstarting from $0.45/hr
Badgr estimated starting price
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: Live marketplace price not available yet
RTX 4090 24GB · Europe
Estimatedstarting from $0.45/hr
Badgr estimated starting price
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: Live marketplace price not available yet
Known limitations
- Memory use increases with context length and concurrency.