Qwen · Qwen 3
Qwen 3 0.6B FP8
Verified by Badgr · RunQwen 3 0.6B FP8 is a 0.6B chat and text-generation model available through supported Badgr execution routes.
Last validated or meaningfully updated: 2026-10-08
✓ Verified by Badgr · vLLM · Run
Last verified 2026-10-08
badgr run --image vllm/vllm-openai:nightly --gpu RTX_5090 --max-cost 1 --max-runtime 20 --cmd 'VLLM_USE_DEEP_GEMM=0 python3 -c "
from vllm import LLM, SamplingParams
llm = LLM(\"Qwen/Qwen3-0.6B-FP8\", max_model_len=2048)
print(\"SMOKE 60261-workaround\", llm.generate([\"Hello\"], SamplingParams(max_tokens=8)))
"'- GPU
- NVIDIA GeForce RTX 5090Reported by the provider, not confirmed inside the container
- Command runtime
- 74s
- Deployment
- dep-902d628dee
- ✓ Workload container started
- ✓ Command exited with code 0
- ✓ Job finished (status = succeeded)
- ✓ Teardown completed
Block-FP8 model on Blackwell (sm_120) with DeepGEMM disabled.
GitHub issue this run was for: vllm-project/vllm#60261
Raw evidence
deployment=dep-902d628dee status=succeeded user_command_exit=0 command_runtime=74s
Run from Badgr's local development environment. Documents that this job works end-to-end for this model.
Availability
Not yet Badgr AI API
Not currently validated
Not yet Dedicated endpoint
Not currently validated
✓ Custom GPU deployment
Available through Badgr
Deployment profile
- Architecture
- Qwen3ForCausalLM
- Licence
- Apache-2.0
- Minimum VRAM
- ~6GB
- Context
- 40K tokens
Recommended: RTX 4090 24GB. VRAM is an estimate and increases with context, cache, concurrency, and runtime overhead.
Model identity
- Base model
- Qwen/Qwen3-0.6B
- Updated
- 1 years ago
- HuggingFace revision
- e5be080333
Popularity and trust
Downloads (last month)
234.6K
Likes
63
Spaces using this model
4
Files and formats
Weight formats
Safetensors
Repository files
10
Tokenizer
BPE
Chat template
Available
GPU deployment scenarios
Estimated from parameter count and quantisation. Will switch to “Verified by Badgr” once a real run backs a tier.
Testing
RTX 4090 24GB
Short context, low concurrency
Small production
L40S 48GB
8K context, 1-4 concurrent requests
Higher throughput
A100 80GB
32K context, continuous batching
Large production
RTX 4090 24GB
Higher concurrency
Estimated — multimodal inputs, long context, and concurrency all increase real VRAM use beyond this estimate.
Also runs well on L40S 48GB, A100 40GB. Available in United States, Europe.
Serving compatibility
| Runtime | Status | Notes |
|---|---|---|
| vLLM | Verified for Run | Confirmed by a real badgr run job (deployment dep-902d628dee); serving is not verified |
| Transformers | Declared by source | Basic fallback |
| llama.cpp | Requires GGUF conversion | No GGUF weights found in repo |
| SGLang | Unknown | Not evaluated |
| Ollama | Unknown | Not evaluated |
| Diffusers | Unknown | Not evaluated |
| PyTorch | Unknown | Not evaluated |
| TensorRT-LLM | Unknown | Conversion may be required |
| TGI | Unknown | Not evaluated |
Estimated pricing
Provider estimate: cold start
2–5 min
Provider estimate: endpoint cost
$0.19 – $0.45/hr
Idle cost
$0 when stopped
Per-token throughput cost is not shown here: Badgr has not benchmarked this model yet, and this page does not display figures it cannot back with real data.
Recommended Badgr routes
RTX 4090 24GB · United States
Availablefrom $0.19/hr
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: just now
RTX 4090 24GB · United States
Availablefrom $0.43/hr
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: just now
RTX 4090 24GB · United States
Availablefrom $0.44/hr
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: just now
RTX 4090 24GB · Europe
Estimatedstarting from $0.45/hr
Badgr estimated starting price
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: Live marketplace price not available yet
RTX 4090 24GB · Europe
Estimatedstarting from $0.45/hr
Badgr estimated starting price
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: Live marketplace price not available yet
RTX 4090 24GB · Europe
Estimatedstarting from $0.45/hr
Badgr estimated starting price
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: Live marketplace price not available yet
Known limitations
- Memory use increases with context length and concurrency.