Xianjun · Quokka
Quokka 7B Instruct
Quokka 7B Instruct is a 7B chat and text-generation model available through supported Badgr execution routes.
Last validated or meaningfully updated: 2026-07-27
✓ Verified by Badgr
Last verified 2026-10-01
badgr serve Xianjun/Quokka-7b-instruct --max-cost 1- Provider
- RunPod
- GPU
- NVIDIA L40S
- Time to verified
- 593.3s
- Deployment
- dep-cb6538f94e
- ✓ vLLM started
- ✓ 1-token /v1/completions request returned HTTP 200
- ✓ final_state = VERIFIED
- ✓ teardown completed
Real cross-provider fallback: 1 other GPU offer didn’t come up in time before this run landed on the NVIDIA L40S above and verified. The same automatic retry a live customer job gets.
Raw evidence
verify_command model=Xianjun/Quokka-7b-instruct ok=True final_state=VERIFIED failure_class=None provider=runpod providers_tried=2 spend_usd=0.0000 elapsed=593.3s
Run from Badgr's local development environment. Documents that this deployment path works end-to-end for this model.
Availability
✓ Badgr AI API
Available through Badgr
✓ Dedicated endpoint
Available through Badgr
✓ Custom GPU deployment
Available through Badgr
Deployment profile
- Architecture
- LlamaForCausalLM
- Licence
- Apache-2.0
- Minimum VRAM
- ~18GB
- Context
- 2K tokens
Recommended: RTX 4090 24GB. VRAM is an estimate and increases with context, cache, concurrency, and runtime overhead.
A real Badgr run for this model found no capacity on RTX 4090 24GB; it verified successfully on NVIDIA L40S instead (see “Verified by Badgr” above).
Model identity
- Updated
- 1 years ago
- HuggingFace revision
- ef3a0b9a9b
Popularity and trust
Downloads (last month)
49
Likes
1
Spaces using this model
0
Files and formats
Weight formats
Safetensors, PyTorch bin
Repository files
12
Tokenizer
SentencePiece
Chat template
Available
GPU deployment scenarios
Verified by Badgr for the tier marked below; other tiers remain estimated from parameter count and quantisation.
Testing
RTX 4090 24GB
Short context, low concurrency
Small production · Verified by Badgr
L40S 48GB
8K context, 1-4 concurrent requests
Higher throughput
A100 80GB
32K context, continuous batching
Large production
RTX 4090 24GB
Higher concurrency
Estimated — multimodal inputs, long context, and concurrency all increase real VRAM use beyond this estimate.
Also runs well on L40S 48GB, A100 40GB. Available in United States, Europe.
Serving compatibility
| Runtime | Status | Notes |
|---|---|---|
| vLLM | Verified by Badgr | Confirmed by a real badgr serve run (deployment dep-cb6538f94e) |
| Transformers | Declared by source | Basic fallback |
| llama.cpp | Requires GGUF conversion | No GGUF weights found in repo |
| SGLang | Unknown | Not evaluated |
| TensorRT-LLM | Unknown | Conversion may be required |
| TGI | Unknown | Not evaluated |
Call this model through Badgr
curl https://aibadgr.com/v1/chat/completions \
-H "Authorization: Bearer $BADGR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"Xianjun/Quokka-7b-instruct","messages":[{"role":"user","content":"Hello"}]}'Ready-to-run Badgr configurations
Quick test
badgr serve Xianjun/Quokka-7b-instruct --max-cost 2Dedicated endpoint command (auto GPU selection, see verified run above)
badgr serve Xianjun/Quokka-7b-instruct --max-cost 10Advanced configuration (defaults to automatic)
- Runtime
- Quantisation
- GPU and GPU count
- Context length
- Maximum concurrency
- Region
- Maximum hourly spend
- Persistent or capped runtime
Estimated pricing
Provider estimate: cold start
2–5 min
Provider estimate: endpoint cost
$0.44 – $0.47/hr
Idle cost
$0 when stopped
Per-token throughput cost is not shown here: Badgr has not benchmarked this model yet, and this page does not display figures it cannot back with real data.
Recommended Badgr routes
RTX 4090 24GB · United States
Availablefrom $0.44/hr
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: just now
RTX 4090 24GB · Europe
Estimatedstarting from $0.45/hr
Badgr estimated starting price
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: Live marketplace price not available yet
RTX 4090 24GB · Europe
Estimatedstarting from $0.45/hr
Badgr estimated starting price
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: Live marketplace price not available yet
RTX 4090 24GB · Europe
Estimatedstarting from $0.45/hr
Badgr estimated starting price
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: Live marketplace price not available yet
RTX 4090 24GB · United States
Availablefrom $0.47/hr
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: just now
RTX 4090 24GB · United States
Availablefrom $0.47/hr
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: just now
Known limitations
- Memory use increases with context length and concurrency.