IBM · Granite 4.0
Granite 4.0 H Tiny
Verified by Badgr · ServeGranite 4.0 H Tiny is a 6.9B chat and text-generation model available through supported Badgr execution routes.
Last validated or meaningfully updated: 2026-10-09
✓ Verified by Badgr · SGLang · Serve
Last verified 2026-10-09
badgr serve --image lmsysorg/sglang:nightly-dev-20261005-f70e8c68 --gpu auto --port 30000 --health-path /health --max-cost 2 --startup-timeout 40 --cmd 'python3 -m sglang.launch_server --model-path ibm-granite/granite-4.0-h-tiny --host 0.0.0.0 --port 30000 --enable-deterministic-inference'- GPU
- NVIDIA GeForce RTX 3090Reported by the provider, not confirmed inside the container
- Time to verified
- 414.9s
- Deployment
- dep-6d3e5d79c9
- ✓ SGLang server started
- ✓ /health returned HTTP 200
- ✓ A repetition_penalty completion request returned text
- ✓ final_state = VERIFIED
- ✓ Teardown completed
Served with --enable-deterministic-inference. A completion request with repetition_penalty 1.06 (the trigger in sgl-project/sglang#43061) returned normally and the server stayed healthy on this image.
GitHub issue this run was for: sgl-project/sglang#43061
Raw evidence
final_state=VERIFIED failure_class=None providers_tried=1 spend_usd=0.0329 elapsed=414.9s
Run from Badgr's local development environment. Documents that this deployment path works end-to-end for this model.
Availability
Not yet Badgr AI API
Not currently validated
✓ Dedicated endpoint
Available through Badgr
✓ Custom GPU deployment
Available through Badgr
Deployment profile
- Architecture
- GraniteMoeHybridForCausalLM
- Licence
- Apache-2.0
- Minimum VRAM
- ~18GB
- Context
- 128K tokens
Recommended: RTX 4090 24GB. VRAM is an estimate and increases with context, cache, concurrency, and runtime overhead.
Model identity
- Updated
- 11 months ago
- HuggingFace revision
- 791e0d3d28
Popularity and trust
Downloads (last month)
52.4K
Likes
210
Spaces using this model
5
Files and formats
Weight formats
Safetensors
Repository files
15
Tokenizer
BPE
Chat template
Available
GPU deployment scenarios
Estimated from parameter count and quantisation. Will switch to “Verified by Badgr” once a real run backs a tier.
Testing
RTX 4090 24GB
Short context, low concurrency
Small production
L40S 48GB
8K context, 1-4 concurrent requests
Higher throughput
A100 80GB
32K context, continuous batching
Large production
RTX 4090 24GB
Higher concurrency
Estimated — multimodal inputs, long context, and concurrency all increase real VRAM use beyond this estimate.
Also runs well on L40S 48GB, A100 40GB. Available in United States, Europe.
Serving compatibility
| Runtime | Status | Notes |
|---|---|---|
| vLLM | Estimated | Used for Badgr dedicated endpoints |
| Transformers | Declared by source | Basic fallback |
| llama.cpp | Requires GGUF conversion | No GGUF weights found in repo |
| SGLang | Verified by Badgr | Confirmed by a real badgr serve run (deployment dep-6d3e5d79c9) |
| Ollama | Unknown | Not evaluated |
| Diffusers | Unknown | Not evaluated |
| PyTorch | Unknown | Not evaluated |
| TensorRT-LLM | Unknown | Conversion may be required |
| TGI | Unknown | Not evaluated |
Ready-to-run Badgr configurations
Quick test
badgr serve ibm-granite/granite-4.0-h-tiny --max-cost 2Validated dedicated endpoint command
badgr serve ibm-granite/granite-4.0-h-tiny --gpu RTX 4090 --max-cost 10Advanced configuration (defaults to automatic)
- Runtime
- Quantisation
- GPU and GPU count
- Context length
- Maximum concurrency
- Region
- Maximum hourly spend
- Persistent or capped runtime
Estimated pricing
Provider estimate: cold start
2–5 min
Provider estimate: endpoint cost
$0.18 – $0.45/hr
Idle cost
$0 when stopped
Per-token throughput cost is not shown here: Badgr has not benchmarked this model yet, and this page does not display figures it cannot back with real data.
Recommended Badgr routes
RTX 4090 24GB · United States
Availablefrom $0.18/hr
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: just now
RTX 4090 24GB · United States
Availablefrom $0.43/hr
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: just now
RTX 4090 24GB · United States
Availablefrom $0.44/hr
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: just now
RTX 4090 24GB · Europe
Estimatedstarting from $0.45/hr
Badgr estimated starting price
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: Live marketplace price not available yet
RTX 4090 24GB · Europe
Estimatedstarting from $0.45/hr
Badgr estimated starting price
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: Live marketplace price not available yet
RTX 4090 24GB · Europe
Estimatedstarting from $0.45/hr
Badgr estimated starting price
Best for
ComfyUI, Batch Inference, Qwen 2.5 7B Instruct
Startup: 2–5 min
Reliability: Standard
Last checked: Live marketplace price not available yet
Known limitations
- Memory use increases with context length and concurrency.