Serve

Serve AI when you need it.

Start a model, keep it available for as long as needed, then stop it when you're done.

Automatic computer selection · Stable address · Health checks · Pay only while your model is running · Stop anytime

VLLM MODEL SERVINGSCALE PREVIEW
> Serve Qwen3-32B --duration 20m --max-cost 10
MODEL
Qwen/Qwen3-32B
MATCHED SETUP
GPUH100 80GB
Imagebadgr/vllm-cuda12.4
EnginevLLM 0.8.5
Context32,768
Model and GPU compatible
Model loaded successfully
Test inference verified
Address ready
TRAFFIC
Requests/min6,200
Response time540 ms
Health100%
Addresshttps://dep-k7f.aibadgr.com/v1
Cost $1.14 / $10Auto-stop in 18:41

Who it's for

Built for teams who need a model available on demand

Badgr chooses the computer, provides an address and tracks the cost.

vLLM and open-model serving teams

You want a specific open model available to your app on a matched GPU, image and engine, not whatever a shared API offers.

Start the model you chose and Badgr picks the compatible computer, image and engine for it.

Teams with temporary or bursty inference

You only need the model available for part of the week, and always-on capacity keeps billing even when idle.

Start it when you need it, watch cost while it runs, and stop it when you're done.

What you get

What Badgr adds to anything you serve

Start a model and Badgr handles the computer, the address, the health checks and the cost.

Automatic computer selection

Choose a model and Badgr recommends the right computer for it automatically.

A stable address

Badgr gives what you started an address your app can call.

Health checks

Badgr checks the model answers correctly before telling you it's live.

Cost you can see

Cost is tracked while it runs, and stops when you stop it.

How it works

How serving works

Three steps from picking a model to stopping it.

1

Choose a model

Pick a model from the catalogue and Badgr recommends the computer to run it on.

2

Start it

Badgr starts the model, checks it's healthy and gives you an address.

3

Stop when done

Stop it from your dashboard whenever you no longer need it.

Starting

Computer chosen

Live

Health check passed

Open

Address ready

Cost

Tracked while running

Stopped

Billing ends

What's running

What's running

Open it, check its logs, see what it's costing you, or stop it — anytime.

AI Badgr — what's runningLive
model
Qwen/Qwen2.5-7B-Instruct
address
https://dep-k7f.aibadgr.com/v1
health
passing
cost so far
$1.84 · $20.00 maximum

Open · View logs · View cost · Stop

View what's running →

Serve AI when you need it

Pick a model, start it, and stop it when you're done — Badgr handles the computer and the cost.