Guide · llama.cpp · RTX 4090
OpenJev is Loop AI’s open decision model for browser and desktop agents. Instead of generating free text, you give it a page state, a question and a set of options, and it returns a probability for each option. That makes it a good fit for the “what should the agent do next?” step, where you want a structured, repeatable answer.
This guide serves the Q4_K_M GGUF build on a single rented RTX 4090 with one Badgr command, checks it with cURL, and then queries it from Python through Loop AI’s own helper/shim.py decision API. We are not affiliated with Loop AI; for benchmarks and intended use, see the model card.
Run one command. These are the settings from the model card (-ngl 999 -c 16384 -np 2):
badgr serve --runtime llama.cpp \ --hf-repo openjev/openjev-GGUF --hf-file OpenJev-Q4_K_M.gguf \ --gpu RTX_4090 --max-cost 1.5 \ --env LLAMA_ARG_N_GPU_LAYERS=999 --env LLAMA_ARG_CTX_SIZE=16384 --env LLAMA_ARG_N_PARALLEL=2
A GGUF repo needs --runtime llama.cpp and the exact file named. --max-cost 1.5 caps what the launch can spend.
Badgr rents the GPU, downloads the 16.5 GB file, starts llama.cpp, and reports the endpoint ready only after it has returned a real inference result. In our run that took 6 minutes (360 s) from launch to verified, and the start-up cost to that point was $0.044. If the CLI stops waiting before the download finishes, it prints the deployment id and the launch keeps going on the server. Find the endpoint URL with:
badgr status # ● running dep-xxxxxxxxxx endpoint RTX_4090 $0.44/hr # URL: https://<id>-8000.proxy.runpod.net
Save that URL and confirm the server is up:
export ENDPOINT_URL="https://<id>-8000.proxy.runpod.net" curl -s $ENDPOINT_URL/health
It should return {"status":"ok"}.
The endpoint is OpenAI-compatible. Send the prompt format from the model card: a state, a question, lettered options, and an instruction to answer with the letter only.
curl -s $ENDPOINT_URL/v1/chat/completions -H 'Content-Type: application/json' -d '{
"messages": [{"role": "user", "content": "State:\nA checkout page shows: Subtotal $40, Shipping $5, a Place order button, and a Coupon field.\n\nQuestion: Which action completes the purchase?\nOptions:\n[A] click_coupon\n[B] click_place_order\n[C] edit_shipping\n\nAnswer with the letter of the best option only."}],
"max_tokens": 4, "temperature": 0, "chat_template_kwargs": {"enable_thinking": false}}'The response content is B, the “Place order” option.
Letters are fine for a demo, but agents want probabilities. helper/shim.py from Loop AI’s openjev/openjevrepository exposes /v1/systemone on top of any OpenAI-compatible server and returns a probability for every option. Point it at your Badgr endpoint:
pip install openai transformers requests curl -O https://huggingface.co/openjev/openjev/raw/main/helper/shim.py VLLM=$ENDPOINT_URL/v1 python shim.py --port 8765
Then ask a question:
curl -s localhost:8765/v1/systemone -H 'Content-Type: application/json' -d '{
"state": "A checkout page shows: Subtotal $40, Shipping $5, a Place order button, and a Coupon field.",
"questions": {"q1": {"type": "choice", "instructions": "Which action completes the purchase?",
"criteria": {"click_coupon_field": "Open the coupon field", "click_place_order": "Click the Place order button", "edit_shipping": "Change the shipping option"}}}}'Real response from our run:
{
"id": "shim-1791636121448",
"answers": {
"q1": {
"type": "choice",
"choice": "click_place_order",
"probabilities": {"click_coupon_field": 0.0005, "click_place_order": 0.9994, "edit_shipping": 0.0001},
"confidence": 0.9991
}
},
"usage": {"input_tokens": 96, "output_tokens": 0}
}The shim supports three question types: choice (pick one option), noul (a yes/no probability) and score(a position on an ordered scale). This script asks one of each about a product review:
"""Ask OpenJev three kinds of question through Loop AI's helper/shim.py running against a Badgr endpoint."""
import json
import time
import requests
SHIM = "http://localhost:8765"
def wait_until_ready(timeout=300):
"""The shim answers /v1/version as soon as it is up."""
deadline = time.time() + timeout
while time.time() < deadline:
try:
if requests.get(f"{SHIM}/v1/version", timeout=5).ok:
return
except requests.RequestException:
pass
time.sleep(3)
raise TimeoutError("the shim did not become ready")
def ask(state, questions):
r = requests.post(f"{SHIM}/v1/systemone", json={"state": state, "questions": questions}, timeout=120)
r.raise_for_status()
return r.json()["answers"]
STATE = (
"Product review for noise-cancelling headphones, 4 stars: "
"'Great sound and the battery lasts two days. The ear cups get warm after an hour, and the case feels flimsy.'"
)
QUESTIONS = {
"sentiment": {
"type": "choice",
"instructions": "What is the overall sentiment of this review?",
"criteria": {"positive": "Mostly positive", "mixed": "Both positive and negative", "negative": "Mostly negative"},
},
"recommend": {
"type": "noul",
"instructions": "Would the reviewer recommend these headphones to a friend?",
"criteria": {"true": "The reviewer would recommend them.", "false": "The reviewer would not recommend them."},
},
"rating": {
"type": "score",
"instructions": "How satisfied is the reviewer with comfort?",
"criteria": ["Very unhappy", "Unhappy", "Neutral", "Happy", "Very happy"],
},
}
if __name__ == "__main__":
wait_until_ready()
started = time.time()
answers = ask(STATE, QUESTIONS)
print(json.dumps(answers, indent=2))
print(f"answered {len(answers)} questions in {time.time() - started:.1f}s")
Output from our run:
{
"sentiment": {"type": "choice", "choice": "mixed",
"probabilities": {"positive": 0.2665, "mixed": 0.7326, "negative": 0.0008}, "confidence": 0.599},
"recommend": {"type": "noul", "noul": 0.6628},
"rating": {"type": "score", "score": 1.4852,
"legend": {"0": "Very unhappy", "1": "Unhappy", "2": "Neutral", "3": "Happy", "4": "Very happy"},
"probabilities": {"0": 0.0123, "1": 0.4979, "2": 0.4827, "3": 0.0065, "4": 0.0006}, "confidence": 0.5752}
}
answered 3 questions in 2.4sAll three came back from one request in 2.4 s. The review is genuinely mixed, and the model says so: the sentiment choice is “mixed” at 0.73, the recommendation probability is 0.66, and the comfort score of 1.49 sits between “Unhappy” and “Neutral”. Our earlier smoke test of three hand-written checkout questions was answered correctly each time, at probabilities of 0.998 to 0.999. A handful of questions is a smoke test, not an accuracy measurement; see the model card for Loop AI’s numbers. We did not test the shim’s READOUT_TARGETED=1 mode, which relies on vLLM-specific options.
badgr down <deployment-id>
The GPU is released and billing stops. A whole test like this one, from launch to teardown, cost a few cents.