Issue evidence · Serve · ggml-org/llama.cpp

uloading model to a GH200 takes ages

GitHub issue: https://github.com/ggml-org/llama.cpp/issues/30322

Not run

Loading a 30B GGUF onto a GH200 takes hours, with one CPU thread at 100% and no GPU activity.

What Badgr ran

Nothing was run: GH200 (aarch64 Grace Hopper) is not a GPU Badgr's providers rent.

What came back

  • The report has no logs and no reproduction beyond the command line.

A timed load on a different GPU would not test the Grace Hopper memory path.

Checked 2026-10-11 in Badgr’s local development environment.

← All issue evidence