Issue evidence · Serve · NVIDIA/TensorRT-LLM
GitHub issue: https://github.com/NVIDIA/TensorRT-LLM/issues/20082
Not run
On one 8x H200 node with the Ray orchestrator, attention data parallelism and the NCCL_EP communication method, each rank opens thousands of listening TCP sockets during warmup.
Nothing was run. The reproduction needs 8x H200 (RunPod's H200 8-GPU node costs about $27 an hour) plus the TensorRT-LLM 1.3.0rc28 wheel and a local Qwen3-30B-A3B copy, which does not fit a $4 budget for one issue.
The check itself is short (count listening sockets per rank after warmup), so it is a good fit for a funded 8x H200 run.
Checked 2026-10-11 in Badgr’s local development environment.
← All issue evidence