Issue evidence · Serve · NVIDIA/TensorRT-LLM

Ray orchestrator with attention DP and TRTLLM_FORCE_COMM_METHOD=NCCL_EP: each rank accumulates thousands of listening TCP sockets during warmup

GitHub issue: https://github.com/NVIDIA/TensorRT-LLM/issues/20082

Not run

On one 8x H200 node with the Ray orchestrator, attention data parallelism and the NCCL_EP communication method, each rank opens thousands of listening TCP sockets during warmup.

What Badgr ran

Nothing was run. The reproduction needs 8x H200 (RunPod's H200 8-GPU node costs about $27 an hour) plus the TensorRT-LLM 1.3.0rc28 wheel and a local Qwen3-30B-A3B copy, which does not fit a $4 budget for one issue.

What came back

  • The reporter runs an AWS p5en.48xlarge node with TensorRT-LLM 1.3.0rc28, Ray 2.58 or 2.59, and NCCL 2.30.7.
  • A sibling issue on the same code path (TensorRT-LLM#20083, EP communicator built from MPI_COMM_WORLD) is also unrun for the same reason.

The check itself is short (count listening sockets per rank after warmup), so it is a good fit for a funded 8x H200 run.

Checked 2026-10-11 in Badgr’s local development environment.

← All issue evidence