From the source
H Company engineers investigated GPU memory leaks on Kubernetes nodes where VRAM appeared allocated without an owning process.
They traced the issue to containers that remained in RUNNING state after pod deletion, with threads stuck in D-state inside the FUSE kernel driver (`request_wait_answer`), blocking on a FUSE reply that never arrived.
The post documents the debugging methodology and identifies the root cause as a FUSE mount hang rather than a CUDA or NCCL issue.



