GPU Infrastructure Engineer
- Джерело:
- djinni.co
Що робити
- Design and operate distributed GPU training infrastructure
- Validate cluster topology, RDMA/InfiniBand, and NCCL performance
- Standardize Kubernetes or Slurm scheduling, GPU images, and software versions
- Diagnose issues across training workloads, networking, storage, and GPU hosts
- Build monitoring, benchmarks, runbooks, and reliability standards
Що очікуємо
- Location in Europe
- Production experience with multi-node GPU training infrastructure
- Strong Linux, containers, CUDA, and NVIDIA-GPU-stack knowledge
- Hands-on experience with NCCL and InfiniBand or RoCE/RDMA troubleshooting
- Deep experience with Kubernetes or Slurm
Схожі вакансії
З блогу Trackr
Усі статті →Знайдено через trackr.help/jobs · Канал: @trackrhelp · Бот для персональних сповіщень: @trackrhelpBot


