
Technical Lead - GPU Infrastructure
Jobgether
Technical Lead - GPU Infrastructure
Technical Lead - GPU Infrastructure needed for a fully remote role based in Australia, responsible for architecting and delivering a large-scale GPU infrastructure platform. Requires deep hands-on expertise in Slurm, Kubernetes, NVIDIA GPU operations on bare metal, and proven technical leadership of distributed engineering teams. The role combines systems architecture, team management, and partner-facing technical ownership.
Technical Lead - GPU Infrastructure
Technical Lead - GPU Infrastructure needed for a fully remote role based in Australia, responsible for architecting and delivering a large-scale GPU infrastructure platform. Requires deep hands-on expertise in Slurm, Kubernetes, NVIDIA GPU operations on bare metal, and proven technical leadership of distributed engineering teams. The role combines systems architecture, team management, and partner-facing technical ownership.
Salary
Core Qualifications
Technical (Must-have)
Soft Skills
Preferred Qualifications
Technical (Nice-to-have)
Key Responsibilities
- Own the end-to-end platform architecture, including architecture proposals, high-level and low-level designs, technical reviews, and ongoing architecture documentation.
- Lead and line-manage a distributed engineering team across backend, frontend, DevOps, QA, and documentation.
- Establish engineering standards, oversee code and design reviews, manage release gates, conduct one-to-ones, and provide growth and performance feedback.
- Design, build, and operate a managed Slurm service supporting research and model-training workloads.
- Own Slurm controllers, accounting, partitions, login nodes, node onboarding, acceptance testing, driver and CUDA baselines, upgrades, stalled-job detection, node health, draining, autohealing, storage visibility, identity, and workload isolation.
- Lead GPU infrastructure operations on bare-metal environments, including NVIDIA drivers, CUDA, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance processes.
- Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, including NVIDIA GPU Operator and Network Operator.
- Oversee GPU isolation using technologies such as KubeVirt and VFIO and manage day-two infrastructure operations, upgrades, backup, recovery, and node replacement.
- Define managed inference architecture covering serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute capabilities.
- Establish observability across the control plane, GPU fleet, and application layers through metrics, logging, alerting, and SLOs.
- Lead incident response, post-incident reviews, and the development of an on-call model that is sustainable for a lean engineering organization.
- Act as the primary technical interface with infrastructure partners and vendors, translating requirements into written specifications and acceptance tests.
- Manage technical escalations with partners through resolution and contribute to capacity planning and hardware sourcing decisions.
- Work directly with research, model-training, and product teams to translate workloads into platform requirements and manage capacity constraints.
- Hire and develop members of the platform team while maintaining a high technical bar.
- Contribute to architecture decisions involving distributed systems, high-performance computing, networking, storage, virtualization, and GPU workloads.