/#jobs
#Platform#About Us#Employers#AI Jobs
Get Started
Australia
5 days ago
Technical Lead - GPU Infrastructure logo

Technical Lead - GPU Infrastructure

Jobgether

Australia
5 days ago
Apply

Technical Lead - GPU Infrastructure

Technical Lead - GPU Infrastructure needed for a fully remote role based in Australia, responsible for architecting and delivering a large-scale GPU infrastructure platform. Requires deep hands-on expertise in Slurm, Kubernetes, NVIDIA GPU operations on bare metal, and proven technical leadership of distributed engineering teams. The role combines systems architecture, team management, and partner-facing technical ownership.

Core AIRemoteFull-timeSeniorSlurmKubernetes

Technical Lead - GPU Infrastructure

Technical Lead - GPU Infrastructure needed for a fully remote role based in Australia, responsible for architecting and delivering a large-scale GPU infrastructure platform. Requires deep hands-on expertise in Slurm, Kubernetes, NVIDIA GPU operations on bare metal, and proven technical leadership of distributed engineering teams. The role combines systems architecture, team management, and partner-facing technical ownership.

Apply
Core AIRemoteFull-timeSeniorSlurm

Salary

Not specified

Work Location

Australia, AU

Work Model

Fully remote, working location between UTC and UTC+5:30, occasional travel to partner sites and team events

Experience Required

8 years

Employment Type

Full-time

Experience Level

8+ years hands-on engineering experience, including at least 3 years leading infrastructure platform teams

Core Qualifications

Technical (Must-have)
SlurmKubernetesNVIDIA GPUCUDAFabric ManagerNVSwitchDCGMMIGInfiniBandRDMASR-IOVNCCLLinuxJavaScriptNode.js
Soft Skills
Technical leadershipTeam managementCommunicationWritten communicationDecision-makingPeople leadershipCross-functional collaborationMentoringProblem solving

Preferred Qualifications

Technical (Nice-to-have)
SoperatorSlinkyKueueVolcanoKAIKubeflow TrainervLLMSGLangTensorRT-LLMKubeVirtKata ContainersQEMU/KVMFirecrackerIntel TDXAMD SEV-SNP

Key Responsibilities

  • •Own the end-to-end platform architecture, including architecture proposals, high-level and low-level designs, technical reviews, and ongoing architecture documentation.
  • •Lead and line-manage a distributed engineering team across backend, frontend, DevOps, QA, and documentation.
  • •Establish engineering standards, oversee code and design reviews, manage release gates, conduct one-to-ones, and provide growth and performance feedback.
  • •Design, build, and operate a managed Slurm service supporting research and model-training workloads.
  • •Own Slurm controllers, accounting, partitions, login nodes, node onboarding, acceptance testing, driver and CUDA baselines, upgrades, stalled-job detection, node health, draining, autohealing, storage visibility, identity, and workload isolation.
  • •Lead GPU infrastructure operations on bare-metal environments, including NVIDIA drivers, CUDA, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance processes.
  • •Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, including NVIDIA GPU Operator and Network Operator.
  • •Oversee GPU isolation using technologies such as KubeVirt and VFIO and manage day-two infrastructure operations, upgrades, backup, recovery, and node replacement.
  • •Define managed inference architecture covering serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute capabilities.
  • •Establish observability across the control plane, GPU fleet, and application layers through metrics, logging, alerting, and SLOs.
  • •Lead incident response, post-incident reviews, and the development of an on-call model that is sustainable for a lean engineering organization.
  • •Act as the primary technical interface with infrastructure partners and vendors, translating requirements into written specifications and acceptance tests.
  • •Manage technical escalations with partners through resolution and contribute to capacity planning and hardware sourcing decisions.
  • •Work directly with research, model-training, and product teams to translate workloads into platform requirements and manage capacity constraints.
  • •Hire and develop members of the platform team while maintaining a high technical bar.
  • •Contribute to architecture decisions involving distributed systems, high-performance computing, networking, storage, virtualization, and GPU workloads.
GPU InfrastructureTechnical LeadSlurmKubernetesHPCAI InfrastructureRemoteFull-timeBare MetalDistributed Systems
/#jobs

Your gateway to a successful career. Show your growth. Be ready for your next step. Capture and seize the best opportunities.

  • Data
  • FAQ
  • Articles
  • AI Jobs
  • Platform
  • Employers
  • About Us
  • Legal
© 2026/#jobsAll rights reserved.

For queries/support, email jobs.support@slashhash.ai