Netherlands
2 days ago

Senior Site Reliability Engineer — Token Factory (Inference Platform)
Jobgether
Netherlands
2 days ago
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer needed for a large-scale AI inference platform in the Netherlands. Requires deep Kubernetes, Terraform, Prometheus/Grafana, and Python/Bash expertise, with a focus on reliability, observability, and GPU-heavy workloads.
Core AIHybridFull-timeSeniorKubernetesPrometheus
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer needed for a large-scale AI inference platform in the Netherlands. Requires deep Kubernetes, Terraform, Prometheus/Grafana, and Python/Bash expertise, with a focus on reliability, observability, and GPU-heavy workloads.
Core AIHybridFull-timeSeniorKubernetes
Salary
Not specified
Core Qualifications
Technical (Must-have)
KubernetesPrometheusGrafanaTerraformPythonBashDistributed SystemsSLOsInfrastructure as CodeGPU WorkloadsvLLMTritonRayMLOpsIncident Management
Soft Skills
TroubleshootingCollaborationProactive MindsetOwnershipContinuous ImprovementRoot Cause Analysis
Preferred Qualifications
Technical (Nice-to-have)
Model HostingAI InfrastructureMachine Learning Platforms
Key Responsibilities
- Own the reliability, performance, and observability of the inference platform and its supporting infrastructure.
- Design, implement, and continuously improve telemetry pipelines covering metrics, logs, and traces.
- Build monitoring and observability solutions capable of processing large volumes of production signals and converting them into actionable insights.
- Configure and optimize Kubernetes infrastructure for high availability, scalability, and efficient resource utilization.
- Tune Kubernetes autoscaling mechanisms to improve the efficiency and utilization of GPU resources.
- Develop and maintain Terraform modules and infrastructure-as-code patterns that embed resilience and reliability into new clusters and services.
- Design and improve request-routing, retry, and failure-handling mechanisms to minimize the impact of transient infrastructure or service failures.
- Develop automation and operational tooling to detect, isolate, and remediate incidents quickly.
- Create, maintain, and improve runbooks for incident response and operational procedures.
- Participate in production incident management, troubleshooting issues and restoring services within demanding reliability objectives.
- Lead or contribute to post-mortem processes and implement corrective actions to prevent recurring incidents.
- Define and improve reliability practices for high-throughput APIs, including alerting strategies and Service Level Objectives (SLOs).
- Investigate distributed-system failures and performance issues across infrastructure and application layers.
- Optimize systems from the kernel and infrastructure layer through to the application layer.
- Support and improve the operation of GPU-intensive inference workloads and accelerator-based infrastructure.
- Contribute to scaling the inference platform while balancing performance, reliability, and infrastructure costs.
- Collaborate closely with software engineers to incorporate reliability and operational excellence into product and platform development.
- Promote automation, self-healing capabilities, and engineering practices that reduce operational overhead and improve system resilience.
Site Reliability EngineerAI Inference PlatformKubernetesTerraformGPU InfrastructureObservabilityDistributed SystemsSeniorNetherlandsHybrid