Netherlands
3 days ago
Senior Site Reliability Engineer — Token Factory (Inference Platform) logo

Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether

Senior Site Reliability Engineer — Token Factory (Inference Platform)

Senior Site Reliability Engineer needed for a large-scale AI inference platform in the Netherlands. Requires deep Kubernetes, Terraform, Prometheus/Grafana, and Python/Bash expertise, with a focus on reliability, observability, and GPU-heavy workloads.

Core AIHybridFull-timeSeniorKubernetesPrometheus

Salary

Not specified

Work Location

Netherlands, NL

Work Model

Hybrid

Employment Type

Full-time

Experience Level

Senior

Core Qualifications

Technical (Must-have)
KubernetesPrometheusGrafanaTerraformPythonBashDistributed SystemsSLOsInfrastructure as CodeGPU WorkloadsvLLMTritonRayMLOpsIncident Management
Soft Skills
TroubleshootingCollaborationProactive MindsetOwnershipContinuous ImprovementRoot Cause Analysis

Preferred Qualifications

Technical (Nice-to-have)
Model HostingAI InfrastructureMachine Learning Platforms

Key Responsibilities

  • Own the reliability, performance, and observability of the inference platform and its supporting infrastructure.
  • Design, implement, and continuously improve telemetry pipelines covering metrics, logs, and traces.
  • Build monitoring and observability solutions capable of processing large volumes of production signals and converting them into actionable insights.
  • Configure and optimize Kubernetes infrastructure for high availability, scalability, and efficient resource utilization.
  • Tune Kubernetes autoscaling mechanisms to improve the efficiency and utilization of GPU resources.
  • Develop and maintain Terraform modules and infrastructure-as-code patterns that embed resilience and reliability into new clusters and services.
  • Design and improve request-routing, retry, and failure-handling mechanisms to minimize the impact of transient infrastructure or service failures.
  • Develop automation and operational tooling to detect, isolate, and remediate incidents quickly.
  • Create, maintain, and improve runbooks for incident response and operational procedures.
  • Participate in production incident management, troubleshooting issues and restoring services within demanding reliability objectives.
  • Lead or contribute to post-mortem processes and implement corrective actions to prevent recurring incidents.
  • Define and improve reliability practices for high-throughput APIs, including alerting strategies and Service Level Objectives (SLOs).
  • Investigate distributed-system failures and performance issues across infrastructure and application layers.
  • Optimize systems from the kernel and infrastructure layer through to the application layer.
  • Support and improve the operation of GPU-intensive inference workloads and accelerator-based infrastructure.
  • Contribute to scaling the inference platform while balancing performance, reliability, and infrastructure costs.
  • Collaborate closely with software engineers to incorporate reliability and operational excellence into product and platform development.
  • Promote automation, self-healing capabilities, and engineering practices that reduce operational overhead and improve system resilience.
Site Reliability EngineerAI Inference PlatformKubernetesTerraformGPU InfrastructureObservabilityDistributed SystemsSeniorNetherlandsHybrid