/#jobs
#Platform#About Us#Employers#AI Jobs
Get Started
Netherlands
2 days ago
Senior Site Reliability Engineer — Token Factory (Inference Platform) logo

Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether

Netherlands
2 days ago
Apply

Senior Site Reliability Engineer — Token Factory (Inference Platform)

Senior Site Reliability Engineer needed for a large-scale AI inference platform in the Netherlands. Requires deep Kubernetes, Terraform, Prometheus/Grafana, and Python/Bash expertise, with a focus on reliability, observability, and GPU-heavy workloads.

Core AIHybridFull-timeSeniorKubernetesPrometheus

Senior Site Reliability Engineer — Token Factory (Inference Platform)

Senior Site Reliability Engineer needed for a large-scale AI inference platform in the Netherlands. Requires deep Kubernetes, Terraform, Prometheus/Grafana, and Python/Bash expertise, with a focus on reliability, observability, and GPU-heavy workloads.

Apply
Core AIHybridFull-timeSeniorKubernetes

Salary

Not specified

Work Location

Netherlands, NL

Work Model

Hybrid

Employment Type

Full-time

Experience Level

Senior

Core Qualifications

Technical (Must-have)
KubernetesPrometheusGrafanaTerraformPythonBashDistributed SystemsSLOsInfrastructure as CodeGPU WorkloadsvLLMTritonRayMLOpsIncident Management
Soft Skills
TroubleshootingCollaborationProactive MindsetOwnershipContinuous ImprovementRoot Cause Analysis

Preferred Qualifications

Technical (Nice-to-have)
Model HostingAI InfrastructureMachine Learning Platforms

Key Responsibilities

  • •Own the reliability, performance, and observability of the inference platform and its supporting infrastructure.
  • •Design, implement, and continuously improve telemetry pipelines covering metrics, logs, and traces.
  • •Build monitoring and observability solutions capable of processing large volumes of production signals and converting them into actionable insights.
  • •Configure and optimize Kubernetes infrastructure for high availability, scalability, and efficient resource utilization.
  • •Tune Kubernetes autoscaling mechanisms to improve the efficiency and utilization of GPU resources.
  • •Develop and maintain Terraform modules and infrastructure-as-code patterns that embed resilience and reliability into new clusters and services.
  • •Design and improve request-routing, retry, and failure-handling mechanisms to minimize the impact of transient infrastructure or service failures.
  • •Develop automation and operational tooling to detect, isolate, and remediate incidents quickly.
  • •Create, maintain, and improve runbooks for incident response and operational procedures.
  • •Participate in production incident management, troubleshooting issues and restoring services within demanding reliability objectives.
  • •Lead or contribute to post-mortem processes and implement corrective actions to prevent recurring incidents.
  • •Define and improve reliability practices for high-throughput APIs, including alerting strategies and Service Level Objectives (SLOs).
  • •Investigate distributed-system failures and performance issues across infrastructure and application layers.
  • •Optimize systems from the kernel and infrastructure layer through to the application layer.
  • •Support and improve the operation of GPU-intensive inference workloads and accelerator-based infrastructure.
  • •Contribute to scaling the inference platform while balancing performance, reliability, and infrastructure costs.
  • •Collaborate closely with software engineers to incorporate reliability and operational excellence into product and platform development.
  • •Promote automation, self-healing capabilities, and engineering practices that reduce operational overhead and improve system resilience.
Site Reliability EngineerAI Inference PlatformKubernetesTerraformGPU InfrastructureObservabilityDistributed SystemsSeniorNetherlandsHybrid
/#jobs

Your gateway to a successful career. Show your growth. Be ready for your next step. Capture and seize the best opportunities.

  • Data
  • FAQ
  • Articles
  • AI Jobs
  • Platform
  • Employers
  • About Us
  • Legal
© 2026/#jobsAll rights reserved.

For queries/support, email jobs.support@slashhash.ai