/#jobs
#Platform#About Us#Employers#AI Jobs
Get Started
Amsterdam
2 weeks ago
Staff Site Reliability Engineer — AI Platform logo

Staff Site Reliability Engineer — AI Platform

Manychat

Amsterdam
2 weeks ago
Apply

Staff Site Reliability Engineer — AI Platform

Manychat is hiring a Staff Site Reliability Engineer to own the reliability, performance, and cost of its AI infrastructure. The role involves designing and evolving the AI Gateway, building observability for AI systems, driving cost optimization, and scaling AI expertise across the organization. Candidates need 5+ years of SRE/platform engineering experience and hands-on experience operating LLM-backed systems in production.

Core AIOn-siteFull-timePrincipalSREPlatform Engineering

Staff Site Reliability Engineer — AI Platform

Manychat is hiring a Staff Site Reliability Engineer to own the reliability, performance, and cost of its AI infrastructure. The role involves designing and evolving the AI Gateway, building observability for AI systems, driving cost optimization, and scaling AI expertise across the organization. Candidates need 5+ years of SRE/platform engineering experience and hands-on experience operating LLM-backed systems in production.

Apply
Core AIOn-siteFull-timePrincipalSRE

Salary

Not specified

Work Location

Amsterdam, North Holland, Netherlands, NL

Work Model

On-site

Experience Required

5 years

Employment Type

Full-time

Experience Level

Staff-level

Core Qualifications

Technical (Must-have)
SREPlatform EngineeringInfrastructure EngineeringLLMAmazon BedrockAzure OpenAIAnthropicAWSKubernetesTerraformCI/CDPrometheusGrafanaOpenTelemetrySLO
Soft Skills
InfluenceCoachingAutonomyOwnership

Preferred Qualifications

Technical (Nice-to-have)
LiteLLMKong AI GatewayGPUvLLMTGITritonEval Pipelines

Key Responsibilities

  • •Own reliability and performance of AI infrastructure: AI Gateway, inference services, and integrations with Amazon Bedrock, Azure OpenAI, and other LLM providers.
  • •Design and evolve the AI Gateway: routing, failover between providers, rate limiting, caching, and guardrails.
  • •Build observability for AI systems: latency/throughput/error SLOs per model and provider, token-level metrics, quality and drift signals.
  • •Drive cost optimization and FinOps for AI workloads: per-feature cost visibility, model right-sizing, caching strategies, provider mix.
  • •Run capacity planning and incident response for inference services; write and improve runbooks and postmortems.
  • •Scale AI expertise across the org: set standards, review designs, and coach teams shipping LLM-backed features.
SREAI PlatformLLMStaff EngineerAmsterdamOn-siteFull-timeAWSKubernetesFinOps
/#jobs

Your gateway to a successful career. Show your growth. Be ready for your next step. Capture and seize the best opportunities.

  • Data
  • FAQ
  • Articles
  • AI Jobs
  • Platform
  • Employers
  • About Us
  • Legal
© 2026/#jobsAll rights reserved.

For queries/support, email jobs.support@slashhash.ai