
Staff Site Reliability Engineer — AI Platform
Manychat
Staff Site Reliability Engineer — AI Platform
Manychat is hiring a Staff Site Reliability Engineer to own the reliability, performance, and cost of its AI infrastructure. The role involves designing and evolving the AI Gateway, building observability for AI systems, driving cost optimization, and scaling AI expertise across the organization. Candidates need 5+ years of SRE/platform engineering experience and hands-on experience operating LLM-backed systems in production.
Staff Site Reliability Engineer — AI Platform
Manychat is hiring a Staff Site Reliability Engineer to own the reliability, performance, and cost of its AI infrastructure. The role involves designing and evolving the AI Gateway, building observability for AI systems, driving cost optimization, and scaling AI expertise across the organization. Candidates need 5+ years of SRE/platform engineering experience and hands-on experience operating LLM-backed systems in production.
Salary
Core Qualifications
Technical (Must-have)
Soft Skills
Preferred Qualifications
Technical (Nice-to-have)
Key Responsibilities
- Own reliability and performance of AI infrastructure: AI Gateway, inference services, and integrations with Amazon Bedrock, Azure OpenAI, and other LLM providers.
- Design and evolve the AI Gateway: routing, failover between providers, rate limiting, caching, and guardrails.
- Build observability for AI systems: latency/throughput/error SLOs per model and provider, token-level metrics, quality and drift signals.
- Drive cost optimization and FinOps for AI workloads: per-feature cost visibility, model right-sizing, caching strategies, provider mix.
- Run capacity planning and incident response for inference services; write and improve runbooks and postmortems.
- Scale AI expertise across the org: set standards, review designs, and coach teams shipping LLM-backed features.