
Senior Site Reliability Engineer (SRE & AI Platform Operations)
HelloPrint
Senior Site Reliability Engineer (SRE & AI Platform Operations)
HelloPrint is seeking a Senior Site Reliability Engineer to take full technical ownership of production reliability, distributed observability, deployment safety, cost optimization, and AI runtime infrastructure. The role involves defining SLOs, evolving CI/CD pipelines, managing AI platform operations, and driving FinOps practices across Google Cloud. Candidates should have extensive GCP experience, strong diagnostic skills, and a FinOps mindset.
Senior Site Reliability Engineer (SRE & AI Platform Operations)
HelloPrint is seeking a Senior Site Reliability Engineer to take full technical ownership of production reliability, distributed observability, deployment safety, cost optimization, and AI runtime infrastructure. The role involves defining SLOs, evolving CI/CD pipelines, managing AI platform operations, and driving FinOps practices across Google Cloud. Candidates should have extensive GCP experience, strong diagnostic skills, and a FinOps mindset.
Salary
Core Qualifications
Technical (Must-have)
Soft Skills
Key Responsibilities
- Define, track, and enforce SLOs, SLIs, and error-budget policies across core customer journeys and critical services.
- Expand distributed observability and telemetry across microservices, Laravel Horizon queue workers, and Google Cloud infrastructure.
- Evolve CI/CD pipelines with canary traffic shifting, automated SLO-driven rollbacks, and automated health gates in GitHub Actions.
- Architect, monitor, and scale runtime infrastructure supporting AI agents, semantic pipelines, and background automation.
- Take ownership of runtime cost control, model and token budget tracking, latency profiles, rate limits, queue backpressure, and provider availability.
- Drive continuous FinOps practices across Google Cloud workloads, optimizing compute/storage footprint and enforcing budget guardrails.
- Lead on-call incident response and blameless post-mortems, turning root causes into automated tests and guardrails.
- Drive capacity forecasting, dependency isolation, automated load testing, and disaster recovery validations against RTO/RPO targets.
- Own declarative infrastructure workflows using Terraform and Google Cloud Run, ensuring strict IAM least privilege and Secret Manager.
- Build internal tooling, runbooks, and self-service deployment primitives to eliminate firefighting.