/#jobs
#Platform#About Us#Employers#AI Jobs
Get Started
Rotterdam
1 day ago
Senior Site Reliability Engineer (SRE & AI Platform Operations) logo

Senior Site Reliability Engineer (SRE & AI Platform Operations)

HelloPrint

Rotterdam
1 day ago
Apply

Senior Site Reliability Engineer (SRE & AI Platform Operations)

HelloPrint is seeking a Senior Site Reliability Engineer to take full technical ownership of production reliability, distributed observability, deployment safety, cost optimization, and AI runtime infrastructure. The role involves defining SLOs, evolving CI/CD pipelines, managing AI platform operations, and driving FinOps practices across Google Cloud. Candidates should have extensive GCP experience, strong diagnostic skills, and a FinOps mindset.

Core AIOn-siteFull-timeSeniorSLOsSLIs

Senior Site Reliability Engineer (SRE & AI Platform Operations)

HelloPrint is seeking a Senior Site Reliability Engineer to take full technical ownership of production reliability, distributed observability, deployment safety, cost optimization, and AI runtime infrastructure. The role involves defining SLOs, evolving CI/CD pipelines, managing AI platform operations, and driving FinOps practices across Google Cloud. Candidates should have extensive GCP experience, strong diagnostic skills, and a FinOps mindset.

Apply
Core AIOn-siteFull-timeSeniorSLOs

Salary

Not specified

Work Location

Rotterdam, South Holland, Netherlands, NL

Work Model

On-site

Employment Type

Full-time

Experience Level

Senior

Core Qualifications

Technical (Must-have)
SLOsSLIsError BudgetsDistributed TracingSentryGoogle Cloud MonitoringCI/CDGitHub ActionsCanary DeploymentsAI OperationsFinOpsTerraformGoogle Cloud RunIAMSecret ManagerLinuxRedisPythonTypeScriptJavaScriptPHPLaravelLLMVector DatabasesIncident ManagementPost-MortemsCapacity PlanningDisaster RecoveryInfrastructure as CodeAutomation
Soft Skills
Pragmatic builderHigh technical agencyProblem-solvingOwnershipCollaboration

Key Responsibilities

  • •Define, track, and enforce SLOs, SLIs, and error-budget policies across core customer journeys and critical services.
  • •Expand distributed observability and telemetry across microservices, Laravel Horizon queue workers, and Google Cloud infrastructure.
  • •Evolve CI/CD pipelines with canary traffic shifting, automated SLO-driven rollbacks, and automated health gates in GitHub Actions.
  • •Architect, monitor, and scale runtime infrastructure supporting AI agents, semantic pipelines, and background automation.
  • •Take ownership of runtime cost control, model and token budget tracking, latency profiles, rate limits, queue backpressure, and provider availability.
  • •Drive continuous FinOps practices across Google Cloud workloads, optimizing compute/storage footprint and enforcing budget guardrails.
  • •Lead on-call incident response and blameless post-mortems, turning root causes into automated tests and guardrails.
  • •Drive capacity forecasting, dependency isolation, automated load testing, and disaster recovery validations against RTO/RPO targets.
  • •Own declarative infrastructure workflows using Terraform and Google Cloud Run, ensuring strict IAM least privilege and Secret Manager.
  • •Build internal tooling, runbooks, and self-service deployment primitives to eliminate firefighting.
Site Reliability EngineeringSREAI Platform OperationsGoogle Cloud PlatformTerraformFinOpsSLOCI/CDLaravelFull-time
/#jobs

Your gateway to a successful career. Show your growth. Be ready for your next step. Capture and seize the best opportunities.

  • Data
  • FAQ
  • Articles
  • AI Jobs
  • Platform
  • Employers
  • About Us
  • Legal
© 2026/#jobsAll rights reserved.

For queries/support, email jobs.support@slashhash.ai