
Senior Site Reliability Engineer
TOPdesk
Senior Site Reliability Engineer
TOPdesk is seeking a Senior Site Reliability Engineer to own the reliability of its Azure SaaS estate, setting SLOs, eliminating toil, and building self-healing automation. The role involves leading incident response, consolidating observability, and embedding AI-native practices. Candidates need 5+ years of SRE/DevOps experience with Azure, Kubernetes, Terraform, and strong observability skills.
Senior Site Reliability Engineer
TOPdesk is seeking a Senior Site Reliability Engineer to own the reliability of its Azure SaaS estate, setting SLOs, eliminating toil, and building self-healing automation. The role involves leading incident response, consolidating observability, and embedding AI-native practices. Candidates need 5+ years of SRE/DevOps experience with Azure, Kubernetes, Terraform, and strong observability skills.
Salary
Core Qualifications
Technical (Must-have)
Soft Skills
Preferred Qualifications
Technical (Nice-to-have)
Key Responsibilities
- Define and own SLOs and error budgets across the Azure estate
- Eliminate toil and implement self-healing automation
- Standardise observability (metrics, alerting, tracing) across datacenters
- Lead incident response and blameless postmortems
- Harden CI/CD and progressive delivery (canaries, safe rollouts, automated rollback)
- Model capacity and load-test critical paths
- Bring AI-native reliability (agents, bounded automation) into detection and remediation
- Maintain runbooks linked to automation candidates
- Own capacity and cost forecasting across the multi-cloud estate
- Collaborate proactively with product teams during design phases