How to match your resume to a Site Reliability Engineer job description
We are looking for a Site Reliability Engineer to own reliability targets for production services. You will run incident response and blameless postmortems, and automate away repetitive operational work. Required: hands-on experience with SLI/SLO design, Incident management, Observability, Capacity planning, and working knowledge of Kubernetes, Datadog, PagerDuty. Preferred: CKA or Google Cloud Professional Cloud DevOps Engineer. Success in this role is measured by availability against SLO, MTTR, toil hours removed.
Required skills to mirror
- SLI/SLO design
- Incident management
- Observability
- Capacity planning
- Automation
- Linux internals
- Chaos testing
- Postmortems
Seniority ladder
- SRE
- Senior SRE
- Staff SRE
Title variants
- SRE
- Production Engineer