Site Reliability & On-Call
Responsibilities:
- establish the company’s on-call program, defining service-level indicators and objectives, runbook standards, and a blameless postmortem culture
- implement Slack-integrated alerting tooling to streamline incident triage and reduce time-to-acknowledge
- design observability dashboards spanning deployment pipeline health, environment version tracking, and infrastructure performance
- lead adoption of application performance monitoring and software cataloging to clarify service ownership across the engineering organization
- partner with engineering teams to embed reliability practices into the regular development lifecycle