Site Reliability & On-Call

Responsibilities:

  • establish the company’s on-call program, defining service-level indicators and objectives, runbook standards, and a blameless postmortem culture
  • implement Slack-integrated alerting tooling to streamline incident triage and reduce time-to-acknowledge
  • design observability dashboards spanning deployment pipeline health, environment version tracking, and infrastructure performance
  • lead adoption of application performance monitoring and software cataloging to clarify service ownership across the engineering organization
  • partner with engineering teams to embed reliability practices into the regular development lifecycle

Updated: