Site Reliability Engineer
Apply NowJob PurposeOperate as a hands-on site reliability engineer delivering after-hours incident detection, triage, and runbook-based remediation for production cloud-native environments, to support our North American customers, escalating cleanly when actions exceed agreed authority. Key Responsibilities• Monitor and respond to alerts; perform initial triage to determine severity, scope, and impact, to support our North American customers. • Execute approved runbooks: workload/node restarts, scaling within agreed bounds, rollbacks, and database stabilization. • Investigate and contain availability, performance, scaling, access, and replication issues within defined permissions. • Engage cloud-provider support (AWS / GCP) when required. • Prepare clear escalation summaries and shift-handoff notes. • Contribute to runbook improvements and flag automation and detection gaps. People Management • Individual contributor; no direct reports. Financial Responsibility • No budget responsibility. • Responsible for disciplined, in-scope operations that protect service levels. Key Performance Indicators (KPIs) • Service-level (SLO/SLA) attainment • Runbook-execution accuracy • Quality and timeliness of escalations and handoffs • Ticket hygiene and documentation • Contribution to runbook / automation improvements