Architect & Drive the reliability, scalability, and performance of our multi-cloud provisioning platform across all production stacks.
Architect and implement end-to-end automation pipelines to eliminate manual intervention, actively identifying and reducing technical toil.
Define, monitor, and improve critical system health indicators (SLIs/SLOs), including latency, throughput, error rates, and capacity usage, making data-driven architectural recommendations.
Lead Collaboration with product and cross-functional engineering teams to embed reliability and security considerations early into the software development lifecycle (SDLC).
Own Incident Response Management: Design robust detection mechanisms, triage critical incidents, lead deep Root Cause Analysis (RCA), and implement long-term preventative engineering solutions.
Simplify Complex Systems: Continually audit platform operations to identify bottlenecks, eliminate single points of failure, and reduce structural complexity.