Vice President, Global Production Operations & Reliability

    ~$125,839 - $233,701Market Estimate
    United States
    Full-Time
    Senior (7+ yrs)
    Engineering & Development
    Leadership
    Posted on August 11, 2026
    Everbridge is seeking a Vice President, Global Production Operations & Reliability to lead the strategy, execution, and continuous improvement of our global production operations. Reporting to the Chief Technology Officer, this executive will be responsible for ensuring the reliability, scalability, security, and operational excellence of our cloud-native SaaS platform.
    Everbridge empowers organizations to keep people safe and operations running during critical events. Our Critical Event Management platform is trusted by enterprises, governments, healthcare providers, financial institutions, and public safety organizations around the world to deliver resilient, mission-critical services when they matter most.

    This role will lead global Site Reliability Engineering (SRE) and Development & Reliability Engineering (DRE) teams while driving operational excellence across production operations, incident management, change governance, disaster recovery, observability, and platform engineering. The successful candidate will play a key role in strengthening customer trust by delivering highly available, resilient services on a global scale.

    What you'll do:
    • Define and execute Everbridge's global production operations and reliability strategy.
    • Lead and develop high-performing global SRE and DRE teams responsible for platform reliability and operational engineering.
    • Own the operational excellence of Everbridge's AWS and Kubernetes-based cloud platform, ensuring scalability, resilience, security, and performance.
    • Establish best practices for service reliability, including Service Level Objectives (SLOs), error budgets, production readiness, capacity planning, observability, and operational automation.
    • Drive disciplined incident management, change governance, release management, and post-incident reviews to continuously improve platform stability and reduce operational risk.
    • Lead disaster recovery planning, resilience testing, business continuity initiatives, and operational readiness across global production environments.
    • Partner closely with Engineering, Product, Security, Customer Support, and Customer Success to embed reliability into the software development lifecycle and improve customer outcomes.
    • Champion automation, cloud-native engineering practices, and continuous improvement to enhance operational efficiency and platform performance.
    • Provide executive leadership and reporting on operational health, reliability metrics, customer-impacting incidents, and strategic initiatives.
    What you'll bring:
  1. 15+ years of experience in production operations, cloud infrastructure, platform engineering, SRE, or related technology leadership roles.
  2. Proven experience leading global production operations for large-scale, mission-critical SaaS or cloud platforms.
  3. Deep expertise in AWS, Kubernetes, cloud-native architectures, distributed systems, and high-availability environments.
  4. Strong understanding of SRE principles, incident management, observability, disaster recovery, change management, CI/CD, and Infrastructure as Code.
  5. Demonstrated success building and leading high-performing global engineering and operations teams.
  6. Excellent executive communication, stakeholder management, and cross-functional leadership skills.
  7. Preferred Qualifications
  8. Experience in mission-critical industries such as public safety, critical communications, healthcare, financial services, security, or enterprise SaaS.
  9. Familiarity with modern SRE practices, progressive delivery, platform engineering, service mesh, and cloud cost optimization.
  10. Knowledge of regulatory and operational frameworks including ISO 27001, SOC 2, NIST, or FedRAMP.
  11. Company:  Everbridge

    Provides enterprise software applications that automate organizations' operational response to critical events to keep people safe and businesses running
    1001-5000 employees
    Software & IT Services
    HQ: United States