The Calix platform enables Communication Service Providers (CSPs) of all sizes to transform and future-proof their businesses. Through real-time data, automation, and actionable insights delivered via Calix One — our cloud-first, AI-powered platform — CSPs can simplify operations, collapse cost, and accelerate innovation. Calix One brings together the automation of everything and the experience of one, empowering customers to deliver differentiated subscriber experiences while driving acquisition, loyalty, and revenue growth. This is the Calix mission: to enable CSPs of all sizes to simplify, innovate, and grow, strengthening both their businesses and the communities they serve.
We’re at the forefront of a once in a generational change in the broadband industry. Join us as we innovate, help our customers reach their potential, and connect underserved communities with unrivaled digital experiences.
Staff Engineer -Cloud Platform Engineer – Cloud Platforms
About the Role
At Calix Cloud Platform Engineering, our mission is to build and operate highly scalable, secure, and resilient cloud and telecom infrastructure platforms that power next-generation broadband services for service providers worldwide. We enable engineering teams to rapidly deliver innovative products while ensuring carrier-grade reliability, observability, automation, and operational excellence.
We are seeking an experienced Staff Engineer – Cloud Platform Engineering with strong expertise in cloud-native platforms, telecommunications Cloud infrastructure, networking technologies, and automation. The ideal candidate will possess deep knowledge of GCP, Kubernetes, CI/CD, network services, and broadband access technologies including OLT, ONT, routers, NATS, and HAProxy. This role will collaborate with development, network engineering, SRE, operations, and product teams to build highly available platforms supporting cloud applications at scale.
Key Responsibilities
Platform Architecture & Engineering
- Design, architect, and operate highly available, scalable, secure, and resilient cloud platform infrastructure.
- Drive platform modernization initiatives leveraging cloud-native technologies and automation frameworks.
- Define and implement platform reliability, observability, disaster recovery, and capacity management strategies.
- Troubleshoot end-to-end service delivery issues spanning cloud infrastructure, network services, OLT/ONT environments, and customer-facing applications.
- Collaborate with engineering teams to optimize service availability, latency, throughput, and operational efficiency.
- Support network protocols and technologies
Messaging, Traffic Management & Service Delivery
- Design, deploy, and operate NATS messaging infrastructure supporting distributed microservices architectures.
- Configure and optimize HA Proxy for load balancing, high availability, traffic routing, SSL termination, and performance optimization.
- Build highly scalable service communication frameworks utilizing event-driven architectures.
- Implement service discovery, fault tolerance, and resiliency patterns.
Cloud Infrastructure & Automation
- Design and manage Google Cloud Platform (GCP) services including:
- GKE
- Compute Engine
- Cloud Storage
- BigQuery
- Pub/Sub
- Cloud SQL
- Composer/Airflow
- Implement Infrastructure as Code (IaC) using Terraform and Terragrunt.
- Automate infrastructure provisioning, deployment, compliance, and operational workflows.
- Build reusable platform services that enable engineering teams to rapidly deploy applications.
CI/CD & DevOps
- Design and maintain enterprise-scale CI/CD pipelines.
- Automate software delivery processes using Jenkins, GitLab CI, GitHub Actions, or equivalent tools.
- Improve developer productivity through platform engineering best practices and self-service capabilities.
Reliability, Monitoring & Observability
- Implement and enhance monitoring, logging, tracing, and alerting solutions.
- Configure observability platforms utilizing:
- Prometheus
- Grafana
- ELK/OpenSearch
- VictoriaMetrics/VictoriaLogs
- GCP Monitoring
- Drive proactive incident prevention and operational excellence initiatives.
- Lead root cause analysis (RCA) efforts for critical platform and network incidents.
Cross Functional Leadership
- Partner with Product, Engineering, SRE, Network Operations, TAC, and Customer Success teams.
- Mentor engineers on cloud technologies, networking fundamentals, platform reliability, and telecom architectures.
- Develop technical roadmaps and drive strategic platform initiatives.
- Contribute to standards, best practices, design reviews, and architecture governance.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Telecommunications, or equivalent practical experience.
- 8+ years of experience designing and operating large-scale distributed systems in production environments.
- 5+ years of hands-on experience with public cloud platforms, preferably Google Cloud Platform (GCP).
- Strong experience with Kubernetes and containerized platforms.
- Deep understanding of networking concepts including:
- TCP/IP
- Routing & Switching
- NAT
- DNS
- DHCP
- Firewalls
- VPN Technologies
- Load Balancing
- Hands-on experience with HAProxy configuration, tuning, troubleshooting, and high availability deployments.
- Hands-on experience with NATS messaging systems and event-driven architectures.
- Experience supporting telecom and broadband infrastructure including OLTs, ONTs, routers, and access network technologies.
- Experience implementing Infrastructure as Code using Terraform.
- Proficiency in Python, Go, Bash, or similar scripting/programming languages.
- Experience with CI/CD pipelines and DevOps platforms.
- Experience with monitoring, telemetry, and observability frameworks.
- Strong troubleshooting, problem-solving, and incident management skills.
Preferred Qualifications
- Experience in broadband access networks, PON technologies, and service provider environments.
- Experience with Kubernetes networking and service mesh technologies.
- Experience with Kafka, Redis, Elasticsearch/OpenSearch, and distributed databases.
- Knowledge of security best practices in cloud and telecom environments.
- Google Cloud Professional Certification(s).
- Experience working in SRE, Platform Engineering, or Telecom Cloud Operations teams.
The base pay range for this position varies based on the geographic location. More information about the pay range specific to candidate location and other factors will be shared during the recruitment process. Individual pay is determined based on location of residence and multiple factors, including job-related knowledge, skills and experience. This posting represents an active vacancy. Calix does not use automated or artificial intelligence tools to evaluate or select candidates.
Canada Locations:
141,000 - 240,000 CAD Annual
As a part of the total compensation package, this role may be eligible for a bonus. For information on our benefits click here.