Weekend Site Reliability Engineer

    Global - Remote
    Full-Time
    Mid (3-6 yrs)
    Engineering & Development
    Posted on May 5, 2026

    What you’ll be doing

    • Work with a team of DevOps and DBA professionals; covering Saturday, Sunday and Monday (5 days in total with flexibility in your days off) as a Weekend SRE
    • Improve existing infrastructure and processes across the countries we’re deployed in, as well as streamlining processes to deploy to new countries in the future
    • Continuously improve Kubernetes platform stability and efficiency, with a focus on optimising resource utilisation, reducing costs, and streamlining environment provisioning through GitOps-first practices
    • Monitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and real user monitoring (RUM)
    • Own weekend on-call operations, triaging and responding to production incidents, performing root cause analysis, and driving post-incident reviews
    • Design and manage alert pipelines to ensure actionable signal quality, with attention to preventing alert fatigue, waterfall alerting, and notification flooding
    • Define and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-call prioritisation
    • Take ownership and responsibility for our cloud operation activities
    • Liaise with external security agencies for annual audits as well as perform our own internal security sweeps
    • Aid in reconfiguring existing architecture to allow for rapid deployments to new countries
    • Mentoring less experienced team members

    What you’ll bring

    • 3+ years DevOps / platform engineering experience
    • Must be based in Europe or Asia or LatAM
    • Experience independently leading the planning and deployment of a project
    • Experienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and production
    • Strong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and Helm being highly valued
    • Experience with Infrastructure-as-Code, particularly Terraform
    • Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus
    • Hands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and OpenTelemetry
    • Experience with real user monitoring (RUM), with familiarity in Grafana Faro or OpenTelemetry SDK instrumentation being a plus
    • Proven on-call and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actions
    • Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patterns
    • Experience defining SLIs and SLOs and using them to inform reliability work
    • Familiarity with service mesh concepts is a plus, as we are actively evaluating Cilium-based service mesh in non-production environments
    • Solid networking knowledge, especially the TCP / IP stack and HTTP protocol
    • Experience handling high HTTP request volumes and designing systems for high availability and high traffic environments
    • A strong understanding of cache, including CDN, HTTP cache, Redis / Memcached
    • Excellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous

    Our stack

    • Languages: Java / Spring Boot, Node.js, Python, JavaScript
    • Database: Aurora MySQL & PostgreSQL, MongoDB, MySQL Community
    • Cache: ElastiCache, Redis, Valkey
    • Messaging: Apache RocketMQ, AutoMQ, Kafka
    • Networking & Proxy: Nginx, Kong, Cilium, eBPF
    • Orchestration & GitOps: Docker, Kubernetes (EKS), ArgoCD, Helm
    • Computing & Storage: AWS EC2, VPC, AWS Lambda, EBS, S3
    • CI/CD: Jenkins, GitHub Actions
    • Metrics: Prometheus, Mimir, Grafana, Alertmanager
    • Logs: Loki, Vector
    • Traces: Tempo, OpenTelemetry, Alloy
    • Profiling: Pyroscope
    • RUM: Grafana Faro, OpenTelemetry SDK
    • Infrastructure as Code: Terraform
    • CDN & Edge: Cloudflare, AWS CloudFront
    • AWS CloudWatch

    What’s in it for you

    • Sporty is a remote first company in pursuit of sustainability
    • A competitive salary + individual performance based bonuses every quarter
    • 28 days paid annual leave
    • Our core working hours are 10am-3pm in your local time zone with flexibility outside of this
    • Referral bonuses & flash bonuses
    • Top of the line equipment
    • Annual company retreats to provide great internal networking opportunities

    Interview process

    • Remote video screening with our Talent Acquisition Team
    • Online assessment via Hackerrank
    • Remote video interview with 3 x Team Members (45 mins each, not separate days)

    If you're interested, we encourage you to apply! Every application is reviewed by a member of our team (AI is not used in our recruitment process), and we aim to respond within 48 hours.


    #LI-remote

    Company:  Sporty Group

    Operates a global sports media, gaming, fintech, and entertainment platform serving millions of daily active users across 10+ countries.
    Remote-First Company
    501-1000 employees
    Gaming & Entertainment
    HQ: United Kingdom