Site Reliability Engineer

    United States
    Full-Time
    Mid (3-6 yrs)
    Engineering & Development
    Posted on June 10, 2026
    Role Summary
    As a Site Reliability Engineer at Filevine, you will improve the reliability, scalability, and
    operational maturity of the Filevine platform. You’ll design automation that reduces toil,
    strengthen observability, support reliable deployments at scale, and solve production challenges
    that keep Filevine running for legal teams across the country. This role is built for engineers
    who apply software engineering principles to infrastructure problems, thrive on complex
    technical challenges, and are energized by taking ownership of the systems they build and
    operate — growing into deeper expertise as their platform knowledge expands
    Responsibilities
    Design, build, and maintain the monitoring, logging, distributed tracing, dashboards, and
    alerting that give teams meaningful visibility into production health.
     Build automation, tooling, and CI/CD improvements that increase engineering efficiency,
    reduce toil, and support reliable deployments at scale.
     Design, implement, and maintain reliable systems for building, deploying, testing, and
    operating Filevine products — proactively identifying and resolving reliability,
    performance, scalability, and security risks before they impact customers.
     Participate in a shared 24/7 on-call rotation, using operational insights to drive automation
    and long-term reliability improvements; continuously improve runbooks, documentation,
    and engineering standards.
     Take ownership of technical initiatives from design through implementation, develop deep
    expertise in critical areas of the Filevine platform, and communicate clearly with technical
    and business stakeholders.
    What we are looking for
    4+ years of hands-on experience in software engineering, cloud infrastructure, platform
    engineering, DevOps, or related technical roles, including at least 2 years in a Site Reliability
    Engineering or reliability-focused role.
    -Working knowledge of distributed systems and how applications, infrastructure, and cloud
    services interact in production; demonstrated ability to troubleshoot production issues,
    perform root cause analysis, and drive long-term reliability improvements.
    -Proficiency with Python, Bash, or similar scripting languages; experience building
    production tooling, automation, or CI/CD pipelines and deployment automation.
    -Hands-on experience operating Kubernetes-based workloads and cloud infrastructure in
    AWS or a comparable platform, including compute, container orchestration, networking,
    IAM, object storage, and cloud-native monitoring.
    -Experience with Infrastructure as Code tools such as Terraform, Pulumi, or AWS
    CloudFormation, and familiarity with modern observability practices including monitoring,
    logging, alerting, distributed tracing, and incident response.
    -Experience using AI-assisted engineering tools to improve productivity, accelerate
    troubleshooting, or automate operational tasks; curiosity, ownership, and a passion for
    building reliable systems through continuous improvement.
    -Strong written and verbal communication skills; Bachelor’s degree in Computer Science,
    Information Systems, or a related field, equivalent industry certifications, or comparable
    professional experience.

    Company:  Filevine

    Provides legal practice management software with AI capabilities for law firms and legal professionals.
    501-1000 employees
    Software & IT Services
    HQ: United States