Staff Site Reliability Engineer - Site Experience Job at Reddit, Remote

a1JVVDd6bFo3Y0JIK29FTW1Iai9HdU1EUmc9PQ==
  • Reddit
  • Remote

Job Description

About Reddit: Reddit is an internet-scale platform, a massive global community of communities built on shared interests, deep passions, and authentic open dialogue. Home to over 100,000+ highly active subreddits and serving approximately 126 million daily active unique visitors, Reddit represents one of the largest and most influential sources of real-time information, content curation, and consumer discussion boards on the modern internet.

Position Overview

We are seeking a highly technical, high-agency Staff Site Reliability Engineer to lead reliability engineering initiatives for our critical user-facing systems at absolute internet scale. Sitting at the vital intersection of infrastructure, core product engineering, and user experience, our Site Experience SRE team ensures that every transaction across the web app, mobile platforms, APIs, feed generation loops, and real-time messaging engines remains blazing fast, highly resilient, and stable. In this technical leadership track, you will design architectural safety guards under massive global load, eliminate operational risk, eliminate toil via automation, and influence engineering culture across the entire organization.

Key Responsibilities

  • User Experience Reliability Leadership: Drive operational excellence, scalability, and latency improvements across Reddit’s most business-critical endpoints, search grids, and media delivery layers.
  • Architecture for Hyperscale: Partner with infrastructure and product groups to design highly available distributed networks, guiding architectural decisions around failover paths, cluster redundancy, graceful degradation, and global traffic engineering.
  • Systemic Risk Mitigation: Audit dependencies, microservices, and deployments to uncover systemic bottlenecks, building proactive mitigation playbooks that continuously reduce severe incident counts.
  • Toil Elimination & Automation: Build programmatic remediation tools, deployment safety rails, and reliability guardrails to replace manual or repetitive on-call operational work.
  • Blameless Incident Management: Lead complex, multi-team incident response actions across global outages, orchestrating blameless postmortems, root-cause diagnostics, and structural long-term code fixes.
  • Engineering Standards & Mentorship: Champion company-wide best practices for SLIs/SLOs, capacity management, and release engineering while providing technical leadership and mentorship to SRE and software engineering peers.

Required Skills & Qualifications

  • 8+ years of verified professional history operating as a Site Reliability Engineer, Infrastructure Engineer, or Systems Architect managing large-scale, high-traffic distributed systems.
  • Demonstrated history supporting high-throughput, user-facing production environments with exceptional availability thresholds.
  • Deep systems-level understanding of Linux operating systems, cloud-native container architectures, network routing, and distributed components.
  • Strong programming and scripting capability using systems languages, preferably Go or Python .
  • Advanced operational mastery of telemetry and observability layers, including distributed tracing, logging, structured alerts, and metric aggregation.
  • Location Context: 100% remote-first operational infrastructure flexibility open exclusively to qualified engineering leaders permanently based within the United Kingdom .

Preferred Strategic Indicators (Nice to Have)

  • Production experience orchestrating containers using Kubernetes and managing public cloud hyperscaler environments.
  • Familiarity with distributed infrastructure tools such as **Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra, or Redis**.
  • Prior history optimizing Content Delivery Networks (CDNs), edge reliability nodes, or global traffic management rules.
  • Active contributions to open-source software communities or a history of leading large-scale organizational reliability transformations.

What We Offer

  • The exceptional engineering canvas to shape the performance and availability of one of the internet’s most influential platforms.
  • Highly competitive UK compensation package supplemented by a Group Personal Pension Scheme with matching employer contributions.
  • Comprehensive private medical and dental healthcare schemes paired with income replacement security programs.
  • A remote-first workspace environment providing global lifestyle benefit credits, professional development budgets, and caregiving support.
  • Flexible vacation schedules, paid volunteer time off, and highly generous paid parental leave brackets.
  • Access to premium mental health resources, coaching support networks, and localized perks like the Bike to Work scheme.

Job Tags

Full time, Remote work, Flexible hours

Similar Jobs

Nike Inc.

Senior Payroll Analyst - Canada Job at Nike Inc.

Senior Payroll Analyst - CanadaWHO YOULL WORK WITHAs the Senior Payroll Analyst, you will report to the Americas Payroll Director. This team drives accurate, compliant payroll operations across the region. Youll partner closely with HR, Finance, Legal...

Delta-T Group Inc.

Paralegal / Legal Administrator Job at Delta-T Group Inc.

 ...Location: Bryn Mawr, PA 19010 Date Posted: 06/01/2026 Category: Administrative Education: Bachelor's Degree Delta-T Group is a nationwide provider of services to the social services, special-education, and behavioral fields for over 35 years. Our corporate legal... 

Pit Stop Stores

Assistant Store Manager Job at Pit Stop Stores

 ...Job Description About Pit Stop Stores: Pit Stop Stores located in Mobridge, SD is a busy convenience store. We pride ourselves on...  ...opening and closing duties as required Qualifications: Prior retail or customer service experience Strong communication and... 

Workforce

Box Truck Driver Job at Workforce

 ...Local CDL Class A Truck Driver Location: Erwin, NC Schedule: Tuesday Friday | 4:00 AM start (until route completion) Pay: $18.50 $20.00 per hour Please note that this position could be short term. Job Summary We are seeking a reliable and... 

Netsync Network Solutions

Deployment Intern - IT Summer Projects (Paid) Job at Netsync Network Solutions

 ...expected in a school settings. Minimum Qualifications/Technical and Education Requirements: Ability to pass a background, TX DPS fingerprint, and drug screen. Valid Texas Driver's License may be required for certain positions. This is a paid internship....