Senior Site Reliability Engineer
1 week ago
Reddit is a community of communities. It's built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 116 million daily active unique visitors, Reddit is one of the internet's largest sources of information. For more information, visit www.redditinc.com . Reddit SRE is rapidly innovating and our teams are working to meet the needs of infrastructure and development teams as they evolve our product faster than ever before. This is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet. As a Senior Site Reliability Engineer on Reddit's Infrastructure SRE team, you'll use your knowledge of distributed systems and architecture to improve the reliability and performance of Reddit’s engineering platforms and services. We are looking for someone who thrives at the intersection of infrastructure and software development. This team will work very closely with the Compute, Traffic, and Observability infrastructure teams. They will own a suite of tools for allowing engineers to understand their creations, based primarily on open-source solutions at scale. We’re active users of and contributors to Prometheus, Thanos, Grafana, Vector and more. In this role, you will also take ownership of risk management, ensuring the reliability and performance of our systems. You will collaborate with cross-functional teams to identify, assess, and mitigate risks, implementing best practices to enhance system resilience. Your expertise will drive proactive measures to maintain uptime and optimize service delivery, making a significant impact on our operational excellence. Join us and help build the future of Reddit Responsibilities: Advise Work closely with engineering teams in designing and developing systems that are resilient and highly performant at a tremendous scale, and maintaining the foundational platform for running Reddit’s infrastructure. Amplify Identify and build capabilities into our foundational Infrastructure and Platform services, which are used by Reddit engineering teams to build, deploy, and operate Reddit. Deliver software to improve the availability, scalability, latency, and efficiency of observability components. Identify and engineer away risk across Reddit’s systems. Automate Take repetitive, manual, or risky tasks and automate them out of existence. Build tools and integrate systems to support Reddit’s evolution. Automate critical aspects of the event driven development process Diagnose Draw on your knowledge of distributed systems to identify and fix network, system, and service-level issues. Practice sustainable incident response, and drive structural improvement with blameless postmortem. Share on-call responsibilities. Optimize Observe and improve performance, reduce cost, and improve the experience for millions of users Contribute upstream changes to the open source projects we use Qualifications 5+ years of experience in Software Engineering, Site Reliability Engineering, or a development-focused DevOps role. Proficiency in one or more programming languages. We’re predominantly writing code in Go and Python. Experience with Kubernetes and Cloud systems. Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki) Experience with the development and operation of high-traffic backend systems. A demonstrated ability to debug, fix, and optimize code. Troubleshooting skills that span applications, networking (TCP/IP), and systems. Strong working knowledge of Linux and containers. Excellent communication and collaborative skills. Pension Scheme Private Medical and Dental Scheme Life Assurance, Income Protection Workspace benefit for your home office Family Planning Support Flexible Vacation & Reddit Global Days Off In select roles and locations, the interviews will be recorded, transcribed and summarized by artificial intelligence (AI). You will have the opportunity to opt out of recording, transcription and summarization prior to any scheduled interviews. During the interview, we will collect the following categories of personal information: Identifiers, Professional and Employment-Related Information, Sensory Information (audio/video recording), and any other categories of personal information you choose to share with us. We will use this information to evaluate your application for employment or an independent contractor role, as applicable. We will not sell your personal information or disclose it to any third party for their marketing purposes. We will delete any recording of your interview promptly after making a hiring decision. For more information about how we will handle your personal information, including our retention of it, please refer to our Candidate Privacy Policy for Potential Employees and Contractors . Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve. Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know. #J-18808-Ljbffr
-
Senior Site Reliability Engineer
2 weeks ago
Greater London, United Kingdom Stratospherec Ltd Full timeOverview Senior DevOps Engineer / Senior Site Reliability Engineer Fully Remote working for candidates based in the UK – Salary to £90k + Benefits We are looking for a Senior DevOps Engineer that has strong C# code knowledge combined with strong knowledge of DevOps tools like Kubernetes (EKS or AKS) and Azure or AWS Cloud platforms. We are looking for a...
-
Senior Site Reliability Engineer
1 week ago
Greater London, United Kingdom Cryptio Full timeAbout Cryptio We’re Cryptio. We build infrastructure to bring financial integrity to the crypto economy. Our enterprise-grade back‑office and data platform power mission‑critical accounting, reporting, and operational workflows for institutions, corporates, and crypto‑native organisations. We’re trusted by leaders like Circle, Societe Generale,...
-
Senior Site Reliability Engineer
3 weeks ago
London, United Kingdom Stratospherec Ltd Full timeSenior DevOps Engineer / Senior Site Reliability Engineer Fully Remote working for candidates based in the UK – Salary to £90k + Benefits We are looking for a Senior DevOps Engineer that has strong C# code knowledge combined with strong knowledge of DevOps tools like Kubernetes (EKS or AKS) and Azure or AWS Cloud platforms. We are looking for a DevOps...
-
Senior Site Reliability Engineer
1 week ago
Greater London, United Kingdom Next Matter Full timeLocationUK Employment TypeFull time Location TypeRemote DepartmentEngineering About Cryptio We’re Cryptio. We build infrastructure to bring financial integrity to the crypto economy. Our enterprise‑grade back‑office and data platform power mission‑critical accounting, reporting, and operational workflows for institutions, corporates, and...
-
Senior Site Reliability Engineer
4 days ago
London, United Kingdom Stratospherec Ltd Full timeSenior DevOps Engineer / Senior Site Reliability Engineer Fully Remote working for candidates based in the UK – Salary to £90k + Benefits We are looking for a Senior DevOps Engineer that has strong C# code knowledge combined with strong knowledge of DevOps tools like Kubernetes (EKS or AKS) and Azure or AWS Cloud platforms. We are looking for a DevOps...
-
Senior Senior Site Reliability Engineer
7 days ago
London, United Kingdom Jack & Jill Full timeHe's an AI agent that sends you unmissable jobs and then helps you ace the interview. He'll make sure you are considered for this role, and help you find others if you ask. Senior Site Reliability Engineer Company Description: Well-funded crypto finance platform Join a well-funded crypto finance platform as a Senior Site Reliability Engineer. You'll...
-
Site Reliability Engineer
2 weeks ago
Greater London, United Kingdom TP ICAP Full timeJoin to apply for the Site Reliability Engineer role at TP ICAP. The TP ICAP Group is a world leading provider of market infrastructure. Our purpose is to provide clients with access to global financial and commodities markets, improving price discovery, liquidity, and distribution of data, through responsible and innovative solutions. Through our people and...
-
Lead Site Reliability Engineer
4 days ago
Greater London, United Kingdom Blackfield Associates Full timeLead Site Reliability EngineerAre you ready to take your career to the next level in a role that’s critical to the reliability, scalability, and performance of cutting-edge systems? We’re on the lookout for a Lead Site Reliability Engineer to bring innovation, leadership, and technical excellence to our growing team.What You'll Do:Design and implement...
-
Senior Site Reliability Engineer
60 minutes ago
London, United Kingdom Leap29 Digital Full timeSenior Site Reliability Engineer (SRE) – Contract – UK-Based Are you a proven Site Reliability Engineer with a passion for driving operational excellence and helping teams mature their SRE practices? We are working with a highly regarded cloud consultancy on a short-term assignment supporting a major digital services project. This is a UK-based contract...
-
Senior Site Reliability Engineer
6 days ago
Greater London, United Kingdom Reddit Full timeReddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 116 million daily active unique visitors, Reddit is one...