💼Jobs 📝Blog 🧮Salary Calc 🌍Cost of Living 📋Tax Guide
👤Sign In Post a Job — from $99
Gitlab

Senior Site Reliability Engineer, Tenant Services: Geo

DevOpsFull-TimeSenior
Location
Europe
Job Type
Full-Time
Experience
Senior
Apply Now

Job Description

ABOUT THE ROLE We're seeking a skilled Site Reliability Engineer (SRE) to join our Tenant Services, Geo team at GitLab. As an SRE, you'll play a critical role in ensuring the smooth operation of user-facing services and production systems, applying sound engineering principles, operational discipline, and mature automation to our operating environments and the GitLab codebase. WHAT YOU'LL DO As a member of our team, you'll be responsible for executing Dedicated Geo migrations and cutovers end-to-end, including planning, pre-cutover validation, execution, and post-cutover verification and cleanup. You'll also join our team's shift and weekend coverage rotation for Dedicated cutovers across EMEA and US hours, and participate in the SaaS Site Reliability Engineering (SRE) on-call rotation to respond to incidents that impact GitLab.com availability. In addition, you'll operate and improve the Geo operational surface for Dedicated, including environment preparation and data hygiene checks prior to migrations, execution of replication, validation, and cutover procedures, and handling Geo-related escalations from Support and internal partners. You'll also design, build, and maintain automation, tooling, and runbooks that make migrations, cutovers, and Geo escalations as "boring" and repeatable as possible. WHAT YOU'LL NEED To succeed in this role, you'll need: - Experience operating highly-available distributed systems at scale, ideally in a SaaS environment with customer-facing SLAs - Hands-on experience with at least one major cloud provider (e.g., Google Cloud Platform or Amazon Web Services), including networking, storage, and managed services - Experience with Kubernetes and its ecosystem (e.g., Helm), including deploying and troubleshooting workloads - Experience with infrastructure as code and configuration management tools such as Terraform, Ansible, or Chef - Strong programming skills in at least one general-purpose language (preferably Go or Ruby) and proficiency with scripting (e.g., Shell, Python) - Experience with observability systems (e.g., Prometheus, Grafana, logging stacks) WHY REMOTE As a remote team, we value flexibility and autonomy. We're looking for someone who is self-motivated, can work independently, and is comfortable with remote communication and collaboration. BENEFITS - Competitive salary - Comprehensive benefits package, including medical, dental, and vision insurance - Generous PTO policy and flexible work hours - Opportunities for professional growth and development - Collaborative and dynamic work environment