🏢
Gitlab
Site Reliability Engineer, Environment Automation
Job Description
ABOUT THE ROLE
We're seeking a highly skilled Site Reliability Engineer to join our Dedicated team at GitLab. As a Site Reliability Engineer, you will be focused on Environment Automation, where your work will help power hundreds of isolated GitLab environments for our customers. You will help keep these environments reliable, scalable, secure, and consistent by treating everything as code and contributing to automation across the entire lifecycle, from initial provisioning to day-to-day operations.
WHAT YOU'LL DO
As a Site Reliability Engineer, you will be responsible for defining, deploying, and maintaining GitLab environments across cloud providers using infrastructure as code, deployment packages, and Kubernetes. You will contribute to automation that reduces manual work, assist in building tooling that orchestrates upgrades and configuration changes safely at scale, and support an observability stack that lets us understand and improve the health of every environment. Your work will directly impact how customers experience GitLab Dedicated and other managed offerings, enabling them to focus on building software while we ensure their GitLab environments are always production ready.
Some examples of work you'll do:
- Contribute to the design and evolution of infrastructure automation using Terraform, Ansible, and Kubernetes to provision, upgrade, and operate many GitLab environments with minimal manual effort
- Help debug and resolve production issues across Kubernetes clusters, GitLab components, and cloud services, then assist in building automation and safeguards that prevent similar issues from recurring
- Assist in creating and maintaining deployment and orchestration tools, such as Helm Charts, omnibus-gitlab configurations, and multi-tenant workflows, that make it easy for teams to manage GitLab environments at scale
In addition to the examples above, your responsibilities will include:
- Contributing to automating operational tasks across many GitLab environments, from initial provisioning and configuration updates to upgrades and routine maintenance, helping reduce manual work and improve reliability at scale under the guidance of senior team members
- Helping build and refine the observability stack for multi-tenant GitLab environments so we monitor the right signals across Kubernetes, cloud services, and GitLab applications, supporting early issue detection and basic capacity tracking
- Assisting in responding to platform alerts and incidents, collaborating with Environment Automation SREs and engineering teams to troubleshoot production issues across multiple tenants and document findings
- Supporting planning and implementation of infrastructure changes, capacity expansions, and new service rollouts for Dedicated and other managed GitLab environments, contributing to efforts that improve resource efficiency and environment isolation
- Developing and maintaining scripts, automation tools, and infrastructure-as-code workflows that manage parts of the GitLab environment lifecycle, enabling more repeatable, self-service operations over time
- Applying and helping implement best practices for running GitLab on Kubernetes and cloud platforms, focusing on day-to-day reliability, performance, and security while learning how to keep environments consistent
- Participating in the on-call rotation for production GitLab environments with appropriate support, helping triage and mitigate incidents across clusters and cloud providers and contributing to post-incident reviews
- Documenting operational tasks, runbooks, and lessons learned so they become clear, repeatable processes and can be candidates for future automation, improving shared knowledge and reducing manual toil across the team
WHAT YOU'LL NEED