🏢
Grafana
Staff Software Engineer - Databases SRE | UK | Remote
Job Description
ABOUT THE ROLE
We are seeking a highly skilled Staff Software Engineer - Databases SRE to join our team at Grafana Labs. As a Staff Software Engineer, you will play a critical role in ensuring the reliability of our Grafana Cloud databases, which are based on Mimir, Loki, Tempo, and Pyroscope. You will partner closely with product engineering squads to design and implement automation that scales our reliability practices, ensuring our customers meet our Service Level Objective (SLO) targets.
WHAT YOU'LL DO
As a Staff Software Engineer - Databases SRE, your primary responsibilities will include:
- Partnering closely with product engineering squads to ensure the reliability of our Grafana Cloud databases
- Owning production reliability for high-SLA and complex customer environments
- Designing and implementing automation to scale our reliability practices
- Defining and evolving per-tenant SLOs and reliability models
- Proactively reducing SLO burn to prevent repeat incidents
- Serving as a primary escalation point and on-call for relevant incidents
- Leading customer-impacting incident response and post-incident reviews
- Contributing to design docs and code reviews
- Influencing feature design to ensure production scalability and operability
- Building automation to eliminate toil where needed
- Improving alert quality and reducing noisy escalations
WHAT YOU'LL NEED
To succeed in this role, you will need:
- 8+ years of engineering experience, with 4+ years in SRE/CRE/production engineering
- Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.)
- Strong experience with technical leadership, leading a team through projects, mentoring other engineers on the team, and serving as a force-multiplier
- Experience operating multi-tenant systems in production
- Strong experience designing and implementing SLOs
- Experience with one or more programming languages (e.g. Go, Python, Java, etc)
- Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling
- Excellent problem-solving and troubleshooting skills
- Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high-quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents)
- Ability to reason about performance, scaling, and failure modes
- Comfortable working within an engineering team where individuals are encouraged to have a strong sense of autonomy and self-direction
- Ability to partner deeply with product engineering teams
WHY REMOTE
As a 100% remote company, we offer the flexibility and autonomy to work from anywhere in the world. Our team members are spread across 40+ countries, and we value diversity and inclusivity. We are committed to creating a culture that is open, collaborative, and supportive.
BENEFITS
- Competitive salary
- Access to modern AI coding assistants, such as GPT-Codex 5/3, Claude Opus 4.6, and Gemini 3 Pro
- Company-funded usage budget for AI tools
- Flexible schedule and remote work arrangement
- Comprehensive health plan
- Parental leave and