💼Jobs 📝Blog 🧮Salary Calc 🌍Cost of Living 📋Tax Guide
👤Sign In Post a Job — from $99
Klaviyo

Software Engineer II, Reliability

EngineeringFull-TimeMid-Level
Location
APAC
Job Type
Full-Time
Experience
Mid-Level
Apply Now

Job Description

ABOUT THE ROLE We are seeking a Software Engineer II, Reliability to join our team at Klaviyo. As a reliability-focused software engineer, you will play a critical role in ensuring the reliability, scalability, and sustainability of our critical platforms. You will work closely with our SRE team to build and operate foundational services and infrastructure, reduce operational toil through automation, and continuously improve systems based on real production learnings. WHAT YOU'LL DO As a Software Engineer II, Reliability, you will contribute to the reliability and operational excellence of Klaviyo's platforms by working on well-scoped projects and owning services with support from senior engineers. Your responsibilities will include: - Building, operating, and improving production systems with a focus on reliability, scalability, and performance - Applying software engineering principles to automate operational tasks and reduce manual toil - Contributing to the design and implementation of systems using established SRE best practices - Helping define and measure SLIs and SLOs for services you support - Improving observability through metrics, dashboards, logging, and tracing - Participating in on-call rotations and responding to production incidents with guidance and support - Assisting with incident investigation and contributing to post-incident reviews and follow-up actions - Performing basic analysis around system behavior, capacity usage, and scaling characteristics - Identifying reliability issues or operational pain points and working with teammates to address them - Collaborating with product, platform, and security engineers to ship reliable systems - Writing and maintaining clear operational runbooks and system documentation WHAT YOU'LL NEED To be successful in this role, you will need to have: - Experience operating cloud-native production systems and services - The ability to write production-quality code (e.g. Python, Go, or similar) to automate operations and improve reliability - Understanding of common failure modes in distributed systems, such as dependency failures, resource exhaustion, and partial outages - Experience working with containerized workloads and platforms (e.g. Kubernetes) in production environments - Comfort participating in on-call rotations and diagnosing straightforward production issues - Experience using observability tools and responding to alerts - Familiarity with SRE concepts such as SLIs, SLOs, and error budgets, and the ability to apply them in practice - Hands-on experience with infrastructure as code or declarative configuration (e.g. Terraform, Kubernetes manifests) - Ability to follow incident response processes and contribute meaningfully during outages - Comfort receiving feedback, learning from incidents, and improving systems over time WHY REMOTE As a remote employee at Klaviyo, you will have the flexibility to work from anywhere and will be supported in setting up your home workspace. We value people who take ownership, learn continuously, and collaborate openly, and we are committed to building an inclusive team. BENEFITS - Competitive salary and bonus structure - Comprehensive health plan and parental leave policy - Education stipend and life insurance - Flexible schedule and remote work arrangement - Opportunities for professional growth and development - Access to cutting-edge technologies and tools - Collaborative and inclusive team environment