🏢
Zuora
Site Reliability Engineer II
Job Description
ABOUT THE ROLE
We're seeking a highly skilled Site Reliability Engineer II to join our high-impact Operations team at Zuora. As a key member of our team, you will play a critical role in ensuring the reliability, scalability, and performance of our global production environment. You will be responsible for designing and implementing intelligent automation for infrastructure lifecycle management, applying AI/ML techniques for predictive monitoring and proactive performance optimization, and leading complex incident response efforts.
WHAT YOU'LL DO
As a Site Reliability Engineer II, you will have the opportunity to:
- Design and implement intelligent automation for infrastructure lifecycle management, including self-healing, anomaly detection, and automated remediation using Infrastructure as Code (IaC) and AI-driven tooling.
- Apply AI/ML techniques for predictive monitoring and proactive performance optimization to identify issues before they impact customers.
- Lead complex incident response efforts and root cause analyses, embedding automation and continuous learning into operational processes.
- Improve system reliability through dynamic scaling, telemetry instrumentation, and automated performance tuning.
- Enhance operational runbooks and playbooks by eliminating manual processes through automation.
- Evaluate and adopt emerging AIOps, cloud-native, and distributed systems technologies to continuously improve our platform.
- Partner cross-functionally with Product Engineering, Customer Support, Global Services, Deal Desk, and Sales to deliver exceptional customer experiences.
WHAT YOU'LL NEED
To be successful in this role, you will need:
- 2–4 years of experience in Linux systems administration and/or Python development in production environments.
- Strong Linux administration skills, including troubleshooting, service management, performance tuning, and networking fundamentals.
- Experience developing Python scripts or lightweight applications to automate operational workflows and system management.
- Hands-on experience with Docker and familiarity with Kubernetes concepts, including deployments, services, and scaling.
- At least one year of experience supporting SaaS or cloud-native production environments.
- Working knowledge of messaging platforms and databases such as Kafka, Redis, MySQL, or similar technologies.
- Experience contributing to CI/CD pipelines and deployment automation.
- Hands-on experience with monitoring and observability platforms such as Prometheus, Grafana, or similar tools.
- Experience participating in incident response, post-incident reviews, and root cause analysis.
- A demonstrated passion for automation and improving operational efficiency.
WHY REMOTE
As a remote employee, you will have the flexibility to work from anywhere while still being part of a collaborative and dynamic team. You will be able to work independently and manage your time effectively while also being able to communicate and collaborate with team members across different locations.
BENEFITS
Zuora offers a comprehensive total rewards package designed to support our employees' wellbeing, growth, and flexibility. This includes:
- Competitive compensation, variable bonus and performance-based reward opportunities, and retirement programs
- Medical, dental, and vision insurance
- Generous, flexible time off, plus paid holidays, wellness days, and a company-wide year-end break
- Paid parental leave (including fully paid leave for eligible employees, subject to location)