🏢
MongoDB
Site Reliability Engineer (Senior or Staff)
Job Description
ABOUT THE ROLE
We are seeking a Senior or Staff Site Reliability Engineer to join our Platform Engineering team at MongoDB. As a critical member of our team, you will be responsible for designing and maintaining our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This includes our multi-cloud-provider Kubernetes infrastructure, networking, load balancing, and observability and alerting systems.
WHAT YOU'LL DO
As a Senior or Staff Site Reliability Engineer, you will contribute to developing a world-class continuous deployment experience, enabling the rapid and reliable shipment of MongoDB products. Your responsibilities will include:
- Contributing to developing and maintaining our continuous delivery infrastructure, primarily composed of Argo Workflows and ArgoCD
- Providing tooling that enables clear system ownership and facilitates self-service onboarding for development teams
- Collaborating with other teams within Platform Engineering to ensure a consistent service-onboarding experience
- Providing internal support for our deployment systems, including answering questions and addressing issues
- Participating in a 24/7 on-call rotation to resolve issues involving the deployment infrastructure
- Contributing to open-source projects and engineering software-based approaches like Kubernetes operators to streamline processes
WHAT YOU'LL NEED
To be successful in this role, you will need:
- 6+ years of experience in software development and operating distributed systems
- Proficiency in Python, Go, or a similar language
- Proven experience building and operating large-scale continuous integration and continuous deployment (CI/CD) pipelines
- A customer-focused mindset and a value for efficiency in processes and operations
- Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market
- Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure
- Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing)
WHY REMOTE
As a remote employee, you will have the flexibility to work from anywhere and enjoy a better work-life balance. You will also have the opportunity to work with a global team of talented engineers and contribute to the development of a world-class continuous deployment experience.
BENEFITS
- Base salary range: $144,000 - $200,000 CAD
- Eligibility for equity, participation in the employee stock purchase program, flexible paid time off, 20 weeks fully-paid gender-neutral parental leave, fertility and adoption assistance, Registered Retirement Savings Plan (RRSP) with employer match, mental health counseling, backup child and elder care, and health, dental, and vision benefits offerings.