🏢
MongoDB
Cloud Operations Engineer
Job Description
ABOUT THE ROLE
We are seeking a highly skilled Cloud Operations Engineer to join our global team at MongoDB. As a Cloud Operations Engineer, you will play a critical role in ensuring the success of our MongoDB Atlas customers worldwide. You will be responsible for the day-to-day operation of our cloud infrastructure, including creating and monitoring system alert dashboards, reviewing critical events and system logs, and performing server administration duties.
WHAT YOU'LL DO
As a Cloud Operations Engineer, your primary responsibilities will include:
- Coordinating and collaborating with a global team of Cloud Operations Engineers to ensure uptime guarantees to our Atlas customer base
- Helping to scale the worldwide Cloud Operations Engineering team by implementing and refining new processes and tools
- Assisting in scoping, designing, and deploying systems that reduce Mean Time to Resolve for customer incidents
- Monitoring and detecting emerging customer-facing incidents on the Atlas platform and assisting in their proactive resolution
- Automating routine monitoring and troubleshooting tasks
- Diagnosing live incidents, differentiating between platform issues versus usage issues, and taking the next steps toward resolution
- Assisting in performing root cause analysis after incident recovery and identifying areas for improvement
- Contributing to documentation of corner case scenarios, troubleshooting workflows, and SOPs
- Working alongside product management, cloud engineering, and support organizations to identify areas for improvement in the management applications powering the Atlas infrastructure
- Informing executive leadership and escalation management personnel of major outages
- Participating in a weekly on-call rotation, handling short-term customer incidents
WHAT YOU'LL NEED
To be successful in this role, you will need:
- At least 2 years of experience as an on-call DevOps, SRE, or Cloud Operations engineer
- Expertise in Linux system administration, configuration, and troubleshooting
- Experience in monitoring, system performance data collection and analysis, and reporting
- Knowledge of database operations and concepts
- Expertise in networking technologies like DNS, TCP/IP, etc.
- Familiarity with Amazon Web Services and other Cloud infrastructure platforms (e.g. GCP, Azure)
- Knowledgeable about a wide range of web and internet technologies
- Capability to write small programs/scripts to solve short-term systems problems
- A CS/CE degree or equivalent experience
- At least 1 of the following programming languages: Java, Go, Javascript
- A keen interest in learning new things
WHY REMOTE
As a remote Cloud Operations Engineer, you will have the flexibility to work from anywhere in the world. You will be part of a global team that works at the frontier of cloud services and database systems.
BENEFITS
Our benefits package includes:
- Competitive salary
- Equity
- Pension
- Health insurance
- Regular performance, compensation, and development reviews
- 20 weeks of Maternity & Paternity leave to spend time with new arrivals