🏢
Adyen
Platfrom Monitoring & Incident Engineer
Job Description
ABOUT THE ROLE
We are seeking an experienced Platform Monitoring & Incident Engineer to join our team at Adyen. As a Platform Monitoring & Incident Engineer, you will play a critical role in ensuring the reliability and performance of our platform, which provides payments, data, and financial products to customers like Meta, Uber, H&M, and Microsoft. You will be part of a team that exhibits an unwavering attention to detail and a deep understanding of the platform-wide monitoring implications to all merchants.
WHAT YOU'LL DO
As a Platform Monitoring & Incident Engineer, you will be responsible for monitoring platform performance, coordinating and commanding incidents, communicating with customers, and working on monitoring frameworks. Your primary objectives will be to:
- Monitor platform performance and detect any issues proactively to mitigate risks in partnership with Engineering teams.
- Coordinate the mitigation, recovery, and resolution of high-impact incidents, ensuring a rapid and effective response across teams.
- Communicate with merchants in real-time during incidents and present accurate and updated information to keep them informed.
- Analyze incident trends to identify recurring issues and systemic weaknesses, and partner with engineering and product teams to advocate for long-term fixes over repeated short-term patches.
- Work together with Operations, Product, and Engineering teams to integrate, grow, and continuously improve our monitoring strategy and increase our reliability.
- Investigate alerts and provide feedback to engineering teams to build effective logging and alerts across the platform architecture.
- Mitigate merchant impact risk by actioning on alerts in partnership with Engineering teams, and contribute to the monitoring playbook by documenting your learnings.
- Improve operations by leading/project managing initiatives and developing automation for effective monitoring.
- Focus on prioritizing, automating, and scaling every aspect of our detection capabilities.
WHAT YOU'LL NEED
To be successful in this role, you will need to have:
- At least 5 years of experience with incident management, problem management, incident client communication, and platform monitoring operations.
- Experience with problem management practices, including identifying trends across incidents, conducting root cause investigations, and driving preventative action.
- Solid communication skills and the ability to develop strong working relationships throughout the organization, with the ability to translate technical situations clearly and concisely to a diverse audience.
- Experience with monitoring and logging tools like Prometheus, Grafana, ELK Stack, etc.
- Experience with observability platforms like Datadog, Dynatrace, Splunk.
- Excellent analytical and problem-solving skills, with the ability to analyze complex systems and spot the root cause of issues.
- Ability to thrive in a fast-paced, dynamic environment and participate in the on-call rotation.
WHY REMOTE
As a remote employee, you will have the flexibility to work from anywhere and be part of a global team that values collaboration and diversity. You will be working in a follow the sun model, with shifts from 9:00 AM - 6:00 PM PDT or 11:00 AM - 8:00 PM CDT, with a 6-day workweek at least twice a month.
BENEFITS
- Competitive salary
- Comprehensive health plan
- Parental leave
- Education stipend
- Life insurance
- Equipment stipend
- Flexible schedule