💼Jobs 📝Blog 🧮Salary Calc 🌍Cost of Living 📋Tax Guide
👤Sign In Post a Job — from $99
Adyen

Platfrom Monitoring & Incident Engineer

EngineeringFull-TimeMid-Level
Location
Europe
Job Type
Full-Time
Experience
Mid-Level
Apply Now

Job Description

ABOUT THE ROLE We are seeking an experienced Platform Monitoring & Incident Engineer to join our team at Adyen. As a Platform Monitoring & Incident Engineer, you will play a critical role in ensuring the reliability and performance of our platform, which provides payments, data, and financial products to customers like Meta, Uber, H&M, and Microsoft. You will be part of a team that exhibits an unwavering attention to detail and a deep understanding of the platform-wide monitoring implications to all merchants. WHAT YOU'LL DO As a Platform Monitoring & Incident Engineer, you will be responsible for monitoring platform performance, coordinating and commanding incidents, communicating with customers, and working on monitoring frameworks. Your primary objectives will be to: - Monitor platform performance and detect any issues proactively to mitigate risks in partnership with Engineering teams. - Coordinate the mitigation, recovery, and resolution of high-impact incidents, ensuring a rapid and effective response across teams. - Communicate with merchants in real-time during incidents and present accurate and updated information to keep them informed. - Analyze incident trends to identify recurring issues and systemic weaknesses, and partner with engineering and product teams to advocate for long-term fixes over repeated short-term patches. - Work together with Operations, Product, and Engineering teams to integrate, grow, and continuously improve our monitoring strategy and increase our reliability. - Investigate alerts and provide feedback to engineering teams to build effective logging and alerts across the platform architecture. - Mitigate merchant impact risk by actioning on alerts in partnership with Engineering teams, and contribute to the monitoring playbook by documenting your learnings. - Improve operations by leading/project managing initiatives and developing automation for effective monitoring. - Focus on prioritizing, automating, and scaling every aspect of our detection capabilities. WHAT YOU'LL NEED To be successful in this role, you will need to have: - At least 5 years of experience with incident management, problem management, incident client communication, and platform monitoring operations. - Experience with problem management practices, including identifying trends across incidents, conducting root cause investigations, and driving preventative action. - Solid communication skills and the ability to develop strong working relationships throughout the organization, with the ability to translate technical situations clearly and concisely to a diverse audience. - Experience with monitoring and logging tools like Prometheus, Grafana, ELK Stack, etc. - Experience with observability platforms like Datadog, Dynatrace, Splunk. - Excellent analytical and problem-solving skills, with the ability to analyze complex systems and spot the root cause of issues. - Ability to thrive in a fast-paced, dynamic environment and participate in the on-call rotation. WHY REMOTE As a remote employee, you will have the flexibility to work from anywhere and be part of a global team that values collaboration and diversity. You will be working in a follow the sun model, with shifts from 9:00 AM - 6:00 PM PDT or 11:00 AM - 8:00 PM CDT, with a 6-day workweek at least twice a month. BENEFITS - Competitive salary - Comprehensive health plan - Parental leave - Education stipend - Life insurance - Equipment stipend - Flexible schedule