💼Jobs 📝Blog 🧮Salary Calc 🌍Cost of Living 📋Tax Guide
👤Sign In Post a Job — from $99
Adyen

Monitoring Engineer

EngineeringFull-TimeMid-Level
Location
Worldwide
Job Type
Full-Time
Experience
Mid-Level
Apply Now

Job Description

ABOUT THE ROLE Adyen is a financial technology platform providing payments, data, and financial products to customers like Meta, Uber, H&M, and Microsoft. We're a team of motivated individuals tackling unique technical challenges at scale and delivering innovative and ethical solutions to help businesses achieve their ambitions faster. As a Platform Monitoring Engineer / Incident Manager, you will be part of a team under the Platform Excellence pillar, exhibiting an unwavering attention to detail and a deep understanding of the platform-wide monitoring implications to all merchants. WHAT YOU'LL DO In this role, you will be responsible for on-call monitoring of platform performance, coordinating and commanding incidents, communicating with customers, working on monitoring frameworks, and providing feedback to product engineering teams to improve the reliability of the platform. You will initiate and lead initiatives across our platform offerings, prioritizing merchant impact to proactively detect any issues, inform merchants quickly, and increase the reliability of our platform. Key responsibilities include: - Participating in 24/7 on-call monitoring, observing platform and merchant performance, and detecting any issues proactively to mitigate risks in partnership with Engineering teams. - Coordinating the mitigation, recovery, and resolution of high-impact incidents, ensuring a rapid and effective response across teams. - Communicating with merchants in real-time during incidents, presenting accurate and updated information to keep them informed. - Analyzing incident trends to identify recurring issues and systemic weaknesses, partnering with engineering and product teams to advocate for long-term fixes over repeated short-term patches. - Integrating, growing, and continuously improving our monitoring strategy and increasing our reliability by working together with Operations, Product, and Engineering teams. - Investigating alerts and providing feedback to engineering teams to build effective logging and alerts across the platform architecture. - Mitigating merchant impact risk by actioning on alerts in partnership with Engineering teams, contributing to the monitoring playbook by documenting learnings. - Improving operations by leading/project managing initiatives and developing automation for effective monitoring. - Focusing on ruthlessly prioritizing, automating, and scaling every aspect of our detection capabilities. WHAT YOU'LL NEED To be successful in this role, you will need at least 5 years of experience with incident management, problem management, incident client communication, and platform monitoring operations. You should have experience with problem management practices, including identifying trends across incidents, conducting root cause investigations, and driving preventative action. You will also need solid communication skills, the ability to develop strong working relationships throughout the organization, and experience with monitoring and logging tools like Prometheus, Grafana, ELK Stack, etc. Additionally, you should have experience with observability platforms like Datadog, Dynatrace, Splunk, and excellent analytical and problem-solving skills. WHY REMOTE This role is a great opportunity to work in a fast-paced, dynamic environment, where collaboration is crucial and a global approach is key for successful implementation of processes and projects. You will be part of a diverse team with a unique approach, where different perspectives are essential in helping us maintain our momentum. BENEFITS Adyen offers a competitive salary and benefits package, including a work schedule of 9.00AM - 6.00PM with a 6-day work