🏢
Adyen
Monitoring Engineer
Job Description
ABOUT THE ROLE
Adyen is a financial technology platform providing payments, data, and financial products to customers like Meta, Uber, H&M, and Microsoft. We're a team of motivated individuals tackling unique technical challenges at scale and delivering innovative and ethical solutions to help businesses achieve their ambitions faster. As a Platform Monitoring Engineer / Incident Manager, you will be part of a team under the Platform Excellence pillar, exhibiting an unwavering attention to detail and a deep understanding of the platform-wide monitoring implications to all merchants.
WHAT YOU'LL DO
In this role, you will be responsible for on-call monitoring of platform performance, coordinating and commanding incidents, communicating with customers, working on monitoring frameworks, and providing feedback to product engineering teams to improve the reliability of the platform. You will initiate and lead initiatives across our platform offerings, prioritizing merchant impact to proactively detect any issues, inform merchants quickly, and increase the reliability of our platform.
Key responsibilities include:
- Participating in 24/7 on-call monitoring, observing platform and merchant performance, and detecting any issues proactively to mitigate risks in partnership with Engineering teams.
- Coordinating the mitigation, recovery, and resolution of high-impact incidents, ensuring a rapid and effective response across teams.
- Communicating with merchants in real-time during incidents, presenting accurate and updated information to keep them informed.
- Analyzing incident trends to identify recurring issues and systemic weaknesses, partnering with engineering and product teams to advocate for long-term fixes over repeated short-term patches.
- Integrating, growing, and continuously improving our monitoring strategy and increasing our reliability by working together with Operations, Product, and Engineering teams.
- Investigating alerts and providing feedback to engineering teams to build effective logging and alerts across the platform architecture.
- Mitigating merchant impact risk by actioning on alerts in partnership with Engineering teams, contributing to the monitoring playbook by documenting learnings.
- Improving operations by leading/project managing initiatives and developing automation for effective monitoring.
- Focusing on ruthlessly prioritizing, automating, and scaling every aspect of our detection capabilities.
WHAT YOU'LL NEED
To be successful in this role, you will need at least 5 years of experience with incident management, problem management, incident client communication, and platform monitoring operations. You should have experience with problem management practices, including identifying trends across incidents, conducting root cause investigations, and driving preventative action. You will also need solid communication skills, the ability to develop strong working relationships throughout the organization, and experience with monitoring and logging tools like Prometheus, Grafana, ELK Stack, etc. Additionally, you should have experience with observability platforms like Datadog, Dynatrace, Splunk, and excellent analytical and problem-solving skills.
WHY REMOTE
This role is a great opportunity to work in a fast-paced, dynamic environment, where collaboration is crucial and a global approach is key for successful implementation of processes and projects. You will be part of a diverse team with a unique approach, where different perspectives are essential in helping us maintain our momentum.
BENEFITS
Adyen offers a competitive salary and benefits package, including a work schedule of 9.00AM - 6.00PM with a 6-day work