🏢
Fastly
Senior SRE - Networks
Job Description
ABOUT THE ROLE
We're seeking a highly experienced Senior SRE - Networks to join our Technical Operations team at Fastly. As a key member of our team, you'll be responsible for building, operating, and maintaining our global network footprint, which spans 6 continents with over 100 points of presence and 578 Tbps of connected capacity. You'll play a critical role in ensuring the operational stability of our network and contributing to making the internet better.
WHAT YOU'LL DO
As a Senior SRE - Networks, you'll be responsible for:
- Building, operating, and maintaining the continually growing global network footprint of Fastly's Edge Cloud Platform.
- Responding to significant traffic incidents and leading network incidents, resolving edge cases and failure scenarios with your expertise in IP routing, particularly BGP.
- Writing code that is performant, maintainable, clear, and concise, and contributing to code reviews, improving the codebase and other team processes.
- Partnering in the development and iteration of tools and automation systems that improve how we operate and build the network.
- Innovating new methods for monitoring network performance, focusing on the end-user experience, and proactively addressing potential issues.
- Continual deep-dive of performance-based analytics and close involvement with partner teams to maintain a performant global network.
- Advocating for the operational stability of the network by identifying opportunities and partnering with engineering teams to shape their roadmaps and software solutions.
- Mentoring team members on the complexities of global routing, especially in an anycast-heavy environment.
WHAT YOU'LL NEED
To succeed in this role, you'll need:
- Experience in the protocols and practices that make up the fabric of the global internet, including TCP/IP, BGP, Anycast, and DNS.
- Experience in Automation and coding using languages like Python, or similar.
- Proficiency in Tier 1 Internet service providers, Internet exchanges, and cloud providers.
- Ability to analyze internet traffic patterns across multiple dimensions using flow-based tools.
- Experience working with alerting, monitoring, and visibility tools (such as Graphite/Grafana, Prometheus, or Splunk).
- Knowledge across cloud hosting solutions (i.e., GCP, AWS, and Azure).
- Knowledge of DevOps practices and CI/CD pipelines (ie. Git, Jenkins, Ansible).
- Understanding of Linux/Unix systems administration and network stack optimisation.
- Adept at knowledge sharing and creating comprehensive documentation to empower teams.
- Able to collaborate with cross-functional teams to shape the technical roadmap, prioritizing initiatives to optimize automation tooling and the network.
WHY REMOTE
As a remote employee, you'll have the flexibility to work from anywhere, with scheduled hours from 0800 - 1700. You'll be expected to be available during core business hours and may require occasional nights and weekends as necessary to support on-call coverage. Travel requirements will be determined by your role or manager.
BENEFITS
We care about your well-being and offer a comprehensive benefits package designed to meet your needs. Our benefits may vary depending on the country where you work and are subject to change.