💼Jobs 📝Blog 🧮Salary Calc 🌍Cost of Living 📋Tax Guide
👤Sign In Post a Job — from $99
MongoDB

Site Reliability Engineering, Fabric (Mid, Senior, or Staff)

NewDevOpsFull-TimeLead
Location
Worldwide
Job Type
Full-Time
Experience
Lead
Apply Now

Job Description

ABOUT THE ROLE MongoDB’s Fabric team is the backbone of our global, multi‑cloud network, ensuring secure, reliable communication between services and the public internet. As a Site Reliability Engineer on this team, you will architect, build, and maintain the infrastructure that keeps MongoDB Atlas and other cloud‑native products connected and protected worldwide. Your work will directly influence the performance, security, and resilience of the platform that powers millions of customers, including Fortune 100 enterprises and AI‑native startups. You will collaborate with platform, security, and product engineering groups to design network topologies, service meshes, and edge load‑balancing solutions that span AWS, Azure, and GCP. Your expertise will drive automation of network provisioning, monitoring, and incident response, reducing manual toil and accelerating feature delivery. You will also mentor and guide service‑owning teams on best practices for service‑to‑service connectivity, ensuring that every new microservice adheres to our high standards for availability and compliance. WHAT YOU'LL DO - Design and implement multi‑cloud networking primitives, including VPCs, subnets, VPNs, peering, and private link services, to support MongoDB’s global architecture. - Build and maintain service mesh and edge load‑balancing layers that guarantee low‑latency, secure traffic flow across regions and cloud providers. - Automate network configuration, monitoring, and alerting using IaC, Terraform, and observability tools such as Prometheus and Grafana. - Participate in a 24/7 on‑call rotation, diagnosing and resolving connectivity incidents with minimal downtime. - Partner with security teams to enforce TLS/mTLS, BGP, and SDN policies, ensuring data remains protected in transit. - Continuously evaluate emerging networking technologies and propose enhancements that improve reliability, scalability, and cost efficiency. WHAT YOU'LL NEED - 10+ years of experience designing, deploying, and operating distributed systems with deep knowledge of TCP/IP, IPv6, DNS, TLS/mTLS, BGP, and overlay networking. - Proven expertise in at least one major cloud provider’s networking stack (AWS VPC, Azure VNets, or GCP VPC) and familiarity with multi‑cloud connectivity patterns such as peering, private link, and CDN integration. - Strong background in service mesh (e.g., Istio, Linkerd) and load‑balancing concepts, with hands‑on experience deploying them at scale. - A passion for automation: proficiency with Terraform, Ansible, or equivalent IaC tools, and a track record of reducing manual operations through scripting or tooling. - Excellent problem‑solving skills, a customer‑focused mindset, and the ability to communicate complex technical concepts to both engineering and non‑engineering stakeholders. - Experience with observability, alerting, and incident response in a high‑availability environment. WHY REMOTE MongoDB embraces a fully distributed, asynchronous culture that empowers teams to work from any North American location. Remote employees enjoy the flexibility to balance work with personal commitments while collaborating with colleagues across time zones through robust video, chat, and documentation platforms. Our hybrid model also supports those who prefer an office environment, offering flexible in‑office days and a collaborative workspace in Toronto or Vancouver. BENEFITS MongoDB offers a comprehensive benefits package that includes health, dental, and vision coverage; a generous paid time off policy; 20 weeks of fully paid gender‑neutral parental leave; fertility and adoption assistance; and an RRSP with employer match. Eligible employees receive a remote work stipend, a learning and