💼Jobs 📝Blog 🧮Salary Calc 🌍Cost of Living 📋Tax Guide
👤Sign In Post a Job — from $99
Okta

Staff Site Reliability Engineer - (Infra)

DevOpsFull-TimeLead
Location
Worldwide
Job Type
Full-Time
Experience
Lead
Apply Now

Job Description

ABOUT THE ROLE We are seeking a seasoned Senior Cloud Engineer to join Okta's Infrastructure Engineering team. As a key member of our team, you will design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. This is a critical role that requires a relentless drive to solve complex challenges with real-world stakes. WHAT YOU'LL DO As a Senior Cloud Engineer, you will be responsible for leading major reliability and modernization initiatives, including container platform migrations and microservice enablement across multi-cloud environments. You will serve as a technical authority in Kubernetes, cloud infrastructure, and modern CI/CD practices. Your responsibilities will include: - Designing, building, and operating highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP - Leading major reliability and modernization initiatives, including container platform migrations and microservice enablement across multi-cloud environments - Partnering with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability - Implementing and managing infrastructure as code to automate provisioning, scaling, and configuration management across multiple cloud providers - Driving improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP - Championing SRE best practices, including defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response - Leading complex technical projects from conception to completion, managing timelines, and technical dependencies across teams - Mentoring engineers across teams, fostering a culture of reliability, automation, and continuous learning - Collaborating with security and compliance partners to ensure infrastructure adheres to best practices and standards - Participating in the on-call rotation, using incidents as learning opportunities to enhance systems and processes WHAT YOU'LL NEED To be successful in this role, you will need to bring the following skills and experience: - Strong hands-on experience architecting and operating cloud-native distributed systems (AWS and GCP) - Deep expertise with Kubernetes (EKS and GKE) — design, provisioning, scaling, and advanced troubleshooting in production - Proven experience leading ECS to EKS/GKE migrations and driving microservice enablement initiatives at scale - Proficiency with Infrastructure as Code tools such as Terraform (multi-provider), Ansible, or CloudFormation - Solid coding and scripting ability in Python, Go, or Shell, with a focus on automation, tooling, and operational excellence - Advanced understanding of CI/CD pipelines, Linux systems, and networking fundamentals - Experience managing databases and caching systems in cloud environments - Hands-on experience with observability tools for performance and reliability insights - Working knowledge of container security, secrets management, and compliance in production environments - Strong communication and problem-solving skills, with demonstrated success leading cross-team projects and mentoring peers WHY REMOTE Okta is a global company with a remote-friendly culture. As a Senior Cloud Engineer, you will have the flexibility to work from anywhere and be part of a community that values connection and collaboration. BENEFITS Okta offers a comprehensive benefits package, including: - Supporting Your Well-