🏢
Okta
Principal Site Reliability Engineer
Job Description
ABOUT THE ROLE
We're seeking a Principal Site Reliability Engineer to drive technical leadership and reliability engineering within Okta's Emerging Products Group (EPG). This role extends beyond operating production systems, requiring a deep understanding of technical strategy, platform architecture, and reliability standards. As a technical leader, you will establish reliability strategy, influence platform architecture, and lead transformational initiatives that improve scalability, resilience, security, and operational excellence for one of Okta's fastest-growing product areas.
WHAT YOU'LL DO
As a Principal Site Reliability Engineer, you will:
- Define and drive the reliability strategy for critical product and platform services
- Establish standards for availability, resilience, observability, incident management, and operational readiness
- Lead architecture reviews for critical services and platform initiatives
- Partner with engineering leaders to ensure reliability objectives align with business priorities and customer expectations
- Create frameworks, standards, and operational guardrails that enable engineering teams to operate safely at scale
- Guide service architecture toward simplicity, scalability, resilience, and operational excellence
- Drive major initiatives that improve platform maturity and long-term sustainability
- Own reliability architecture and operational excellence for the Spera / ISPM product area
- Collaborate closely with engineering leadership to establish reliability objectives and technical roadmaps
- Lead large-scale scalability, resiliency, and performance initiatives
- Partner with platform and product engineering teams to build self-service operational capabilities that improve developer productivity while strengthening reliability and security
- Influence technical direction through data-driven recommendations, engineering expertise, and collaborative leadership
- Support highly available, large-scale cloud environments as part of an on-call rotation
- Design, build, and operate large-scale cloud infrastructure and production services
- Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies
- Eliminate operational toil through automation, tooling, and platform engineering
- Improve deployment safety, operational workflows, and platform consistency through GitOps and Infrastructure-as-Code practices
- Collaborate on modernizing existing workloads and aligning them with evolving platform capabilities
- Lead complex engineering initiatives from conception through production rollout and long-term operational ownership
- Mentor Staff and Senior engineers across multiple teams and organizations
- Lead technical reviews, design reviews, and operational readiness assessments
- Build engineering consensus across teams with differing priorities and objectives
- Help develop the next generation of technical leaders within Okta
- Drive adoption of reliability engineering best practices across EPG
- Share patterns, tooling, and operational practices across Workflows, Inbox, PAM, and ISPM teams
- Influence technical direction through expertise, collaboration, and execution rather than organizational authority
WHAT YOU'LL NEED
- A deep understanding of technical strategy, platform architecture, and reliability standards
- Experience leading large-scale engineering initiatives that drive measurable business outcomes
- Strong organizational influence and ability to build engineering consensus across teams with differing priorities and objectives
- Excellent technical leadership and communication skills
- Ability to drive technical direction through data-driven recommendations, engineering expertise, and collaborative leadership
- Experience with Go, Python, Terraform, and related technologies
- Strong understanding of cloud infrastructure and production services
- Ability to eliminate operational toil through automation, tooling, and