🏢
Brex
Senior Software Engineer, Release Infra
Job Description
ABOUT THE ROLE
As a Senior Software Engineer, Release Infra at Brex, you will play a critical role in designing, building, and operating the core systems that power Brex's release, observability, and incident management processes. You will partner closely with product, platform, and operations teams to ensure releases are safe, fast, and reliable, and that our infrastructure scales securely as Brex grows.
WHAT YOU'LL DO
As a Senior Software Engineer, Release Infra, your responsibilities will include:
Designing, building, and maintaining the release infrastructure that powers Brex's deployment pipelines and incident workflows
Driving technical strategy and architecture for release and observability systems, making them more scalable, reliable, and secure
Collaborating with product, engineering, and operations partners to ensure Brex's releases are safe, predictable, and low-friction
Identifying and delivering improvements to the end-to-end release process (from code merge to production) to reduce risk and cycle time
Building and evolving tooling for observability and incident response, enabling fast detection, triage, and resolution
Proactively identifying and mitigating risks in our release and infrastructure stack, including performance, reliability, and security concerns
Defining, instrumenting, and monitoring key metrics for release engineering (e.g., deployment frequency, change failure rate, MTTR) and using them to guide improvements
Partnering with other infrastructure and product teams to debug complex production issues and drive long-term fixes
Contributing to and championing best practices in release engineering, reliability, and operational excellence across the organization
Mentoring other engineers on the team, providing technical guidance and code reviews to elevate the overall quality of our infrastructure
Staying up-to-date on emerging tools and practices in release engineering, observability, and SRE, and bringing relevant ideas into Brex's stack
WHAT YOU'LL NEED
To be successful in this role, you will need:
7+ years of professional experience designing, building, and operating backend or infrastructure systems in production
Strong proficiency in backend programming languages (e.g., Go, Java, Kotlin, or Python) with a focus on reliability and performance
Hands-on experience with CI/CD and release pipelines (e.g., GitHub Actions, CircleCI, Buildkite, Argo, Spinnaker, Jenkins) including build, test, and deployment automation
Experience architecting and operating scalable, high-availability distributed systems on cloud platforms (e.g., AWS, GCP, Azure)
Deep familiarity with containerization and orchestration (e.g., Docker, Kubernetes) and infrastructure-as-code (e.g., Terraform, CloudFormation)
Experience designing and maintaining observability tooling (metrics, logs, tracing) and integrating it into incident response workflows
Strong understanding of reliability and SRE practices, including SLIs/SLOs, error budgets, and incident management best practices
Experience designing and optimizing data storage systems (SQL and/or NoSQL) for operational and observability use cases
Proven track record of improving release processes (e.g., reducing deployment risk, increasing deployment frequency, automating rollbacks)
Comfort working cross-functionally with product and other engineering teams to debug complex production issues and ship changes safely