PB✓
PBridge

Full-time jobsthe United States

Software Engineer, Site Reliability (SRE)

sierra · San Francisco, CA · Full-time

About this role

ABOUT US

At Sierra, we’re creating a platform to help businesses build better, more human customer experiences with AI. We are primarily an in-person company based in San Francisco, with growing offices in Atlanta, New York, London, Paris, Madrid, Munich, Singapore, Tokyo, and Sydney.

We are guided by a set of values that are at the core of our actions and define our culture: Trust, Customer Obsession, Craftsmanship, Intensity, and Family. These values are the foundation of our work, and we are committed to upholding them in everything we do.

Our co-founders are Bret Taylor https://www.linkedin.com/in/brettaylor/ and Clay Bavor https://www.linkedin.com/in/claybavor/. Bret currently serves as Board Chair of OpenAI. Previously, he was co-CEO of Salesforce (which had acquired the company he founded, Quip) and CTO of Facebook. Bret was also one of Google's earliest product managers and co-creator of Google Maps. Before founding Sierra, Clay spent 18 years at Google, where he most recently led Google Labs. Earlier, he started and led Google’s AR/VR effort, Project Starline, and Google Lens. Before that, Clay led the product and design teams for Google Workspace. 

WHAT YOU'LL DO

As a Software Engineer on our Site Reliability team at Sierra, you will be responsible for defining and building the foundation of reliability, observability, and scalability across Sierra’s AI-driven infrastructure. You’ll partner closely with our core engineering and product teams to ensure our systems are highly available, efficient, and built for growth.

- Own Sierra’s observability stack—monitoring, alerting, logging, and tracing—to give engineers clear visibility into system health and performance.

- Partner with product and platform engineers to design systems that are reliable and scalable from day one—not as an afterthought.

- Design and implement scalable, reliable, and secure cloud infrastructure (AWS) using Terraform and modern DevOps tooling.

- Improve the reliability and scalability of our LLM deployments, ensuring robust, performant, and cost-effective operation.

- Lead improvements to deployment pipelines, CI/CD tooling, and incident management processes to reduce downtime and response time.

- Define the foundation of SRE practices at Sierra, influencing culture, tooling, and best practices across the engineering org.

WHAT YOU'LL BRING

- 5+ years of hands-on experience in Site Reliability or Infrastructure engineering roles for complex SaaS or cloud-based systems.

- Experience designing for availability, scalability, and reliability at both infrastructure and application layers.

- Deep experience with Terraform, AWS services, container orchestration, and cloud networking (including IAM and VPC architecture).

- Strong background in observability systems (e.g., Prometheus, Grafana, Datadog, or similar).

- Experience working with enterprise customers and familiarity with their compliance and networking needs along with integration patt

Tired of applying one by one?

Our Career Success Team finds roles in the United States that fit you, tailors your CV to each, and submits the applications — tracked end to end. You just show up to interviews.

We apply, you interview →