PB✓
PBridge

Full-time jobsthe United Kingdom

Software Engineer, GPU Infrastructure- ChatGPT Engineering

openai · London, UK · Full-time

About this role

About the Team

ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance.

As our GPU fleet continues to grow, we're investing in the infrastructure that operates it. Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous. We work across production engineering, distributed systems, capacity management, and AI-powered operational tooling to help researchers and product teams move faster while maximizing the efficiency of every GPU.

This is a unique opportunity to work on infrastructure at the frontier of AI, where small improvements in fleet efficiency, reliability, and automation have an outsized impact on the development and deployment of AGI.

About the Role

We're looking for a Software Engineer with deep experience operating large-scale GPU or compute infrastructure.

You'll design and build the systems that manage GPU clusters at scale—from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention. You'll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization.

This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI.

In This Role, You Will

- Design, build, and operate software that manages large-scale GPU infrastructure supporting ChatGPT inference.

- Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead.

- Improve observability, reliability, and operational efficiency across thousands of GPUs.

- Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response.

- Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance.

- Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform.

- Help establish engineering best practices around operational excellence, automation, and infrastructure reliability.

You Might Thrive in This Role If You

- Have experience operating large-scale production infrastructure, preferably GPU clusters or other compute-intensive distributed systems.

- Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering.

- Have built software that automates operational workflows rather than relying on manual processes.

- Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure.

- Understand infrastructure observability, monitoring

Tired of applying one by one?

Our Career Success Team finds roles in the United Kingdom that fit you, tailors your CV to each, and submits the applications — tracked end to end. You just show up to interviews.

We apply, you interview →