About this role
ABOUT THE TEAM
The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI’s largest-scale AI systems. We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads.
As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical. Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI’s rapidly expanding compute environment.
About the Role
We are seeking a Hardware Technical Program Manager, Infrastructure Partner Operations to lead operational delivery across OpenAI’s third-party infrastructure partners, including major cloud service providers and strategic compute vendors.
In this role, you will serve as the primary operational program manager for external infrastructure partners, driving accountability for service delivery, operational readiness, incident management, performance reporting, and continuous operational improvement. You will work closely with partner engineering and operations teams while coordinating internally across Hardware Engineering, Infrastructure Operations, Capacity Planning, Networking, Supply Chain, Deployment, Reliability Engineering, and executive leadership.
Success in this role requires someone who understands how hyperscale infrastructure organizations operate, can establish strong operational governance with external partners, and is comfortable driving complex technical programs without direct ownership of the underlying infrastructure.
Key Responsibilities
- Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations.
- Develop operational governance frameworks with strategic partners, including business reviews, operational scorecards, escalation processes, executive reporting, and performance improvement plans.
- Define, track, and continuously improve key operational metrics related to infrastructure availability, deployment execution, incident response, operational health, service quality, and partner performance.
- Build dashboards and reporting mechanisms that provide clear visibility into partner operational performance, risks, trends, and areas requiring executive attention.
- Drive cross-functional coordination between OpenAI teams and external infrastructure providers to resolve operational issues, remove execution blockers, and improve delivery outcomes.
- Lead operational escalations involving infrastructure availability, deployment execution, hardware operations, capacity delivery, or service performance, ensuring timely resolution and clear executive communication.
- Establish rep