PB✓
PBridge
Full-timeOtherWorldwide

Technical Program Manager, Core Network & WAN Infrastructure

at Open AI

OpenAI is seeking a Technical Program Manager to lead the end-to-end delivery of WAN infrastructure and large-scale GPU cluster deployments. This hands-on role involves managing complex network builds and cross-functional readiness in a hybrid work environment based in San Francisco.

Job Description

About the Team

The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads.

We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth.

About the Role

You will join an engineer-first TPM team and own end-to-end delivery of OpenAI’s WAN Infrastructure, partnering with engineers to bring clusters online across external providers and partners.

This is a hands-on infrastructure execution role. You’ll run a broad, parallel portfolio spanning hardware, cabling, optics, cloud cross-connects, and port maps—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale.

The right person combines enough technical depth to reason through physical and logical network readiness with the program discipline and ownership to keep complex builds moving and improve how we scale.

This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance.

In this role, you will

  • Lead end-to-end delivery of OpenAI’s WAN build-out and large-scale GPU clusters across an external partner ecosystem
  • Drive multi-threaded network infrastructure bring-up programs across physical and logical readiness—owning plans, dependencies, and critical paths
  • Partner with engineering to turn up large-scale network capacity, reducing bring-up time for our largest GPU supercomputers by working across optics, circuits
  • Identify recurring bottlenecks in WAN and network infrastructure build-out, then drive fixes that make future builds faster, more predictable, and less dependent on tribal knowledge
  • Coordinate cross-functional readiness (security, finance, operations, product/research stakeholders) to ship production-ready compute
  • Manage integration and handoffs across teams and partners—ensuring consistent execution, clear communication, and fast issue resolution
  • Identify bottlenecks and systemic gaps, then drive durable fixes across tooling, process, and partner interfaces
  • Provide crisp executive visibility on progress, tradeoffs, and risks across a large portfolio of concurrent programs

You might thrive in this role if you have

  • Possess a degree in a hard science, or have a demonstrated track record of engineering expertise.
  • Have 5+ years of experience in program management for major projects including capital projects or hyperscaler infrastructure deployment.
  • Are relentlessly resourceful and thrive in ambiguous, fast-paced environments
  • Demonstrate the ability to serve as the go-to person solely responsible for driving and delivering complex projects.
  • Strong technical intuition across physical networking, WAN/backbone infrastructure, colocation environments, cloud interconnects, cross-connects, optics, cabling, routing, and operational readiness.
  • Have experience interfacing with and leading external vendors including engineering firms, equipment suppliers, and/or construction firms.
  • Have expertise in designing and implementing simple, scalable processes that solve complex problems.
  • Have experience managing complicated dependencies such as logistics and/or supply chains.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its

Responsibilities & Requirements

Responsibilities

  • Lead end-to-end delivery of OpenAI’s WAN build-out and large-scale GPU clusters across an external partner ecosystem
  • Drive multi-threaded network infrastructure bring-up programs across physical and logical readiness
  • Partner with engineering to turn up large-scale network capacity
  • Identify recurring bottlenecks in WAN and network infrastructure build-out and drive fixes
  • Coordinate cross-functional readiness across security, finance, operations, and product/research stakeholders
  • Manage integration and handoffs across teams and partners
  • Identify systemic gaps and drive durable fixes across tooling, process, and partner interfaces
  • Provide executive visibility on progress, tradeoffs, and risks

Requirements

  • Degree in a hard science or demonstrated track record of engineering expertise
  • 5+ years of experience in program management for major projects including capital projects or hyperscaler infrastructure deployment
  • Technical intuition across physical networking, WAN/backbone infrastructure, colocation environments, cloud interconnects, cross-connects, optics, cabling, routing, and operational readiness
  • Experience interfacing with and leading external vendors including engineering firms, equipment suppliers, and/or construction firms
  • Expertise in designing and implementing simple, scalable processes
  • Experience managing complicated dependencies such as logistics and/or supply chains
  • Ability to serve as the go-to person responsible for driving and delivering complex projects
  • Relentlessly resourceful in ambiguous, fast-paced environments

Benefits & Perks

  • Relocation assistance
  • Hybrid work model (3 days in office per week)

Skills

Program ManagementWAN InfrastructureGPU ClustersNetwork EngineeringInfrastructure DeploymentVendor ManagementSupply ChainCross-functional LeadershipRisk ManagementTechnical Operations

Tags

Technical Program ManagementTechnical Program Management