About this role
About the Team
OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads.
About the Role
We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network.
The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil.
Key Responsibilities
- Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers.
- Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR).
- Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks.
- Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact.
- Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance.
- Support new AI cluster deployments, data center expansions, and infrastructure migrations in partnership with deployment and engineering teams.
- Partner with cloud service providers (CSPs), colocation providers, Smart Hands teams, and hardware vendors to maintain production infrastructure.
- Perform root-cause analysis (RCA) for production incidents and drive permanent corrective actions that eliminate recurring issues.
- Build and maintain monitoring, telemetry, dashboards, and alerting to improve network observability and proactive issue detection.
- Develop and improve operational runbooks, playbooks, troubleshooting documentation, and standard operating procedures.
- Automate repetitive operational tasks using Python and infrastructure automation frameworks to reduce toil and improve efficiency.
- Continuously identify opportunities to improve service reliability, scalability, operational maturity, and engineering efficiency.
Qualifications
- Bachelo