OpenAI is dedicated to ensuring that artificial general intelligence (AGI) benefits all of humanity. Our mission requires building not only world-class AI models, but also the infrastructure that enables those models to be deployed reliably, efficiently, and at global scale. As demand for AI continues to grow, we are expanding the ways OpenAI can bring high-performance inference capacity online across a diverse hardware ecosystem.
The GPT Infrastructure team builds software that turns advanced inference and optimization research into production products.
What you'd do
- Design, build, and operate durable APIs and control-plane services for multi-hour or multi-day optimization campaigns, including scheduling, retries, budgets, checkpoints, artifact lineage, and observability.
- Build secure partner-side runner and grader software that can compile, execute, verify, and benchmark candidate artifacts on third-party accelerator hardware.
- Integrate hardware profiles, ISA and toolchain context, compilers, runtimes, and inference-serving engines into a repeatable optimization workflow.
- Turn research prototypes into reliable product surfaces with clear contracts, debuggable failure modes, reproducible outputs, and excellent developer ergonomics.
- Develop correctness and performance evaluation systems spanning latency, throughput, memory use, utilization, and cost efficiency.
- Build artifact, provenance, and qualification workflows that make optimized kernels, binaries, configurations, and reports safe to review and deploy.
- Collaborate with Research, Inference Engineering, Infrastructure, Security, Product, and Strategic Partnerships to deliver production-ready solutions.
- Drive technical architecture and execution across ambiguous, cross-functional initiatives that connect OpenAI systems with partner environments.
What they want
- 8+ years of professional software engineering experience building large-scale distributed systems, infrastructure platforms, or cloud services, or equivalent depth of experience.
- Strong programming skills in one or more of C++, Python, Go, or Rust.
- Experience designing and operating highly available backend systems, APIs, job orchestration systems, or durable workflows for production workloads.
- Strong understanding of distributed systems, Linux, networking, storage, containers, and modern cloud architectures.
- Experience debugging complex systems and using measurement, profiling, and benchmarks to guide engineering decisions.
- Proven ability to lead complex technical initiatives as a senior individual contributor and work effectively across organizational boundaries.