软件工程师,计算 - 存储
Software Engineer, Compute - Storage
关于团队
存储基础设施团队构建并运营支撑OpenAI最繁重工作负载的存储基础架构。我们直接与研究团队合作,为快速发展的实验设计存储系统,同时也在大规模生产环境中提供支持。我们全程负责平台:后端系统、面向用户的服务和API,以及管理数据如何随时间放置、移动和保留的控制平面。
我们的技术栈涵盖跨不同工作负载特性的云和自建对象存储,从GPU连接的系统到专用存储硬件。我们还构建了联邦层,将这些后端统一在简单的接口之下,并将每个工作负载路由到合适的存储解决方案。
关于职位
你将帮助构建支撑OpenAI研究和生产系统的存储平台。这是一个需要动手实践的基础设施职位,适合希望在大规模环境下参与深度技术系统并负责其生产的工程师。
你将涉及对象存储、跨区域数据传输、生命周期管理,以及提供多个后端统一接口的联邦层。我们的大部分技术栈运行在Kubernetes上,我们主要使用Rust构建服务。
在此职位中,你将:
- 构建和运营支撑OpenAI研究基础设施的存储服务
- 在云和自建环境中开发对象存储系统
- 构建跨区域数据传输、复制和恢复系统
- 设计保持数据持久性、可用性和成本效益的生命周期管理功能
- 持续演进联邦层,将多个后端系统统一在简单接口之下
- 提升平台的性能、可靠性和运维卓越性
- 与研究人员和基础设施团队紧密合作,支持快速变化的工作负载
你可能适合这个职位,如果你:
- 有在生产环境中构建或运维分布式系统的经验
- 有存储基础设施、对象存储、分布式文件系统或其他数据密集型后端系统的经验
- 喜欢全程负责基础设施,包括调试和长期可靠性改进
- 编写高质量的生产代码,最好是Rust或其他系统导向语言
- 熟悉Kubernetes-based系统
- 有使用Terraform、Grafana或其他基础设施和可观测性工具的经验
关于OpenAI
查看英文原文
ABOUT THE TEAM
The Storage Infrastructure team builds and operates the storage foundation behind OpenAI’s most demanding workloads. We work directly with research to design storage systems for rapidly evolving experiments, while also powering production at scale. We own the platform end to end: backend systems, user-facing services and APIs, and the control planes that manage how data is placed, moved, and retained over time.
Our stack spans cloud and in-house object stores across very different workload profiles, from GPU-attached systems to dedicated storage hardware. We also build the federation layer that unifies these backends behind a simple interface and routes each workload to the right storage solution.
ABOUT THE ROLE
You will help build the storage platform that powers OpenAI’s research and production systems. This is a hands-on infrastructure role for engineers who want to work on deeply technical systems at scale and own them in production.
You’ll work across object storage, cross-region data movement, lifecycle management, and the federation layer that provides a unified interface across multiple backends. Much of our stack runs on Kubernetes, and we primarily build services in Rust.
In this role, you will:
- Build and operate storage services that underpin OpenAI’s research infrastructure
- Develop object storage systems across cloud and in-house environments
- Build systems for cross-region data movement, replication, and recovery
- Design lifecycle management capabilities that keep data durable, available, and cost-effective
- Evolve the federation layer that unifies multiple backend systems behind a simple interface
- Improve performance, reliability, and operational excellence across the platform
- Collaborate closely with researchers and infrastructure teams to support rapidly evolving workloads
You might thrive in this role if you:
- Have experience building or operating distributed systems in production
- Have worked on storage infrastructure, object stores, distributed filesystems, or other data-intensive backend systems
- Enjoy owning infrastructure end to end, including debugging and long-term reliability improvements
- Write strong production code, ideally in Rust or another systems-oriented language
- Are comfortable working with Kubernetes-based systems
- Have experience with tools such as Terraform, Grafana, or similar infrastructure and observability tooling
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf.
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form https://form.asana.com/?d=57018692298241&k=5MqR40fZd7jlxVUh5J-UeA. No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link https://form.asana.com/?k=bQ7w9h3iexRlicUdWRiwvg&d=57018692298241.
OpenAI Global Applicant Privacy Policy https://cdn.openai.com/policies/global-employee-and-contractor-privacy-policy.pdf
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.