数据科学家,推理能力优化
Data Scientist, Inference Capacity Optimization
职位描述
OpenAI的工业计算团队负责确保我们的计算基础设施能够高效扩展,以支持数百万用户和日益复杂的AI模型。
我们正在寻找一名数据科学家,与容量系统工程、基础设施、产品和研究团队紧密合作,优化我们全球GPU集群的推理容量。该职位结合统计建模、大规模数据分析、预测和系统思维,推动关于基础设施投资、性能-效率权衡和客户体验的关键决策。
你将把复杂的运营数据转化为可操作的见解,直接影响OpenAI如何分配和扩展全球最大的AI计算环境。
主要职责
- 构建统计和机器学习模型,用于分析和提升GPU利用率、延迟、吞吐量和整体集群效率。
- 为不同产品、地区和模型系列的推理需求构建预测模型。
- 分析生产工作负载,识别延迟瓶颈和容量限制,突出优化机会。
- 与容量系统工程团队合作,为基础设施规划和长期GPU投资策略提供依据。
- 设计实验和模拟,评估调度策略、服务策略和基础设施权衡。
- 构建仪表盘和运营指标,使管理层能够做出数据驱动的容量决策。
- 与产品、研究、财务和基础设施团队协作,确保计算规划与业务增长和模型路线图保持一致。
- 向工程团队和高管层清晰传达技术发现。
任职要求
- 统计学、计算机科学、运筹学、应用数学、经济学或相关定量学科的硕士或博士学历(或同等行业经验)。
- 在基础设施数据科学领域有5年以上工作经验。
- 精通Python和SQL。
- 有构建预测、优化或预测模型的经验。
- 对实验设计、统计推断和因果分析有深刻理解。
- 有向高管利益相关者传达分析见解的经验。
优先考虑的技能
- 容量规划
- 分布式系统
- AI基础设施
- 数据中心设计与建设
- 排队论
- 时间序列预测
- 运营优化
查看英文原文
About the Role
OpenAI’s Industrial Compute organization is responsible for ensuring our compute infrastructure scales efficiently to support millions of users and increasingly sophisticated AI models.
We’re looking for a Data Scientist to partner closely with Capacity Systems Engineering, Infrastructure, Product, and Research to optimize inference capacity across our global GPU fleet. This role combines statistical modeling, large-scale data analysis, forecasting, and systems thinking to drive critical decisions around infrastructure investments, performance-efficiency trade-offs, and customer experience.
You’ll transform complex operational data into actionable insights that directly influence how OpenAI allocates and scales one of the world’s largest AI compute environments.
Key Responsibilities
- Build statistical and machine learning models to profile and improve GPU utilization, latency, throughput, and overall fleet efficiency.
- Develop forecasting models for inference demand across products, regions, and model families.
- Analyze production workloads to identify latency bottlenecks and capacity constraints, highlighting optimization opportunities.
- Partner with Capacity Systems Engineering to inform infrastructure planning and long-term GPU investment strategies.
- Design experiments and simulations to evaluate scheduling policies, serving strategies, and infrastructure tradeoffs.
- Build dashboards and operational metrics that enable leadership to make data-driven capacity decisions.
- Collaborate with Product, Research, Finance, and Infrastructure teams to align compute planning with business growth and model roadmaps.
- Communicate technical findings clearly to both engineering teams and executive leadership.
Qualifications
- MS or PhD in Statistics, Computer Science, Operations Research, Applied Mathematics, Economics, or related quantitative discipline (or equivalent industry experience).
- 5+ years of experience working in the infrastructure data science space.
- Strong expertise in Python and SQL.
- Experience building forecasting, optimization, or predictive models.
- Strong understanding of experimentation, statistical inference, and causal analysis.
- Experience communicating analytical insights to executive stakeholders.
Preferred Skills
- Capacity planning
- Distributed systems
- AI infrastructure
- Datacenter design and buildout
- Queueing theory
- Time-series forecasting
- Operations research
- Supply-demand modeling
- Reinforcement learning for resource allocation
- Cost optimization
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf.
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form https://form.asana.com/?d=57018692298241&k=5MqR40fZd7jlxVUh5J-UeA. No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link https://form.asana.com/?k=bQ7w9h3iexRlicUdWRiwvg&d=57018692298241.
OpenAI Global Applicant Privacy Policy https://cdn.openai.com/policies/global-employee-and-contractor-privacy-policy.pdf
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.