资深可靠性科学家
Principal Reliability Scientist
**关于我们**
Graphcore 是人工智能计算领域全球领先的创新者之一。
公司正在开发硬件、软件和系统基础设施,以解锁下一代人工智能突破,并推动人工智能解决方案在各个行业的广泛应用。
作为软银集团的一部分,Graphcore 是一家精英公司的成员,这些公司负责世界上一些最具变革性的技术。他们共同拥有一个大胆的愿景:实现人工智能超级智能,并确保其好处对每个人都是可及的。
Graphcore 的团队来自多样化的背景,带来了广泛的专业技能和视角。由人工智能研究专家、芯片设计人员、软件工程师和系统架构师组成,Graphcore 拥有持续学习和不断创新的文化。
**职位概述**
向制造运营部门的质量领导汇报,高级可靠性科学家负责在复杂、高性能系统中领导可靠性活动。与已有的可靠性专家和跨职能团队紧密合作,该职位利用实验数据和高级建模来支持设计决策,验证产品可靠性并优化可服务性策略,包括备件供应。
**团队介绍**
制造运营部门的质量团队负责确保 Graphcore 硬件产品组合的稳健性、可靠性和生命周期性能。该团队包括经验丰富的可靠性专家,并与技术研究、芯片、电路板、系统设计、平台和运营团队密切合作,将可靠性见解转化为产品生命周期中的可操作改进。
**职责与工作内容:**
- 与研究和设计团队合作,在芯片、电路板和系统层面定义和优化可靠性要求
- 将先进的可靠性方法应用于高度创新的系统,包括与液冷架构和流体动力学相关的挑战
- 设计和执行实验,生成高质量的可靠性和性能数据,确保统计严谨性和相关性
- 分析实验、现场和制造数据,量化可靠性指标,如平均无故障时间(MTBF)、平均修复时间(MTTR)、可靠性可用性服务(RAS)特性以及软错误率(SER)
- 利用数据驱动的洞察,支持产品设计权衡、可靠性目标和备件供应
查看英文原文
**About us**
Graphcore is one of the world’s leading innovators in Artificial Intelligence compute.
It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry.
As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone.
Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation.
**Job Summary**
Reporting to the Quality leadership within Manufacturing Operations, the Senior Reliability Scientist is responsible for leading reliability activities across complex, high-performance systems. Working closely with established reliability experts and cross-functional teams, this role uses experimental data and advanced modelling to inform design decisions, validate product reliability and optimise serviceability strategies, including spares provisioning.
**The Team**
The Quality team within Manufacturing Operations is responsible for ensuring product robustness, reliability and lifecycle performance across Graphcore’s hardware portfolio. The team includes experienced reliability specialists and works closely with technology research, chip, board, system design, platform and operations teams to translate reliability insights into actionable improvements across the product lifecycle.
**Responsibilities and Duties:**
- Define and refine reliability requirements across silicon, board and system levels, working in partnership with research and design teams
- Apply advanced reliability methodologies to highly innovative systems, including challenges associated with liquid-cooled architectures and fluid dynamics
- Design and execute experiments to generate high-quality reliability and performance data, ensuring statistical rigour and relevance
- Analyse experimental, field and manufacturing data to quantify reliability metrics such as MTBF, MTTR, RAS characteristics and soft error rates (SER)
- Use data-driven insights to inform product design trade-offs, reliability targets and spares provisioning strategies
- Collaborate with chip, board and system design teams to influence architecture and component selection based on reliability considerations
- Support development of system-level reliability models incorporating thermal, mechanical and fluid behaviour
- Lead complex root cause investigations into reliability issues, driving corrective and preventative actions across teams
- Contribute to the evolution of reliability tools, processes and best practices within the organisation
- Communicate complex reliability concepts, risks and recommendations clearly to a wide range of stakeholders
Qualifications:
- Strong background in reliability engineering or reliability science within semiconductor, hardware or complex systems environments
- Experience of physics-of-failure approaches in high-performance computing, AI hardware or related domains
- Experience with reliability modelling, experimental design and statistical data analysis
- Proven ability to work with and interpret experimental reliability data to drive engineering decisions
- Experience with key reliability metrics such as MTBF, MTTR, RAS and failure rate analysis
- Ability to operate effectively in complex, cross-functional environments with multiple stakeholders
- Strong problem-solving skills with the ability to lead technically challenging investigations independently
- Excellent communication skills, with the ability to influence design and operations teams using data-driven insights
Preferred Qualification:
- Experience with liquid cooling systems, fluid dynamics or thermally complex hardware environments
- Knowledge of soft error mechanisms and SER modeling
- Experience contributing to reliability strategy, processes or tooling improvements