远程工作雷达

高级软件工程师,AI评估

Senior Software Engineer, AI Evals

AI开发工程未标注地域
公司Sentry
薪资$155,000 - $400,000
工作地点San Francisco, California
地域资格未标注地域
时区要求无特别要求
用工类型FullTime
发布时间2026-01-28
数据来源Ashby
前往企业招聘页投递 →

关于Sentry

软件驱动世界,速度比以往任何时候都更快。Sentry帮助开发者在用户察觉之前修复错误和性能问题,让团队能减少灭火时间,更多时间用于构建产品。

被200,000+家机构信赖,Sentry是当今应用监控的标准,我们的团队正在打造其AI原生的未来。

关于该职位

作为Sentry AI/ML团队的高级软件工程师,你将负责构建评估基础设施,以衡量我们AI系统的准确性、可靠性及实际性能。这个角色对于确保我们的调试代理和AI驱动功能在扩展时正确、安全且可预测至关重要。你将设计数据集、基准测试和测试框架,将模糊的AI行为转化为可测量的信号,帮助团队有信心地推出AI产品。

在此职位中,你将:

- 设计并构建稳健的评估框架,用于衡量AI系统中的准确性、可靠性、回归和边缘情况
- 创建并整理高质量的数据集、黄金测试用例和基于真实生产数据的基准测试
- 构建自动化测试框架和指标流水线,持续评估模型、提示词和代理工作流
- 与应用AI工程师和产品负责人紧密合作,定义“良好”的标准,并将其转化为可衡量的标准
- 负责主要AI项目的评估生命周期,从早期实验到生产监控

如果你:

- 非常重视AI系统中的正确性、严谨性和度量
- 喜欢将模糊的产品目标和模型行为转化为具体的测试和指标
- 喜欢构建基础架构,为整个AI团队加快迭代速度并提高信心
- 在跨职能环境中茁壮成长,并享受通过更好的评估影响模型设计

资格要求

- 至少5年以上相关工作经验,计算机科学、机器学习或相关领域的学士学位
- 有构建测试、评估或数据基础设施的经验,AI/ML经验优先
- 能够编写生产级代码(我们使用Python和TypeScript)
- 有处理结构化和非结构化数据集、标注流程或数据质量流水线的经验
- 熟悉现代机器学习系统和评估技术(例如,离线指标、在线评估、r

查看英文原文

ABOUT SENTRY

Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building.

Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future.

ABOUT THE ROLE

As a Senior Software Engineer on Sentry’s AI/ML team, you’ll be responsible for building the evaluation infrastructure that measures the accuracy, reliability, and real-world performance of our AI systems. This role is critical to ensuring that our debugging agents and AI-powered features behave correctly, safely, and predictably as they scale. You’ll design datasets, benchmarks, and test harnesses that turn ambiguous AI behavior into measurable signals, helping the team ship AI with confidence.

IN THIS ROLE YOU WILL

- Design and build robust evaluation frameworks to measure accuracy, reliability, regressions, and edge cases in AI systems

- Create and curate high-quality datasets, golden test cases, and benchmarks grounded in real production data

- Build automated test harnesses and metrics pipelines to continuously evaluate models, prompts, and agentic workflows

- Partner closely with applied AI engineers and product leaders to define what “good” looks like and translate it into measurable criteria

- Own the evaluation lifecycle for major AI initiatives, from early experimentation through production monitoring

YOU’LL LOVE THIS JOB IF YOU

- Care deeply about correctness, rigor, and measurement in AI systems

- Enjoy turning fuzzy product goals and model behavior into concrete tests and metrics

- Like building foundational infrastructure that unlocks faster iteration and higher confidence for the entire AI team

- Thrive in cross-functional environments and enjoy influencing model design through better evaluation

QUALIFICATIONS

- Minimum 5+ years of professional experience with a Bachelor’s degree in computer science, machine learning, or a related field

- Experience building testing, evaluation, or data infrastructure for complex systems (AI/ML experience strongly preferred)

- Comfort writing production-quality code (we use Python and TypeScript)

- Experience working with structured and unstructured datasets, labeling workflows, or data quality pipelines

- Familiarity with modern ML systems and evaluation techniques (e.g., offline metrics, online evaluation, regression testing for models or prompts)

- Bonus: experience evaluating LLMs, agentic systems, or AI-assisted developer tools

The base salary range (or hourly wage range, if applicable) that Sentry reasonably expects to pay for this position is $155,000 to $400,000 USD. A successful candidate’s actual base salary (or hourly wage) amount will be determined by a variety of relevant factors including, without limitation, the candidate’s work location, education, work and other relevant experience, skills, and job-related knowledge. A successful candidate will be eligible to participate in Sentry’s employee benefit plans/programs applicable to the candidate’s position (including incentive compensation, equity grants, paid time off, and group health insurance coverage). See Sentry Benefits https://sentry.io/careers/ for more details about the Company’s benefit plans/programs.

EQUAL OPPORTUNITY AT SENTRY

Sentry is committed to providing equal employment opportunities to its employees and candidates for employment regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, or other legally-protected characteristic. This commitment includes the provision of reasonable accommodations to employees and candidates for employment with physical or mental disabilities who require such accommodations in order to (a) perform the essential functions of their jobs, or (b) seek employment with Sentry. We strive to build a diverse team, with an inclusive culture where every teammate can thrive. Sentry is an open-source company because we believe that everyone, everywhere, should have the ability and tools to make great software. Software should be accessible. That starts with making our industry accessible.

If you need assistance or an accommodation due to a disability, you may contact us at accommodations@sentry.io.

Want to learn more about how Sentry handles applicant data? Get the details in our Applicant Privacy Policy https://sentry.io/careers/applicantprivacy/.

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

高级技术客户成功经理

SentrySan Francisco, California$200,000 - $250,000FullTime2026-01-29
市场运营全球可投(据职位描述推断)

计费平台工程经理

SentrySan Francisco, California$220,000 - $450,000FullTime3 天前
开发工程未标注地域

← 返回全部职位