远程工作雷达

软件工程师 - 模型性能系统

Software Engineer - Model Performance Systems

AI开发工程未标注地域
公司Baseten
薪资$165,000 - $330,000
工作地点San Francisco / Toronto / New York / Montreal
地域资格未标注地域
时区要求无特别要求
用工类型FullTime
发布时间2026-01-07
数据来源Ashby
前往企业招聘页投递 →

关于 BASETEN

Baseten 为全球最具活力的 AI 公司提供关键任务推理支持,例如 Cursor、Notion、OpenEvidence、Abridge、Clay、Gamma 和 Writer。通过结合应用 AI 研究、灵活的基础架构和无缝的开发者工具,我们使处于 AI 前沿的公司能够将前沿模型投入生产。我们正在快速成长,并最近完成了 1.5 亿美元的 F 轮融资 https://www.baseten.co/blog/announcing-our-series-f/,由 Altimeter Capital、Conviction Partners 和 Spark Capital 领投。加入我们,帮助构建工程师们用来部署 AI 产品的平台。

职位描述

我们正在寻找软件工程师加入我们的团队。这是一个专业性强、影响力大的职位,位于高性能计算(HPC)和大语言模型(LLM)工程的交叉点。你不仅会构建我们下一代 AI 基础架构的自动化“速度计和诊断”套件;你还将定义路线图,推动关键技术决策,并全面负责这项工作的未来。

职责

- 基准测试:评估、运行并自动化标准 LLM 质量基准(GSM8K、MMLU)以及针对特定工作负载的自定义性能套件(例如,长上下文窗口、KV 缓存重用、解耦服务)。

- 开发体验改进:开发和维护内部 GPU 支持的开发环境(类似于 GitHub Codespaces)。你将确保团队拥有无缝、高性能的“开发机器”,优化模型实验。

- 工具开发:构建并贡献开源工具,如 InferenceMAX 和 genai-bench,以自动化模型评估、基准测试和分析。

- 系统分析:使用 PyTorch Profiler、NVIDIA Nsight Systems 和 py-spy 等分析工具收集性能分析数据,识别瓶颈,并调试计算/网络堆栈。

- 监控与可观测性:开发实时仪表盘和警报,监控系统健康状况、模型启动时间和运行时性能。

- 持续集成:通过 CI/CD 流水线自动化性能测试,捕捉回归问题,并为模型运行时堆栈构建发布流程自动化。

- 优化自动化:构建工具以找到“帕累托前沿”——为给定模型和工作负载确定最佳配置(延迟 vs 成本 vs 质量)。

要求

这是一个中高级、高杠杆的职位。我们关注你的技术深度和良好的沟通能力。

查看英文原文

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.

THE ROLE

We are looking for Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work.

RESPONSIBILITIES

- Benchmarking: Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving).

- DevEx Improvement: Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation.

- Tool Development: Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis.

- System Profiling: Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack.

- Monitoring & Observability: Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance.

- Continuous Integration: Automate performance testing via CI/CD pipelines to catch regressions and build release workflow automation for the model runtimes stack.

- Optimization Automation: Build tools to find the "Pareto frontier"—identifying the absolute best configuration (latency vs. cost vs. quality) for a given model and workload.

REQUIREMENTS

This is a mid-senior, high leverage role. We care about your technical depth, strong communication skills to drive cross-team efforts, ability to navigate vague requirements and mentor other engineers. We want to talk to you if you have:

- A Love for Systems & Hardware: You aren’t just interested in the AI; you want to understand GPU memory subsystems, InfiniBand, and how data moves across a cluster.

- An Automation Mindset: You believe that if a task has to be done twice, it should be scripted. You have a passion for stress-testing and fuzzy testing to find the "breaking point" of a system.

- Mathematical Curiosity: A desire to understand the underlying math of Transformers and how it translates into FLOPs and memory requirements.

- Technical Toolkit: Familiarity with Python, and an eagerness to master the NVIDIA software stack. C++ familiarity is good to have.

WHY THIS ROLE

- Direct Impact: Your tools will be the gatekeeper for what defines "good" performance for our customers.

- Deep Learning (Literally): You will gain world-class expertise in GPU orchestration and LLM inference that few engineers in the industry possess.

- High Ownership: As the lead of a small team, you will have the autonomy to build tools from scratch and contribute to open-source projects.

BENEFITS

- Competitive compensation, including meaningful equity

- (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents

- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

- Paid parental leave

- Fertility and family-building stipend through Carrot

- Company-facilitated 401(k)

- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

AI推理工程师

BasetenSan Francisco / Remote / Toronto / New Yor$165,000 - $330,000FullTime2026-08-03
AI开发工程全球可投

← 返回全部职位