远程工作雷达

软件工程师 - GPU内核

Software Engineer - GPU Kernels

开发工程未标注地域
公司Baseten
薪资$180,000 - $360,000
工作地点San Francisco / Toronto / New York / Montreal
地域资格未标注地域
时区要求无特别要求
用工类型FullTime
发布时间2025-07-17
数据来源Ashby
前往企业招聘页投递 →

ABOUT BASETEN

Baseten 为全球最具活力的 AI 公司提供关键推理支持,如 Cursor、Notion、OpenEvidence、Abridge、Clay、Gamma 和 Writer。通过结合应用 AI 研究、灵活的基础架构和无缝的开发者工具,我们使处于 AI 前沿的公司能够将前沿模型投入生产。我们正在快速成长,并最近完成了 15 亿美元 F 轮融资 https://www.baseten.co/blog/announcing-our-series-f/,由 Altimeter Capital、Conviction Partners 和 Spark Capital 领投。加入我们,帮助构建工程师们用来发布 AI 产品的平台。

THE ROLE

我们正在寻找一名 GPU 内核工程师加入我们的团队,参与 AI 加速的最前沿工作,你的代码将直接影响最先进的机器学习模型的性能。作为 GPU 内核工程师,你将打造支撑现代 AI 工作负载的基础,优化每一微秒的计算,以实现突破性应用。

你将在一个快节奏、充满智力挑战的环境中工作,技术卓越是首要任务,你的贡献将直接影响服务于数百万用户的多个产品中的生产系统。这个职位为对底层优化和高影响力系统工作充满热情的工程师提供了出色的成长机会。

EXAMPLE INITIATIVES

你将作为我们模型性能团队的一员,参与以下类型的项目:

- Baseten 嵌入式推理:目前最快的嵌入式解决方案 https://www.baseten.co/blog/introducing-baseten-embeddings-inference-bei/

- Baseten 推理栈 https://www.baseten.co/resources/guide/the-baseten-inference-stack/

- 推动模型性能优化 https://www.baseten.co/blog/driving-model-performance-optimization-2024-highlights/

RESPONSIBILITIES

核心工程职责

- 为关键 ML 操作(包括矩阵乘法、注意力机制和专家混合路由)设计和实现高性能 GPU 内核

- 使用 CUDA、PTX 汇编和特定架构技术编写和优化代码

- 应用高级性能优化方法,如内存合并、warp 级编程、张量核心加速和计算/内存重叠

性能与创新

- 实现量化(FP8/FP4)、稀疏性以及计算/通信重叠等前沿功能

- 使用 Nsight Systems、Nsight Compute 等工具识别并解决性能瓶颈

查看英文原文

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.

THE ROLE

We’re seeking a GPU Kernel Engineer to join our team at the cutting edge of AI acceleration, where your code directly impacts the performance of state-of-the-art machine learning models. As a GPU Kernel Engineer, you'll craft the foundation that powers modern AI workloads, optimizing every microsecond of computation to enable breakthrough applications.

You'll work in a fast-paced, intellectually stimulating environment where technical excellence is paramount and your contributions directly influence production systems serving millions of users across numerous products. This role offers exceptional growth potential for engineers passionate about low-level optimization and high-impact systems work.

EXAMPLE INITIATIVES

You'll get to work on these types of projects as part of our Model Performance team:

- Baseten Embeddings Inference: The fastest embeddings solution available https://www.baseten.co/blog/introducing-baseten-embeddings-inference-bei/

- The Baseten Inference Stack https://www.baseten.co/resources/guide/the-baseten-inference-stack/

- Driving model performance optimization https://www.baseten.co/blog/driving-model-performance-optimization-2024-highlights/

RESPONSIBILITIES

Core Engineering Responsibilities

- Design and implement high-performance GPU kernels for key ML operations, including matrix multiplications, attention mechanisms, and mixture-of-experts routing

- Write and optimize code using CUDA, PTX assembly, and architecture-specific techniques

- Apply advanced performance optimization methods such as memory coalescing, warp-level programming, tensor core acceleration, and compute/memory overlap

Performance & Innovation

- Implement cutting-edge features like quantization (FP8/FP4), sparsity, and compute/communication overlap

- Identify and resolve performance bottlenecks using tools like Nsight Systems, Nsight Compute, and Torch Profiler

- Collaborate with research teams to productionize theoretical advancements

Impact & Collaboration

- Contribute to internal and open-source GPU libraries

- Present technical contributions at industry conferences (e.g., NVIDIA GTC, AWS re:Invent)

REQUIREMENTS

- Strong understanding of GPU architecture and programming paradigms:

- Memory hierarchy (global, shared, registers, L1/L2 cache)

- Thread/block/grid organization

- Synchronization techniques and race condition mitigation

- Proficient in C++ and GPU performance profiling tools

- Knowledge of:

- CUDA C++ API

- Memory access patterns and bandwidth optimization

- Numerical precision and quantization strategies

- Modern GPU features (e.g., tensor cores, async operations)

NICE TO HAVE

- Experience with Transformer models and attention optimization (e.g., Flash Attention)

- Familiarity with GPU kernel libraries: Cutlass, Triton, Thrust, CUB

- Background in GEMM tuning and distributed/multi-GPU compute

- Contributions to open-source GPU projects

- Research publications or conference presentations on GPU performance

BENEFITS

- Competitive compensation, including meaningful equity

- (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents

- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

- Paid parental leave

- Fertility and family-building stipend through Carrot

- Company-facilitated 401(k)

- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

AI推理工程师

BasetenSan Francisco / Remote / Toronto / New Yor$165,000 - $330,000FullTime2026-08-03
AI开发工程全球可投

← 返回全部职位