软件工程师 - 模型产品
Software Engineer - Model Products
关于Baseten
Baseten为全球最具活力的AI公司提供关键任务的推理服务,如Cursor、Notion、OpenEvidence、Abridge、Clay、Gamma和Writer。通过结合应用AI研究、灵活的基础设施和无缝的开发者工具,我们使处于AI前沿的公司能够将最前沿的模型投入生产。我们正在快速成长,并最近完成了15亿美元的F轮融资,由Altimeter Capital、Conviction Partners和Spark Capital领投。加入我们,帮助构建工程师们用来发布AI产品的平台。
职位描述
Baseten的模型性能(MP)团队负责确保运行在我们平台上的模型速度快、可靠且成本高效。作为该团队的一员,你将专注于模型API——为最新开源模型提供托管API端点的基础架构。这项工作涵盖分布式系统、模型服务和开发者体验。你将加入一个小型但影响深远的团队,在产品、模型性能和基础设施的交汇点上工作,帮助定义开发者如何大规模与AI模型交互。
职责
- 设计、构建和运营模型API界面,重点关注高级推理功能:结构化输出(JSON模式、语法约束生成)、工具/函数调用和多模态服务
- 分析和优化TensorRT-LLM内核,分析CUDA内核性能,实现自定义CUDA操作符,调整内存分配模式以实现最大吞吐量,并优化多GPU设置中的通信模式
- 在深入理解其内部机制的基础上,将性能改进产品化:推测解码实现、结构化输出的引导生成、高性能服务的自定义调度和路由算法
- 构建全面的基准测试框架,衡量不同模型架构、批处理大小、序列长度和硬件配置下的实际性能
- 将性能改进产品化到各种运行时(例如TensorRT、TensorRT-LLM):推测解码、量化、批处理和KV缓存复用
- 实现深度可观测性(指标、追踪、日志),并构建可重复的基准测试,以衡量速度、可靠性和质量
- 实现平台基础功能:API版本控制、验证、使用计量、配额和认证
- 与其他团队紧密合作以交付成果
查看英文原文
ABOUT BASETEN
Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.
THE ROLE
Baseten’s Model Performance (MP) team is responsible for ensuring the models running on our platform are fast, reliable, and cost‑efficient. As part of this team, you’ll focus on Model APIs — the infrastructure powering our hosted API endpoints for the latest open‑source models. This work spans distributed systems, model serving, and developer experience. You’ll join a small, high‑impact team operating at the intersection of product, model performance, and infra, helping to define how developers interact with AI models at scale.
RESPONSIBILITIES
- Design, build, and operate the Model APIs surface with focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling and multi-modal serving
- Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, implement custom CUDA operators, tune memory allocation patterns for maximum throughput and optimize communication patterns across multi-GPU setups
- Productionize performance improvements across runtimes with deep understanding of their internals: speculative decoding implementations, guided generation for structured outputs, custom scheduling and routing algorithms for high-performance serving
- Build comprehensive benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware configurations
- Productionize performance improvements across runtimes (e.g.TensorRT, TensorRT‑LLM): speculative decoding, quantization, batching, and KV‑cache reuse.
- Instrument deep observability (metrics, traces, logs) and build repeatable benchmarks to measure speed, reliability, and quality.
- Implement platform fundamentals: API versioning, validation, usage metering, quotas, and authentication.
- Collaborate closely with other teams to deliver robust, developer‑friendly model serving experiences.
REQUIREMENTS
- 3+ years experience building and operating distributed systems or large‑scale APIs.
- Prior ML or LLM experience is not required. We’re looking for strong systems engineers with performance instincts - experience with model serving or inference systems is a plus
- Proven track record of owning low‑latency, reliable backend services (rate‑limiting, auth, quotas, metering, migrations).
- Infra instincts with performance sensibilities: profiling, tracing, capacity planning, and SLO management.
- Comfortable debugging performance and reliability issues across complex systems, from application-level behavior to runtime and infrastructure internals.
- Strong written communication; able to produce clear design docs and collaborate across functions.
NICE TO HAVE
- Experience with LLM runtimes (vLLM, SGLang, TensorRT‑LLM) or contributions to open-source inference engines (vLLM, TensorRT-LLM, SGLang, TGI)
- Knowledge of Kubernetes, service meshes, API gateways, or distributed scheduling.
- Background in developer‑facing infrastructure or open‑source APIs.
BENEFITS
- Competitive compensation, including meaningful equity
- (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.
We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).