远程工作雷达

前向部署工程工程经理(大语言模型)

Engineering Manager, Forward Deployed Engineering (LLM)

AI未标注地域
公司Baseten
薪资$260,000 - $380,000
工作地点San Francisco / Toronto / New York / Montreal
地域资格未标注地域
时区要求无特别要求
用工类型FullTime
发布时间2026-05-08
数据来源Ashby
前往企业招聘页投递 →

ABOUT BASETEN

Baseten 为全球最具活力的 AI 公司提供关键推理支持,如 Cursor、Notion、OpenEvidence、Abridge、Clay、Gamma 和 Writer。通过结合应用 AI 研究、灵活的基础架构和无缝的开发者工具,我们使处于 AI 前沿的公司能够将前沿模型投入生产。我们正在快速成长,并最近完成了 1.5 亿美元的 F 轮融资 https://www.baseten.co/blog/announcing-our-series-f/,由 Altimeter Capital、Conviction Partners 和 Spark Capital 领投。加入我们,帮助构建工程师们用来交付 AI 产品的平台。

THE ROLE

作为工程经理(玩家与教练),你将领导并指导一支专注于为 Baseten 客户构建、扩展和优化 LLM 推理工作负载的 Forward Deployed Engineers 团队。通过技术上的亲自负责和管理领导力,你将引导团队完成在 Baseten 平台上设计、部署和管理高性能、低延迟 AI 应用的流程。Baseten 的 FDE 不是销售职能——我们是工程、产品和客户架构师的混合体,他们参与核心 Baseten 代码库的开发,推动我们功能路线图的大部分内容,并执行复杂的客户项目。

你还将与产品、基础架构和其他客户工程团队合作,确保大型语言模型(LLM)和其他生成式 AI 系统在生产环境中实现最佳性能、可靠性和成本效率。

EXAMPLE INITIATIVES

看看我们 Forward Deployed Engineering 团队成员撰写的这些博客文章:

- Forward Deployed Engineering on the frontier of AI https://www.baseten.co/blog/forward-deployed-engineering/

- The fastest, most accurate Whisper transcription https://www.baseten.co/blog/the-fastest-most-accurate-and-cost-efficient-whisper-transcription/

- Deploy production-ready model servers from Docker images https://www.baseten.co/blog/deploy-production-model-servers-from-docker-images/

- Deploy custom ComfyUI workflows as APIs https://www.baseten.co/blog/deploying-custom-comfyui-workflows-as-apis/

RESPONSIBILITIES

Leadership & Team Management

- 领导、指导和培养一支 Forward Deployed Engineers 团队,提供技术方向、项目执行和职业发展的指导。

- 设定明确的目标,并确保在多个面向客户的项目中按时高质量地交付。

查看英文原文

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.

THE ROLE

As an Engineering Manager (Player & Coach), you will lead and mentor a team of Forward Deployed Engineers focused on building, scaling, and optimizing LLM inference workloads for Baseten customers. Applying both hands-on technical ownership and managerial leadership, you will guide your team through the processes of designing, deploying, and managing high performance, low latency AI applications on Baseten’s platform. FDE at Baseten is not a sales function – we are a mix of engineering, product, and customer architects who contribute to the core Baseten codebase, drive large portions of our feature roadmap, and execute on complicated customer engagements.

You will also partner with product, infrastructure, and other customer engineering teams to ensure that large language models (LLMs) and other generative AI systems deliver best-in-class performance, reliability, and cost efficiency in production environments.

EXAMPLE INITIATIVES

Take a look at these blog posts written by members of our Forward Deployed Engineering team:

- Forward Deployed Engineering on the frontier of AI https://www.baseten.co/blog/forward-deployed-engineering/

- The fastest, most accurate Whisper transcription https://www.baseten.co/blog/the-fastest-most-accurate-and-cost-efficient-whisper-transcription/

- Deploy production-ready model servers from Docker images https://www.baseten.co/blog/deploy-production-model-servers-from-docker-images/

- Deploy custom ComfyUI workflows as APIs https://www.baseten.co/blog/deploying-custom-comfyui-workflows-as-apis/

RESPONSIBILITIES

Leadership & Team Management

- Lead, mentor, and grow a team of Forward Deployed Engineers, providing guidance on technical direction, project execution, and professional development.

- Set clear goals and ensure timely, high-quality delivery across multiple customer-facing projects involving LLM deployment and inference optimization.

- Collaborate with leadership to align team priorities with company and customer goals, balancing short-term delivery, widely varying customer priorities, and long-term technical initiatives.

- Player-coach – While much of this role will be leading the team, you will also be expected to be a key driver on strategic product initiatives and customer engagements. The best managers derive credibility from being able to be hands-on when needed.

Technical Ownership

- Develop and maintain software systems and product features using one or more general-purpose programming languages in a production-level environment, with a preference for Python due to its relevance in ML projects.

- Drive customer impact by designing, implementing, and deploying Baseten solutions end-to-end (problem framing → evaluation → production deployment → monitoring). This involves working with customers’ engineering teams at every stage of the customer journey including: sales, implementation, and expansion.

- Deliver with velocity: turn vague objectives into clear specs and well-defined PoCs so we can rapidly ship well-tested services and outcomes for our customers

- Optimize and enhance AI/ML projects, contributing to the continuous improvement of our technical stack. This includes developing features and PRDs with other engineering and product orgs.

- Own products and customer projects end-to-end, functioning as both an engineer, project manager, and product manager, with a focus on user empathy, project specification, and end-to-end execution.

REQUIREMENTS

- Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, or related field.

- 4+ years of professional software engineering experience, including 1+ year in a leadership or mentorship capacity.

- Strong programming skills in Python, with production experience in building or optimizing ML inference systems.

- Proven experience with LLMs, inference optimization, or serving frameworks (e.g., vLLM, TensorRT, Triton, Hugging Face, Ray Serve).

- Familiarity with observability, profiling, and cost/performance tradeoffs in production ML systems.

- Excellent communication and collaboration skills—able to lead cross-functional efforts and drive outcomes in ambiguous, fast-paced environments.

NICE TO HAVE

- Experience leading customer-facing engineering teams or working directly with enterprise partners.

- Deep understanding of GPU infrastructure, distributed inference, or model compression techniques.

BENEFITS

- Competitive compensation, including meaningful equity

- (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents

- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

- Paid parental leave

- Fertility and family-building stipend through Carrot

- Company-facilitated 401(k)

- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

AI推理工程师

BasetenSan Francisco / Remote / Toronto / New Yor$165,000 - $330,000FullTime2026-08-03
AI开发工程全球可投

← 返回全部职位