研究工程师 - 推理
Research Engineer - Inference
关于ElevenLabs
ElevenLabs是一家人工智能研究和产品公司,正在改变我们与技术互动的方式。
我们于2023年1月成立,推出了首个类人AI语音模型。如今,我们为数百万用户和数千家企业提供服务——从快速增长的初创公司到德意志电信和Meta等大型企业。我们的投资者包括安德里森·霍罗维茨、ICONIQ Growth和红杉资本等全球知名机构。我们已筹集7.81亿美元资金,最新估值为110亿美元——始终与11相关。
我们已从语音扩展到三个主要平台:
- ElevenAgents使企业能够提供无缝且智能的客户体验,具备集成、测试、监控和可靠性,以大规模部署语音和聊天代理。
- ElevenCreative赋予创作者和营销人员能力,可在70多种语言中生成和编辑语音、音乐、图像和视频。
- ElevenAPI为开发者提供对我们的领先AI音频基础模型的访问权限。
我们所做的一切都是团队创造力和奉献精神的结果——建设者们正在做他们人生中最出色的工作。我们是研究人员、工程师和运营人员。IOI奖牌获得者和前创始人。如果你希望努力工作并创造持久的积极影响,我们期待你的加入。
我们如何工作
- 高速度:快速实验、精简自主团队和最小化官僚主义。
- 以影响力而非职位名称来衡量:我们没有职位名称。而是看你的影响力。没有任务高于或低于你。
- 以AI为先:我们使用AI以更高的质量更快地完成工作。我们在整个公司都这样做——从工程到增长再到运营。
- 无处不在的卓越:我们所做的每件事都应该与我们的AI模型质量相匹配。
- 全球团队:我们优先考虑你的才能,而不是你的位置。
我们提供的福利
- 创新文化:你将参与定义AI发展轨迹的历史性机会,身边有不断突破可能性边界的团队。
- 成长路径:加入ElevenLabs意味着加入一个充满活力的团队,有无数机会推动影响力——超越你的直接角色和职责。
- 学习与发展:ElevenLabs主动通过年度可支配津贴支持职业发展。
- 社交旅行:我们还提供年度可支配津贴,让你每年以自己喜欢的方式与同事见面。
- 年度公司集训:每年,我们都会在新的地点将整个团队聚集在一起。
查看英文原文
ABOUT ELEVENLABS
ElevenLabs is an AI research and product company transforming how we interact with technology.
We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world's most prominent, including Andreessen Horowitz, ICONIQ Growth and Sequoia. We've raised $781M in funding and our last valuation was $11B - multiples of 11, always.
We have expanded from voice into three main platforms:
- ElevenAgents enables businesses to deliver seamless and intelligent customer experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale.
- ElevenCreative empowers creators and marketers to generate and edit speech, music, image, and video across 70+ languages.
- ElevenAPI gives developers access to our leading AI audio foundational models.
Everything we do is the result of the creativity and commitment of our team - builders doing the best work of their lives. We are researchers, engineers, and operators. IOI medalists and ex-founders. If you want to work hard and create lasting positive impact, we want to hear from you.
HOW WE WORK
- High-velocity: Rapid experimentation, lean autonomous teams, and minimal bureaucracy.
- Impact not job titles: We don’t have job titles. Instead, it’s about the impact you have. No task is above or beneath you.
- AI first: We use AI to move faster with higher-quality results. We do this across the whole company—from engineering to growth to operations.
- Excellence everywhere: Everything we do should match the quality of our AI models.
- Global team: We prioritize your talent, not your location.
WHAT WE OFFER
- Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible.
- Growth paths: Joining ElevenLabs means joining a dynamic team with countless opportunities to drive impact - beyond your immediate role and responsibilities.
- Learning & development: ElevenLabs proactively supports professional development through an annual discretionary stipend.
- Social travel: We also provide an annual discretionary stipend to meet up with colleagues each year, however you choose.
- Annual company offsite: Each year, we bring the entire team together in a new location - past offsites have included Croatia and Italy.
- Co-working: If you’re not located near one of our main hubs, we offer a monthly co-working stipend.
ABOUT THE ROLE
We are looking for a Research Engineer to join the research team at ElevenLabs, focused on deploying and optimizing our frontier AI models in production. The quality of our models only matters if they can be served fast, reliably, and at scale. You will own the systems that turn research breakthroughs into real-time products used by millions. You will thrive in this role if you enjoy:
- Deploying state-of-the-art models to production and owning the path from research checkpoint to serving infrastructure.
- Optimizing inference performance across the stack, including latency, throughput, and cost, using techniques such as quantization, distillation, KV-cache optimization, batching strategies, and custom kernels.
- Building and tuning high-performance serving systems for real-time, streaming workloads where every millisecond matters.
- Creating tooling and infrastructure that lets researchers ship new models to production quickly, safely, and with confidence in their performance characteristics.
REQUIREMENTS
We do not require any formal certifications or degrees. Instead, we are seeking enthusiastic engineers who can showcase solving impressively hard problems with artifacts such as past projects, designs, or GitHub contributions. Ideally, you bring:
- Experience deploying and serving ML models in production, ideally for latency-sensitive or real-time applications.
- Strong engineering skills in GPU programming and inference optimization (e.g., CUDA, Triton, TensorRT, or serving frameworks such as vLLM or SGLang).
- The capacity to autonomously profile, diagnose, and eliminate bottlenecks across the serving stack, from model architecture to kernels to orchestration, and to build the tooling to measure it.
LOCATION
This role is remote and can be executed globally. If you prefer, you can work from our offices in London, New York, San Francisco, and Warsaw.
#LI-Remote
We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or other legally protected statuses.