高级应用研究科学家 - 个性化
Senior Applied Research Scientist - Personalization
个性化团队让每位听众决定接下来听什么变得更简单、更愉快。从Blend到Discover Weekly,我们打造了Spotify最受欢迎的功能之一。我们通过比任何人都更深入地理解音乐和播客世界来实现这些功能。加入我们,你将通过为每一位用户做出优秀的推荐,持续吸引数百万用户收听。
在个性化团队中,Speak团队负责Spotify最先进的语音模型的开发,包括语音识别、语音合成以及语音到语音的模型。我们打造的语音模型能够达到人类级别的情感表达,从而深度与听众互动,并大规模支持创作者。我们在语音合成方面的开创性工作依赖于最先进的深度学习方法和评估技术,高效的数据处理和模型服务,以及从我们的语音人才库中捕捉高质量的音频。
我们正在寻找一位有经验的高级应用研究科学家,具备开发新颖机器学习技术和架构的经验,并对跨完整生产流程的工作有浓厚兴趣,以产出最先进的生成式对话语音到语音模型。你将与我们的工程团队合作,帮助开发我们的生产流程,探索新的想法和方法以提升质量、理解和真实感,同时推动我们的语音技术所能实现的边界。
Spotify是一个平等机会雇主。无论你来自哪里,长什么样,耳机里播放的是什么,你都可以在Spotify找到属于自己的位置。我们的平台属于每个人,我们的工作场所也是如此。我们业务中代表和放大越多的声音,我们就能更加茁壮成长、贡献力量并保持前瞻性!所以,请带来你的个人经历、观点和背景。正是在我们的差异中,我们将找到不断革新世界聆听方式的力量。
在Spotify,我们热衷于包容性,并确保整个招聘过程对每个人都是可访问的。我们在面试过程中提供合理的便利措施,并协助满足你的需求。如果你在申请或面试过程的任何阶段需要便利措施,请告诉我们——我们会尽一切努力支持你。
你将负责
- 开发和实验新的语音合成和语音相关的方法
查看英文原文
he Personalization team makes deciding what to play next easier and more enjoyable for every listener. From Blend to Discover Weekly, we're behind some of Spotify's most-loved features. We built them by understanding the world of music and podcasts better than anyone else. Join us and you'll keep millions of users listening by making great recommendations to each and every one of them.
Within Personalization, the Speak Team owns the development of Spotify's state-of-the-art speech models, contributing to speech recognition, speech synthesis, and speech-to-speech models. We craft voice models that match human-level emotional expressiveness, so we can deeply engage our listeners and support creators at scale. Our groundbreaking work on speech synthesis relies on state-of-the-art deep learning methods and evaluation techniques, highly efficient data processing and model serving, and capturing audio of outstanding quality from our voice talent pool.
We're looking for a senior applied research scientist with experience in developing novel ML techniques and architectures and with a strong interest in working across a full production pipeline to produce state-of-the-art generative conversational speech-to-speech models. You'll collaborate with our engineering teams to help develop our production pipelines, explore new ideas and methods to improve quality, understanding and realism, as well as push the frontiers of what is possible with our speech technology.
Spotify is an equal opportunity employer. You are welcome at Spotify for who you are, no matter where you come from, what you look like, or what’s playing in your headphones. Our platform is for everyone, and so is our workplace. The more voices we have represented and amplified in our business, the more we will all thrive, contribute, and be forward-thinking! So bring us your personal experience, your perspectives, and your background. It’s in our differences that we will find the power to keep revolutionizing the way the world listens.
At Spotify, we are passionate about inclusivity and making sure our entire recruitment process is accessible to everyone. We have ways to request reasonable accommodations during the interview process and help assist in what you need. If you need accommodations at any stage of the application or interview process, please let us know - we’re here to support you in any way we can.
What You'll Do
- Develop and experiment with new methods for speech synthesis and speech recognition, along with end-to-end approaches, building on the latest research and ideas.
- Work towards the expansion of our speech use-cases targeting different markets and products.
- Be part of a highly motivated research team dedicated to building and creating models at scale to power the Spotify platform.
- Champion best practices for research and development, sharing your knowledge and experience with other researchers within Speak.
- Collaborate with our engineering and data teams on ideas requiring new infrastructure or new high-quality data, as well as to help improve our speech recognition and speech synthesis pipelines, and help turn proven ideas into scalable products.
Who You Are
- You have a strong background in ML (PhD degree on top of professional experience), and
- experience in working with any of the following: transformers, GANs, diffusion models, flow matching, VAEs, audio codecs.
- You have experience in developing generative models for speech synthesis, speech recognition, audio/music, natural language processing, or computer vision.
- You have strong experience with Python, particularly PyTorch.
- You have strong communication skills and the ability to explain technical ideas with clarity to technical and non-technical people alike.
- You have experience in an academic or professional setting conducting high-quality research.
Where You'll Be
- This role is based in London or Stockholm.
- We offer you the flexibility to work where you work best! There will be some in person meetings, but still allows for flexibility to work from home