高级应用机器学习工程师(语音与音频)
Senior Applied ML Engineer (Speech & Audio)
项目概述
加入一个前沿项目,专注于为阿拉伯语市场构建先进的AI语音基础设施。该项目涉及开发最先进的阿拉伯语语音技术,包括:
· 自然文本转语音(TTS)
· 实时自动语音识别(ASR)
· 端到端语音对话系统
解决方案针对地区性阿拉伯语方言,包括埃及语、海湾语、黎凡特语等。
职位描述
我们正在寻找一位具有深厚语音和音频技术经验的高级应用机器学习工程师。在此职位中,您将设计、微调并优化用于阿拉伯语音应用的先进机器学习模型。您将参与完整的开发生命周期,从数据管道构建和模型实验到推理优化和生产部署。
此职位适合热衷于将前沿研究转化为可扩展、低延迟系统的工程师。
主要职责
· 使用阿拉伯语专用测试集对TTS和ASR模型进行基准测试和评估,测量词错误率(WER)、自然度和方言覆盖范围等指标。
· 微调生成模型用于语音克隆、零样本说话人适应和语音合成。
· 构建和维护以阿拉伯语为重点的数据管道,包括:
· 音频收集和预处理
· 声调标注(Tashkil)
· 数据清洗和增强
- 在生产环境中使用以下方法优化模型推理:
- 量化
- KV缓存调优
- 流式推理技术
- 集成并评估完整的语音到语音对话流程。
- 根据最新研究论文进行实验,并将成果转化为生产就绪的解决方案。
- 与工程和产品团队合作,部署稳健且可扩展的语音系统。
所需资格
· 5年以上机器学习、应用AI或AI研究经验。
· 精通Python编程。
· 在PyTorch和Hugging Face生态系统中有丰富的实践经验。
· 有训练和微调神经模型的经验,包括:
· 文本转语音(TTS)
· 自动语音识别(ASR)
· 音频编解码器
- 深入理解现代语音架构,如:
- Whisper
- Conformer
- HiFi-GAN
- 基于扩散的模型
- 有音频处理技术经验,包括:
- 语音活动检测(VAD)
- 说话人识别
查看英文原文
Project Overview
Join a cutting-edge initiative focused on building advanced AI voice infrastructure for Arabic-speaking markets. The project involves developing state-of-the-art Arabic speech technologies, including:
· Natural Text-to-Speech (TTS)
· Real-Time Automatic Speech Recognition (ASR)
· End-to-End Speech-to-Speech Conversational Systems
The solutions are tailored to regional Arabic dialects, including Egyptian, Gulf, Levantine, and others.
Job Description
We are seeking a highly skilled Senior Applied Machine Learning Engineer with deep expertise in speech and audio technologies. In this role, you will design, fine-tune, and optimize advanced machine learning models for Arabic voice applications. You will work across the full development lifecycle, from data pipeline construction and model experimentation to inference optimization and production deployment.
This position is ideal for engineers who are passionate about transforming cutting-edge research into scalable, low-latency systems that support natural and accurate Arabic speech interactions.
Key Responsibilities
· Benchmark and evaluate TTS and ASR models using Arabic-specific test sets, measuring metrics such as Word Error Rate (WER), naturalness, and dialect coverage.
· Fine-tune generative models for voice cloning, zero-shot speaker adaptation, and speech synthesis.
· Build and maintain Arabic-focused data pipelines, including:· Audio collection and preprocessing
· Diacritization (Tashkil)
· Data cleaning and augmentation
- Optimize model inference for production environments using:· Quantization
- KV-cache tuning
- Streaming inference techniques
- Integrate and evaluate complete speech-to-speech conversational pipelines.
- Conduct experiments based on recent research papers and convert findings into production-ready solutions.
- Collaborate with engineering and product teams to deploy robust and scalable speech systems.
Required Qualifications
· 5+ years of experience in Machine Learning, Applied AI, or AI Research.
· Strong programming skills in Python.
· Extensive hands-on experience with PyTorch and the Hugging Face ecosystem.
· Proven experience training and fine-tuning neural models for:· Text-to-Speech (TTS)
· Automatic Speech Recognition (ASR)
· Audio codecs
- Deep understanding of modern speech architectures such as:· Whisper
- Conformer
- HiFi-GAN
- Diffusion-based models
- Experience with audio processing techniques including:· Voice Activity Detection (VAD)
- Speaker Diarization
- Neural Vocoders
- Demonstrated ability to implement and adapt research papers into practical production experiments.
- Strong understanding of Arabic language challenges, including:· Diacritization (Tashkil)
- Dialectal variations
- Code-switching
- Experience with inference optimization techniques such as:· Quantization
- Streaming inference
- NVIDIA TensorRT
Preferred Qualifications
· Experience developing custom NVIDIA CUDA kernels for high-performance model inference.
· Familiarity with speculative decoding and other advanced acceleration techniques.
· Experience deploying models at scale in cloud or GPU-based production environments.
· Contributions to open-source speech or machine learning projects.
WHY YOU’LL LOVE US
· All employees benefits for free (our famous games room, daily breakfast, fruits, coffee and other hot drinks, soft drinks and juices, company days out and parties…)
· Social insurance
· Open-door management policy
· Full Medical insurance
· Accommodation and Transportation Allowance
· Friendly environment that values innovation and efficiency
· Exciting opportunities for career growth and talent development
· Feedback encouragement
· Recognition and reward programs
· Competitive salaries and incentives
· Friendly environment
· Flexible and Comfortable schedule
· Fun committees
· Monetary rewards
· Fun, smart and creative people
· Career possibilities with growing team
· Paid vacations
· Social benefits
For more information about Nile Bits, please visit our website:
https://www.nilebits.com