AI工程师-分类器、媒体智能与语音研发
AI Engineer - Classifiers, Media Intelligence & Voice R&D
这是一个远程职位。
我们的客户正在寻找一位有创新精神且充满动力的AI工程师加入他们的团队。他们是一家媒体智能和AI驱动内容创作的领先企业,最近在AI语音和图像技术领域进行了扩展,推动下一代尖端产品的开发。该职位将专注于大规模AI生成媒体的创建、分类和组织,同时领导AI语音和音频生成以及高级图像智能能力的研发。
职位描述
职责:
- 设计、训练和部署内容流水线的分类模型,包括风格检测、质量评分、内容审核、过滤和生成媒体的语义分类。
- 开发和维护媒体库的自动化标签和组织系统:提取属性、检测视觉特征、聚类相似内容,并实现智能搜索功能。
- 构建和优化训练数据流水线:创建标注工具,整理数据集,建立主动学习循环,并确保高质量的标注数据。
- 领导AI语音和音频生成的研发工作,包括语音克隆、文本转语音和音频合成;原型集成并从研究到功能建立可生产使用的路径。
- 研究和原型化图像智能技术,如人脸/身体分析、姿态估计、风格迁移和图像到图像的一致性。
- 开发评估框架,以衡量分类器的准确性、生成模型的质量以及随时间变化的模型漂移。
- 优化推理流水线,提高性能、降低成本和延迟——包括批处理、量化、缓存和模型服务策略。
- 与GPU计算基础设施集成,并通过生产API交付模型。
要求
- 3年以上在生产环境中构建和部署机器学习模型的经验,特别是在分类、标签或内容理解方面。
- 具有模型训练的实际经验,包括数据集整理、架构实验、超参数调优和调试。
- 在图像分类和计算机视觉技术方面有扎实的背景(例如CNN、视觉Transformer、CLIP)。
- 有语音/音频AI经验或表现出浓厚兴趣(例如文本转语音、语音克隆、音频分类)。
- 精通Python,有PyTorch或TensorFlow的经验。
- 有构建数据流水线的经验。
查看英文原文
This is a remote position.
Our client is looking for an innovative and driven AI Engineer to join their team. A leader in media intelligence and AI-driven content creation, they have recently expanded their work in AI voice and image technologies, driving the development of the next generation of cutting-edge products. This role will focus on the creation, classification, and organization of massive volumes of AI-generated media, along with spearheading R&D into AI voice and audio generation and advanced image intelligence capabilities.
Job Description
Responsibilities:
- Design, train, and deploy classification models for content pipeline, including style detection, quality scoring, content moderation, filtering, and semantic categorization of generated media.
- Develop and maintain automated tagging and organization systems for the media library: extracting attributes, detecting visual features, clustering similar content, and enabling intelligent search.
- Build and optimize training data pipelines: create annotation tooling, curate datasets, establish active learning loops, and ensure high-quality labeled data.
- Lead R&D into AI voice and audio generation, including voice cloning, text-to-speech, and audio synthesis; prototype integrations and create a production-ready pathway from research to features.
- Research and prototype image intelligence technologies such as face/body analysis, pose estimation, style transfer, and image-to-image consistency.
- Develop evaluation frameworks to measure the accuracy of classifiers, the quality of generation models, and model drift over time.
- Optimize inference pipelines for performance, cost, and latency—incorporating batching, quantization, caching, and model serving strategies.
- Integrate with GPU compute infrastructure and deliver models via production APIs.
Requirements
- 3+ years of experience building and deploying machine learning models in production, particularly in classification, tagging, or content understanding.
- Hands-on experience with model training, including dataset curation, experimenting with architectures, tuning hyperparameters, and debugging.
- Strong background in image classification and computer vision techniques (e.g., CNNs, vision transformers, CLIP).
- Experience or demonstrated interest in voice/audio AI (e.g., text-to-speech, voice cloning, audio classification).
- Proficiency in Python, with experience in PyTorch or TensorFlow.
- Experience with building data labeling pipelines, annotation workflows, or active learning systems.
- Understanding of model serving in production environments, including REST APIs and latency optimization.
Qualifications:
- Bachelor’s degree or higher in Computer Science, Engineering, or related field.
- Experience in AI/ML, particularly in content classification, tagging, and media organization systems.
- Proven experience with Python and ML frameworks like PyTorch or TensorFlow.
- Strong communication skills to collaborate with R&D teams and integrate new technologies into production.
Benefits
- 100% remote, full-time role.
- Flexible work hours.
- Competitive salary and comprehensive benefits package.
- Opportunity for career advancement and personal development.
- Work on high-impact projects with cutting-edge AI technologies.
Originally posted on Himalayas