远程工作雷达

AI工程师-分类器、媒体智能与语音研发

AI Engineer - Classifiers, Media Intelligence & Voice R&D

AI开发工程限定地区(需当地身份)
公司WOW Remote Teams
薪资未公开
工作地点United States
地域资格限定地区(需当地身份)
时区要求日间重叠约 9 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 United States 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

这是一个远程职位。
我们的客户正在寻找一位有创新精神且充满动力的AI工程师加入他们的团队。他们是一家媒体智能和AI驱动内容创作的领先企业,最近在AI语音和图像技术领域进行了扩展,推动下一代尖端产品的开发。该职位将专注于大规模AI生成媒体的创建、分类和组织,同时领导AI语音和音频生成以及高级图像智能能力的研发。

职位描述

职责:

  • 设计、训练和部署内容流水线的分类模型,包括风格检测、质量评分、内容审核、过滤和生成媒体的语义分类。
  • 开发和维护媒体库的自动化标签和组织系统:提取属性、检测视觉特征、聚类相似内容,并实现智能搜索功能。
  • 构建和优化训练数据流水线:创建标注工具,整理数据集,建立主动学习循环,并确保高质量的标注数据。
  • 领导AI语音和音频生成的研发工作,包括语音克隆、文本转语音和音频合成;原型集成并从研究到功能建立可生产使用的路径。
  • 研究和原型化图像智能技术,如人脸/身体分析、姿态估计、风格迁移和图像到图像的一致性。
  • 开发评估框架,以衡量分类器的准确性、生成模型的质量以及随时间变化的模型漂移。
  • 优化推理流水线,提高性能、降低成本和延迟——包括批处理、量化、缓存和模型服务策略。
  • 与GPU计算基础设施集成,并通过生产API交付模型。

要求

  • 3年以上在生产环境中构建和部署机器学习模型的经验,特别是在分类、标签或内容理解方面。
  • 具有模型训练的实际经验,包括数据集整理、架构实验、超参数调优和调试。
  • 在图像分类和计算机视觉技术方面有扎实的背景(例如CNN、视觉Transformer、CLIP)。
  • 有语音/音频AI经验或表现出浓厚兴趣(例如文本转语音、语音克隆、音频分类)。
  • 精通Python,有PyTorch或TensorFlow的经验。
  • 有构建数据流水线的经验。
查看英文原文

This is a remote position.
Our client is looking for an innovative and driven AI Engineer to join their team. A leader in media intelligence and AI-driven content creation, they have recently expanded their work in AI voice and image technologies, driving the development of the next generation of cutting-edge products. This role will focus on the creation, classification, and organization of massive volumes of AI-generated media, along with spearheading R&D into AI voice and audio generation and advanced image intelligence capabilities.
Job Description

Responsibilities:

  • Design, train, and deploy classification models for content pipeline, including style detection, quality scoring, content moderation, filtering, and semantic categorization of generated media.
  • Develop and maintain automated tagging and organization systems for the media library: extracting attributes, detecting visual features, clustering similar content, and enabling intelligent search.
  • Build and optimize training data pipelines: create annotation tooling, curate datasets, establish active learning loops, and ensure high-quality labeled data.
  • Lead R&D into AI voice and audio generation, including voice cloning, text-to-speech, and audio synthesis; prototype integrations and create a production-ready pathway from research to features.
  • Research and prototype image intelligence technologies such as face/body analysis, pose estimation, style transfer, and image-to-image consistency.
  • Develop evaluation frameworks to measure the accuracy of classifiers, the quality of generation models, and model drift over time.
  • Optimize inference pipelines for performance, cost, and latency—incorporating batching, quantization, caching, and model serving strategies.
  • Integrate with GPU compute infrastructure and deliver models via production APIs.

Requirements

  • 3+ years of experience building and deploying machine learning models in production, particularly in classification, tagging, or content understanding.
  • Hands-on experience with model training, including dataset curation, experimenting with architectures, tuning hyperparameters, and debugging.
  • Strong background in image classification and computer vision techniques (e.g., CNNs, vision transformers, CLIP).
  • Experience or demonstrated interest in voice/audio AI (e.g., text-to-speech, voice cloning, audio classification).
  • Proficiency in Python, with experience in PyTorch or TensorFlow.
  • Experience with building data labeling pipelines, annotation workflows, or active learning systems.
  • Understanding of model serving in production environments, including REST APIs and latency optimization.

Qualifications:

  • Bachelor’s degree or higher in Computer Science, Engineering, or related field.
  • Experience in AI/ML, particularly in content classification, tagging, and media organization systems.
  • Proven experience with Python and ML frameworks like PyTorch or TensorFlow.
  • Strong communication skills to collaborate with R&D teams and integrate new technologies into production.

Benefits

  • 100% remote, full-time role.
  • Flexible work hours.
  • Competitive salary and comprehensive benefits package.
  • Opportunity for career advancement and personal development.
  • Work on high-impact projects with cutting-edge AI technologies.

Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

SEO专员

WOW Remote TeamsUnited StatesFull Time今天
市场运营限定地区(需当地身份)

会计

WOW Remote TeamsUnited StatesPart Time2 天前
职能支持限定地区(需当地身份)

初级估算员

WOW Remote TeamsUnited StatesFull Time3 天前
其他限定地区(需当地身份)

← 返回全部职位