技术成员,研究工程
Member of Technical Staff, Research Engineering
职位名称:技术成员,研究工程
职位类型:全职
工作地点:远程
职位描述
我们正在寻找一名研究工程师,在强化学习(RL)的前沿领域工作,开发新的环境、训练流程和评估系统,以提升现代人工智能模型的能力。该职位位于研究与生产之间,将实验性想法转化为可扩展、高性能的系统。
你将参与的工作
- 设计自包含的强化学习环境,以捕捉复杂的真实任务,包括奖励函数、验证器和评估逻辑。
- 设计并扩展剧集流程和多组件训练流程(MCP),以支持可重复的实验。
- 构建自动化数据生成系统,利用合成数据加速训练周期,同时不牺牲质量。
- 开发并集成基于人工智能的评估和质量保障系统,用于自动化评分、验证和反馈循环。
- 使用内部生成的数据集和自定义训练策略对开源强化学习模型进行微调和优化。
- 建立基准框架,以衡量模型在不同任务中的能力、鲁棒性和数据质量。
- 为内部和外部基准平台(例如 micro1 基准)的评估发布和分析做出贡献。
我们寻找的人选
- 在强化学习方面有深入经验,包括环境设计和训练动态。
- 有构建和扩展强化学习系统、流程或实验框架的出色记录。
- 精通自动化和数据生成,包括合成数据流程。
- 熟悉自动化评估系统、模型验证和质量保障流程。
- 有微调和评估开源机器学习模型的经验。
- 沟通清晰,技术写作能力强。
- 能够在快节奏、以研究为导向且高度协作的环境中工作。
优先考虑
- 有发布基准、评估或研究成果的经验。
- 熟悉评估生态系统(例如 micro1 基准或其他类似框架)。
- 有大规模强化学习实验的可扩展基础设施背景。
薪酬与福利说明
该全职职位的全国薪资范围为 180,000 至 260,000 美元。所有员工均有资格获得股权补偿,员工还可能根据绩效获得奖金,具体取决于公司政策。
查看英文原文
Job Title: Member of Technical Staff, Research Engineering
Job Type: Full-time
Location: Remote
The Role
We are seeking a Research Engineer to operate at the frontier of Reinforcement Learning (RL), developing novel environments, training pipelines, and evaluation systems that advance the capabilities of modern AI models. This role sits at the intersection of research and production, translating experimental ideas into scalable, high-performance systems.
What You’ll Work On
- Architect self-contained RL environments that capture complex, real-world tasks, including reward functions, verifiers, and evaluation logic.
- Design and scale episode pipelines and multi-component training processes (MCPs) to support reproducible experimentation.
- Build automated data generation systems, leveraging synthetic data to accelerate training cycles without compromising quality.
- Develop and integrate AI-driven evaluation and quality assurance systems for automated grading, validation, and feedback loops.
- Fine-tune and optimize open-source RL models using internally generated datasets and custom training strategies.
- Establish benchmarking frameworks to measure model capability, robustness, and data quality across tasks.
- Contribute to the release and analysis of evaluations on internal and external benchmark platforms (e.g., micro1 benchmarks).
What We're Looking For
- Deep experience in Reinforcement Learning, including environment design and training dynamics.
- Strong track record of building and scaling RL systems, pipelines, or experimentation frameworks.
- Proficient in automation and data generation, including synthetic data pipelines.
- Familiar with automated evaluation systems, model validation, and quality assurance workflows.
- Experienced in fine-tuning and evaluating open-source ML models.
- Clear, concise communicator with strong technical writing skills.
- Comfortable operating in fast-paced, research-driven, and highly collaborative environments.
Preferred
- Experience publishing benchmarks, evaluations, or research artifacts.
- Familiarity with evaluation ecosystems (e.g., micro1 benchmarks or similar frameworks).
- Background in scalable infrastructure for large-scale RL experimentation.
Compensation & Benefits Notice
The national pay range for this full-time position is base salary of $180,000 –$260,000 USD. All employees are eligible for equity compensation, and employees may also receive performance-based bonuses, dependent on role and subject to company policies. micro1 provides a comprehensive benefits package, including up to 100% reimbursement for health-insurance premiums, paid time off, a 401(K) plan with a company match, and additional benefits designed to support a high-performing, remote-first workforce.
micro1 is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, or gender identity), national origin, age, disability, genetic information, veteran status, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance and/or a reasonable accommodation during the application process, reach out to .
Our hiring process utilizes artificial intelligence tools to assist in candidate screening and assessment. Our AI tools are designed to complement, not replace, human decision-making.
Disclaimer
The information contained in this job posting, including but not limited to role responsibilities, qualifications, compensation, and benefits, is provided for informational purposes only and does not constitute a binding offer of employment. micro1 reserves the right to amend, modify, or withdraw any portion of this posting at its sole discretion and without prior notice. All employment decisions are made in accordance with applicable laws and regulations.
Originally posted on Himalayas