技术职员,编码研究
Member of Technical Staff, Coding Research
职位名称:技术员工,编码研究方向
职位类型:全职
工作地点:远程
职位描述
我们正在寻找一名技术员工,帮助推进前沿编码代理的评估和开发。你将处于人工智能研究、软件工程和模型评估的交汇点,设计用于衡量和改进下一代编码模型的基准、方法和数据系统。
你将负责的工作
- 设计并主导编码代理的评估框架,包括基准规范、评分方法、评分标准和质量标准。
- 领导端到端的研究项目,专注于在各种软件工程任务中衡量和提升编码模型的性能。
- 开发高质量的数据集、黄金示例和评估协议,以实现对前沿编码系统的可靠评估。
- 分析模型行为和失败模式,识别系统性弱点,并将发现转化为训练和评估的改进建议。
- 构建工具和基础设施,支持大规模实验、数据生成、评审工作流和评估流程。
- 建立编码代理评估的最佳实践,确保方法严谨性、可复现性和测量质量。
- 与研究人员、工程师和应用AI团队紧密合作,设计实验并评估新兴模型能力。
- 为技术报告、基准研究和面向客户的科研项目做出贡献,传达模型性能和洞察。
我们寻找的人选
- 具备扎实的软件工程背景,精通Python、C++或类似编程语言。
- 3年以上软件工程、机器学习、AI研究、评估或相关技术领域的经验。
- 有设计、审查或验证技术评估、基准测试、编码任务或评估方法的经验。
- 熟悉大语言模型、编码代理、强化学习、模型评估或相关AI系统。
- 有构建工具、自动化工作流并通过系统化实验改进技术流程的实证能力。
- 强大的分析能力,能够调查模型行为并从复杂的系统中得出见解。
- 出色的书面和口头沟通能力,能够向不同受众清晰地阐述技术发现。
- 中国国籍
查看英文原文
Job Title: Member of Technical Staff, Coding Research
Job Type: Full-time
Location: Remote
The Role
We are seeking a Member of Technical Staff to help advance the evaluation and development of frontier coding agents. Sitting at the intersection of AI research, software engineering, and model evaluation, you will design the benchmarks, methodologies, and data systems that shape how next-generation coding models are measured and improved.
What You'll Do
- Design and own evaluation frameworks for coding agents, including benchmark specifications, scoring methodologies, rubrics, and quality standards.
- Lead end-to-end research initiatives focused on measuring and improving coding model performance across diverse software engineering tasks.
- Develop high-quality datasets, golden examples, and evaluation protocols that enable reliable assessment of frontier coding systems.
- Analyze model behavior and failure modes, identifying systematic weaknesses and translating findings into actionable improvements for training and evaluation.
- Build tooling and infrastructure that support large-scale experimentation, data generation, review workflows, and evaluation pipelines.
- Establish best practices for coding-agent assessment, ensuring methodological rigor, reproducibility, and measurement quality.
- Partner closely with researchers, engineers, and applied AI teams to design experiments and evaluate emerging model capabilities.
- Contribute to technical reports, benchmark studies, and client-facing research initiatives that communicate model performance and insights.
What We're Looking For
- Strong software engineering background with expertise in Python, C++, or comparable programming languages.
- 3+ years of experience in software engineering, machine learning, AI research, evaluation, or related technical disciplines.
- Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies.
- Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems.
- Proven ability to build tooling, automate workflows, and improve technical processes through systematic experimentation.
- Strong analytical skills with the ability to investigate model behavior and derive insights from complex technical systems.
- Excellent written and verbal communication skills, including the ability to clearly articulate technical findings to diverse audiences.
- Comfortable operating in fast-moving research environments with significant ambiguity and evolving priorities.
Preferred
- Experience working on frontier AI systems, coding agents, or model evaluation research.
- Deep interest in understanding how data, evaluations, and feedback mechanisms influence model capabilities.
- Track record of independently driving ambiguous technical or research projects from conception to execution.
- Experience designing benchmarks or datasets for machine learning systems at scale.
- Familiarity with agentic workflows, tool use, reinforcement learning, or post-training methodologies.
- Publications, open-source contributions, or demonstrated technical leadership in AI, machine learning, or software engineering.
Compensation & Benefits Notice
The national pay range for this full-time position is base salary of $200,000 –$260,000 USD. All employees are eligible for equity compensation, and employees may also receive performance-based bonuses, dependent on role and subject to company policies. micro1 provides a comprehensive benefits package, including up to 100% reimbursement for health-insurance premiums, paid time off, a 401(K) plan with a company match, and additional benefits designed to support a high-performing, remote-first workforce.
micro1 is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, or gender identity), national origin, age, disability, genetic information, veteran status, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance and/or a reasonable accommodation during the application process, reach out to .
Our hiring process utilizes artificial intelligence tools to assist in candidate screening and assessment. Our AI tools are designed to complement, not replace, human decision-making.
Originally posted on Himalayas