人工智能研究同行评审评估员(机器学习/人工智能)
AI Research Peer Review Evaluator (ML/AI)
Lightly AG 是一家位于苏黎世的 AI 公司,由 ETH/HSG 衍生而来,获得 Y Combinator 和顶级投资机构的支持。我们的机器学习和计算机视觉技术被自动驾驶、医学影像和视觉检测领域的全球领先企业所信赖。
我们正在寻找具有扎实机器学习/AI 背景的研究人员,以支持一个专注于科学同行评审的 AI 评估项目。你将评估由代理型 AI 系统生成的评审,并将其与 ML/AI 研究论文的专家人工同行评审进行对比。
这是一份远程的、基于项目的合同职位,工作时间灵活。
任务
你将从事的工作
- 阅读并分析 ML/AI 研究论文,理解其核心贡献、方法、实验和主张
- 阅读原始的人工同行评审,为每篇论文建立专家基准
- 使用结构化的评分标准,将 AI 生成的同行评审与该基准进行对比
- 评估每份 AI 评审的技术准确性、分析深度、建设性价值以及新颖性/重要性评估
- 识别 AI 评审中的幻觉、无根据的主张、遗漏的技术问题或 AI 评审提出的有价值见解
- 将两份 AI 生成的评审并排比较,判断哪一份提供更强或更有用的分析
- 使用 Google Scholar、arXiv 或 Semantic Scholar 等来源搜索并验证相关学术文献,包括检查引用的先前工作是否在论文提交日期之前已存在
- 提供简洁、有证据支持的解释,说明你的评估决策,并一致地应用项目评分标准
该评估特别关注代理型 AI 评审员是否能提供超越专家人工评审的有意义价值,例如:发现人类评审遗漏的相关前期文献、质疑重要的假设,或通过证据解决不一致之处。
要求
如果你符合以下条件,将是强有力的候选人:
- 拥有机器学习、人工智能、计算机科学、统计学或相关技术领域的硕士、博士学历,或正在攻读研究生课程
- 至少参与过一篇科研论文的撰写,最好是第一作者,但共同作者或其他重要贡献者也欢迎
- 有批判性阅读 ML/AI 研究论文的经验,包括评估方法、实验设计、结果、局限性和科学主张
- 熟悉
查看英文原文
Lightly AG is a Zurich-based AI company and ETH/HSG spin-off, backed by Y Combinator and top-tier investors. Our machine learning and computer vision technology is trusted by global leaders in autonomous driving, medical imaging, and visual inspection.
We’re looking for researchers with strong Machine Learning / AI backgrounds to support an AI evaluation project focused on scientific peer review. You’ll evaluate reviews generated by agentic AI systems and compare them against expert human peer reviews of ML/AI research papers.
This is a remote, project-based contractor opportunity with flexible working hours.
Tasks
What you'll be doing
- Read and scan ML/AI research papers to understand their core contributions, methodology, experiments, and claims
- Review the original human peer reviews to establish an expert baseline for each paper
- Evaluate AI-generated peer reviews against that baseline using a structured scoring rubric
- Assess the technical accuracy, analytical depth, constructive value, and novelty/significance assessment of each AI review
- Identify hallucinations, unsupported claims, missed technical issues, or valuable insights surfaced by the AI reviewers
- Compare two AI-generated reviews side-by-side and determine where one provides stronger or more useful analysis
- Search and verify relevant academic literature using sources such as Google Scholar, arXiv, or Semantic Scholar, including checking whether cited prior work was available before the paper’s submission date
- Provide concise, evidence-based rationales explaining your evaluation decisions and consistently apply the project rubric
The evaluation specifically looks at whether agentic AI reviewers can provide meaningful value beyond expert human reviewers—for example, by identifying relevant prior literature that humans missed, questioning important assumptions, or resolving inconsistencies using evidence.
Requirements
You're a strong candidate if you:
- Have a Master’s, PhD, or are currently pursuing graduate study in Machine Learning, Artificial Intelligence, Computer Science, Statistics, or a closely related technical field
- Have contributed to at least one scientific/research paper, ideally as a first author, although co-authors and other substantial contributors are also welcome
- Have experience critically reading ML/AI research papers, including evaluating methodology, experimental design, results, limitations, and scientific claims
- Are familiar with major ML/AI research venues, such as NeurIPS, ICML, ICLR, ACL, CVPR, or comparable conferences and journals
- Have prior academic peer-review experience, ideally for an ML/AI conference or journal — strongly preferred
- Are comfortable conducting academic literature searches and verifying prior work, publication dates, citations, and novelty claims
- Have strong analytical and written communication skills and can distinguish meaningful technical concerns from superficial criticism
- Can provide clear, concise, evidence-based rationales for your decisions
- Can consistently apply detailed evaluation guidelines and scoring rubrics across multiple papers and reviews
- Have strong attention to detail, particularly when identifying factual inaccuracies or hallucinated technical claims
Benefits
- Fully remote and flexible — work from anywhere
- Part-time contractor role with flexible hours
- Work directly on the evaluation of cutting-edge agentic AI systems for scientific research
- Apply your ML/AI research expertise to help measure and improve the quality of AI-generated scientific peer review
Send us your CV along with a brief note about your research background and areas of expertise. Please include any relevant publications, as well as previous peer-review experience for conferences, journals, workshops, or similar academic venues.
If applicable, we'd also love to know which ML/AI research areas and conferences you’re most familiar with.
We look forward to hearing from you!
Originally posted on Himalayas