高级评估算法工程师
Senior Evaluation Algorithm Engineer
Tags: Web3 职位 • 区块链机器学习职位 • Web3 Python 职位 • Web3 全职职位 • Web3 职位
Binance 是一个领先的全球区块链生态系统,是交易量和注册用户最多的加密货币交易所。我们受到 100 多个国家 3 亿多用户的信任,因为我们具有行业领先的安全性、用户资金透明度、交易引擎速度、深度流动性以及无与伦比的数字资产产品组合。Binance 的服务范围从交易和金融到教育、研究、支付、机构服务、Web3 功能等。我们利用数字资产和区块链的力量,构建一个包容性的金融生态系统,以推进货币自由并改善全球人们的金融可及性。
职位简介
在人工智能时代,大型语言模型正在重塑对话和交易等核心业务场景。模型能力的迭代依赖于科学且可信的评估体系——衡量模型质量并指导研发方向的“标尺”。我们正在寻找一位具有算法背景的评估专家,构建覆盖对话、金融交易等场景的 LLM 评估能力,使用专业的评估方法量化模型性能,定位问题,并推动模型持续改进。
职责
为对话和金融交易等业务场景设计端到端的 LLM 评估方案。构建评估指标体系和评分标准,将主观的模型性能判断转化为可量化、可复现、可解释的评估结论。
主导评估数据集的设计和构建。定义评估维度和场景覆盖范围,建立高质量的数据标注指南和质量控制流程,构建真实反映业务需求且具备区分力的基准。
基于评估结果分析模型的能力边界和失败模式。产出可操作的改进建议,并与算法和产品团队协作推动模型迭代,使评估成为研发循环中的关键环节。
推动评估工作流的自动化和扩展。构建可持续的评估平台和工具链,支持快速模型迭代过程中的高频、稳定评估需求。
与算法、产品和数据团队协作,将业务和模型目标转化为清晰的评估标准,并将评估结果转化为实际的优化行动。
查看英文原文
Tags: Web3 Jobs • Cryptocurrency Machine Learning Jobs • Cryptocurrency Python Jobs • Cryptocurrency Full Time Jobs • Web3 Web3 JobsBinance is a leading global blockchain ecosystem behind the world’s largest cryptocurrency exchange by trading volume and registered users. We are trusted by 300+ million people in 100+ countries for our industry-leading security, user fund transparency, trading engine speed, deep liquidity, and an unmatched portfolio of digital-asset products. Binance offerings range from trading and finance to education, research, payments, institutional services, Web3 features, and more. We leverage the power of digital assets and blockchain to build an inclusive financial ecosystem to advance the freedom of money and improve financial access for people around the world.About the RoleIn the AI era, large language models are reshaping core business scenarios such as dialogue and trading. Model capability iteration relies on a scientific and trustworthy evaluation system — the "ruler" that measures model quality and guides R&D direction. We are seeking an evaluation expert with an algorithmic background to build LLM evaluation capabilities covering dialogue, financial trading, and other scenarios, using professional evaluation methods to quantify model performance, pinpoint issues, and drive continuous model improvement.ResponsibilitiesDesign end-to-end LLM evaluation plans for business scenarios such as dialogue and financial trading. Build evaluation metric systems and rubrics, transforming subjective model performance judgments into quantifiable, reproducible, and explainable evaluation conclusions.Lead the design and construction of evaluation datasets. Define evaluation dimensions and scenario coverage, establish high-quality data annotation guidelines and quality control processes, and build benchmarks that authentically reflect business needs and have discriminative power.Analyze model capability boundaries and failure modes based on evaluation results. Produce actionable improvement recommendations and collaborate with algorithm and product teams to drive model iteration, making evaluation a critical component of the R&D loop.Drive the automation and scaling of evaluation workflows. Build sustainable evaluation platforms and toolchains to support high-frequency, stable evaluation needs during rapid model iteration.Collaborate with algorithm, product, and data teams to translate business and model objectives into clear evaluation standards, and turn evaluation findings into concrete R&D directions and drive their implementation.RequirementsMaster's degree or above in Computer Science, Artificial Intelligence, Mathematics, Statistics, or related fields, with a solid algorithmic foundation and understanding of LLM principles, training, and fine-tuning processes.Hands-on LLM evaluation experience at a large tech company, with participation in commercial deployment evaluation (not purely academic or offline benchmarking). Familiar with the full pipeline from evaluation data preparation and rubrics design to evaluation-driven R&D.Familiar with mainstream evaluation methods (human evaluation, model-based automatic evaluation / LLM-as-a-judge, metric computation) and their applicable boundaries. Able to define appropriate evaluation dimensions for different business scenarios and write clear, actionable, and discriminative rubrics.Systematic control over evaluation data representativeness, annotation consistency, and result reliability, ensuring scientific and trustworthy evaluation conclusions.Proficient in Python, with experience in evaluation workflow automation, benchmark construction, or evaluation platform development. Able to independently handle data processing, evaluation script writing, and result analysis.Strong business understanding and communication skills, able to translate evaluation findings into clear improvement directions and effectively drive cross-team collaboration.Bonus QualificationsExperience evaluating dialogue systems, AI Agents, or financial/trading LLMs.Experience building high-quality AI training/evaluation data or data annotation systems.Familiarity with RLHF, reward models, or preference data-related work.Why Binance• Shape the future with the world’s leading blockchain ecosystem• Collaborate with world-class talent in a user-centric global organization with a flat structure• Tackle unique, fast-paced projects with autonomy in an innovative environment• Thrive in a results-driven workplace with opportunities for career growth and continuous learning• Competitive salary and company benefits• Work-from-home arrangement (the arrangement may vary depending on the work nature of the business team)Binance is committed to being an equal opportunity employer. We believe that having a diverse workforce is fundamental to our success.By submitting a job application, you confirm that you have read and agree to our Candidate Privacy Notice.We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.Apply here 👉 Senior Evaluation Algorithm Engineer