高级机器学习运维工程师
Senior Machine Learning Operations Engineer
Mercury在风险决策中使用机器学习的范围和重要性正在快速增长。模型越来越多地驱动关于欺诈和金融犯罪的实时决策,Machine Learning Platform(MLP)团队的职责是构建从训练好的模型到可靠生产部署的顺畅路径,加快迭代速度,并确保细粒度的生产可观测性。
MLP负责生产环境中的机器学习生命周期:从模型注册、部署、实时推理、可观测性到重新训练的系统。我们的数据科学同事负责创建和训练模型。我们构建平台,让他们能够注册、部署和观察这些模型在生产环境中的表现,而无需承担运营负担。我们还为依赖这些模型的决策引擎提供低延迟、高可用性的评分。该平台广泛支持业务决策,我们的第一个用例集中在欺诈风险结果上。
在Mercury,我们致力于为初创公司打造卓越的银行体验。我们的团队专注于确保我们的产品创造一个安全的环境,满足客户、管理员和监管机构的需求。
* Mercury是一家金融科技公司,不是FDIC承保的银行。银行服务由Choice Financial Group和Column N.A.提供,均为FDIC成员。
作为该职位的一部分,您将:
- 构建并维护为风险决策引擎评分模型的实时推理服务,低延迟和高可用性是首要要求
- 负责模型部署基础设施:注册和版本控制、带有性能、偏差和一致性检查的CI/CD、影子模式和分阶段发布
- 构建模型可观测性:可用性、延迟和错误监控,以及作为重新训练触发器的漂移检测
- 与风险数据科学团队合作,从干净的开发到生产交接,再到MLP负责的生产运行
- 实现实验能力,如冠军/挑战者和金丝雀路由,以及SHAP属性等可解释性输出
- 拥有强烈的产品责任感,并积极寻求责任。我们在小型和中型项目上自我组织,我们需要一位热衷于帮助塑造和构建全新平台团队的人
该职位的理想候选人具备:
- 5年以上机器学习工程、后端软件工程、MLOps或相关领域经验
- 生产环境机器学习服务经验:部署、服务
查看英文原文
Mercury's use of machine learning in risk decisioning is growing fast in scope and in stakes. Models increasingly drive real-time decisions about fraud and financial crime, and the Machine Learning Platform (MLP) team exists to build a paved path from a trained model to a reliable production deployment, speeding up iteration, and ensuring granular production observability.
MLP owns the production ML lifecycle: the systems that take a model from registry through deployment, real-time inference, observability, and retraining. Our Data Science colleagues author and train the models. We build the platform that lets them register, deploy, and observe those models in production without carrying the operational burden themselves. We also serve low-latency, highly available scores to the decision engine that depends on them. The platform supports business decisioning broadly, with our first use cases focused on fraud risk outcomes.
At Mercury, we are committed to crafting an exceptional banking* experience for startups. Our team is passionately focused on ensuring our products create a safe environment that meets the needs of our customers, administrators, and regulators.
* Mercury is a fintech company, not an FDIC-insured bank. Banking services provided through Choice Financial Group and Column N.A., Members FDIC.
As part of this role, you will:
- Build and operate the real-time inference service that scores models for the risk decision engine, with low latency and high availability as first-class requirements
- Own model deployment infrastructure: registry and versioning, CI/CD with performance, bias, and consistency checks, shadow mode, and staged rollouts
- Build model observability: availability, latency, and error monitoring, plus drift detection as a retraining trigger
- Partner with Risk Data Science to take models from a clean development-to-production handoff through to production operation under MLP ownership
- Implement experimentation capabilities such as champion/challenger and canary routing, and explainability outputs like SHAP attributions
- Feel a strong sense of product ownership and actively seek responsibility. We self-organize on small and medium projects, and we want someone excited to help shape and build a brand-new platform team
The ideal candidate for the role has:
- 5+ years in machine learning engineering, backend software engineering, MLOps, or a closely related field
- Production ML service experience: deploying, serving, and operating models in low-latency, high-availability contexts
- Strong backend engineering fundamentals in Python, with API frameworks like FastAPI or Flask
- Experience with model deployment and lifecycle tooling: model registries, CI/CD for models, versioning, and staged rollout patterns (shadow, canary, champion/challenger)
- Experience building observability and alerting for production services: latency, errors, and ideally model-specific signals like drift
- Comfort with the data layer ML depends on: SQL, key-value/low-latency stores (Redis, DynamoDB, or equivalent), and streaming pipelines (Kafka, Kinesis, Redpanda, or equivalent)
Nice to have:
- Familiarity with a modern data stack (Snowflake, dbt, Dagster, Airflow, or similar)
- Experience operating in a regulated, audit-sensitive, or compliance-adjacent environment
- Exposure to functional languages or willingness to work across a stack that includes Haskell, React, and TypeScript
Mercury values diversity & belonging and is proud to be an Equal Employment Opportunity employer. All individuals seeking employment at Mercury are considered without regard to race, color, religion, national origin, age, sex, marital status, ancestry, physical or mental disability, veteran status, gender identity, sexual orientation, or any other legally protected characteristic. We are committed to providing reasonable accommodations throughout the recruitment process for applicants with disabilities or special needs. If you need assistance, or an accommodation, please let your recruiter know once you are contacted about a role.
#LI-GC1
Total Rewards
The total rewards package at Mercury includes base salary, equity (stock options/RSUs), and benefits.
Our salary and equity ranges are highly competitive within the SaaS and fintech industry and are updated regularly using the most reliable compensation survey data for our industry. New hire offers are made based on a candidate’s experience, expertise, geographic location, and internal pay equity relative to peers.
Our target new hire base salary ranges for this role are the following:
US employees (any location):
$166,600—$208,300 USD
Canadian employees (any location):
$157,400—$196,800 CAD