高级软件工程师 - Mercury Command
Senior Software Engineer - Mercury Command
1965年,一名苏格兰工程师被赋予了一个看似普通但棘手的任务——银行希望在周六营业并服务客户,但不知道如何解决正确用户的认证问题。詹姆斯·古德费洛(James Goodfellow)在史密斯工业公司工作时发现,你需要拥有某样东西(一张卡)和知道某样东西(一个PIN码)。因为有人告诉他他们只能记住4位数字而不是他提出的6位,今天超过300万台ATM机仍在使用卡+4位PIN码,这一方式已经持续了60多年。
就像ATM一样,Mercury正在构建长期推进金融界面的技术。Command是Mercury推出的基于大语言模型的财务助手,于2026年6月向所有客户推出。它让用户用自然语言了解自己的财务状况并采取行动,从询问现金流到发送付款、发行卡片和管理发票。随着产品现在已交付给客户,我们的工作是不断进化它、扩展其功能,并寻找新的方法利用大语言模型,为Mercury客户提供更强大的银行体验。
你将要做的
交付用户喜爱的新功能:
- 设计并交付新的Command技能,即用于指导模型处理如转账、管理发票和理解现金流等流程的领域特定指令集
- 在Command中设计并构建代理流程,定义多步骤代理交互的架构,以便我们为客户代表产品执行更多操作
- 与后端团队合作,定义新功能的工具模式,塑造Mercury业务逻辑与模型之间的数据契约
- 从系统提示到渲染响应的前端组件,全程负责新功能的开发
负责大语言模型层:
- 维护并演进Command的提示架构:系统提示、技能加载系统、会话上下文以及底层的策略和合规层
- 调整模型行为:推理努力程度、提示缓存策略、回退链以及使产品感觉快速的流式传输模式
- 跟踪模型的演变情况,并将这些知识带回Command的构建中
构建质量:
- 编写并扩展Command的评估框架,添加覆盖新功能的案例和能提前检测回归的评分标准
- 与产品和合规团队合作,定义每项新功能“正确工作”的含义,然后构建证明这一点的测试
- 全程负责
查看英文原文
In 1965, an engineer in Scotland was given a mundane task with a tricky problem to it - banks wanted to close on Saturdays and still serve customers, but they didn’t know how to solve the authentication of the right user. James Goodfellow, working at Smiths Industries, uncovered the insights that you needed something you have (a card) and something you know (a PIN). Because someone told him they could only remember 4 digits instead of his proposed 6, today over 3 million ATMs use a card+4 digit PIN over 60 years later.
Like the ATM, Mercury is building technology that pushes forward financial interfaces for the long term. Command is Mercury's LLM-powered financial assistant, launched to all customers in June 2026. It lets users understand their finances and take action in plain language, from asking about cash flow to sending payments, issuing cards, and managing invoices. With the product now in customers' hands, the work is to evolve it, extend its capabilities, and find new ways to leverage LLMs to give Mercury customers a more powerful banking* experience.
What you'll do
Ship new capabilities users love:
- Design and ship new Command skills, the domain-specific instruction sets that teach the model how to handle workflows like sending money, managing invoices, and understanding cash flow
- Design and build agentic workflows in Command, defining the architecture for how multi-step agent interactions should work as we extend what the product can do on a customer's behalf
- Work with backend teams to define tool schemas for new capabilities, shaping the data contracts between Mercury's business logic and the model
- Own new capabilities end to end, from the system prompt to the frontend component that renders the response
Own the LLM layer:
- Maintain and evolve Command's prompt architecture: the system prompt, skill loading system, session context, and the policy and compliance layers underneath
- Tune model behavior: reasoning effort, prompt caching strategy, fallback chains, and the streaming patterns that make the product feel fast
- Stay current with how models are evolving and bring that knowledge back to how Command is built
Build quality in:
- Write and expand Command's eval harness, adding cases that cover new capabilities and scoring rubrics that detect regressions before users do
- Partner with product and compliance teams to define what "working correctly" means for each new capability, then build the tests that prove it
- Own the reliability and quality of what you ship, from initial design through post-launch monitoring
This list is illustrative. Command is a product in motion and priorities will shift as we learn. The right person will help choose the next highest-leverage work.
The ideal candidate
- Has 7 or more years of software engineering experience, with deep technical expertise building and scaling LLM-powered applications in production
- Has gone beyond shipping a first version: you have scaled an LLM-powered product, dealt with the reliability and performance problems that come with real usage, and made it better over time
- Has experience designing agentic systems and has opinions about how to architect multi-step workflows that are reliable, explainable, and safe to run on behalf of real users
- Has built eval infrastructure and can write cases that actually measure whether the product works, not just whether the model outputs something plausible
- Understands the real tradeoffs in LLM deployments: latency, cost, compliance, and what breaks in production that doesn't show up in demos
- Has opinions about what makes an AI product trustworthy, not just impressive, and can build toward that bar
- Is comfortable with TypeScript and willing to learn Haskell for backend tool work, or already comfortable with both
- Can work across the full stack of an AI product, from the system prompt to the streaming frontend
- Has a track record of mentoring engineers and raising the technical bar of their team
If this role interests you, we invite you to explore our public demo at demo.mercury.com.
*Mercury is a fintech company, not an FDIC-insured bank. Banking services provided through Choice Financial Group and Column N.A., Members FDIC.
Mercury values diversity & belonging and is proud to be an Equal Employment Opportunity employer. All individuals seeking employment at Mercury are considered without regard to race, color, religion, national origin, age, sex, marital status, ancestry, physical or mental disability, veteran status, gender identity, sexual orientation, or any other legally protected characteristic. We are committed to providing reasonable accommodations throughout the recruitment process for applicants with disabilities or special needs. If you need assistance, or an accommodation, please let your recruiter know once you are contacted about a role.
#LI-RA1
Total Rewards
The total rewards package at Mercury includes base salary, equity (stock options/RSUs), and benefits.
Our salary and equity ranges are highly competitive within the SaaS and fintech industry and are updated regularly using the most reliable compensation survey data for our industry. New hire offers are made based on a candidate’s experience, expertise, geographic location, and internal pay equity relative to peers.
Our target new hire base salary ranges for this role are the following:
US employees (any location):
$200,700—$250,900 USD
Canadian employees (any location):
$189,700—$237,100 CAD