DevOps资深平台工程师
DevOps Staff Platform Engineer
DevOps Staff Platform Engineer
地点:远程(加拿大)
薪资:135,000 - 170,000 加元 CAD + 奖金 + 有实际意义的股权
关于Orion Digital
Orion Digital Corp(纳斯达克/多伦多证券交易所:ORIO)是一家公开上市的加拿大金融科技公司,运营一个跨多个引擎的金融平台,涵盖借贷(Mogo)、支付(Carta)和财富管理(Intelligent Investing)。
我们已经超越了传统的金融科技模式。资本分配、决策和风险管理是业务创造价值的核心,AI作为赋能层贯穿整个平台,而不是作为独立的产品叙事。这一层是真实存在的,并且正在使用中,现在的工作是让它更快、更安全、更难被破坏。
为什么需要这个职位
这个职位将构建这一层。
在这里,平台工程不是支持职能。多个业务线、受监管的实体以及上市公司报告义务都运行在这支团队所拥有的基础设施上。每一次部署路径、权限边界和审计追踪,要么让更快的决策成为可能,要么默默限制业务的推进速度。
这意味着平台是战略成功或停滞的关键。如果部署需要一天时间,如果代理无法安全地接触生产环境,如果没人能知道发生了什么以及为什么,赋能层就停留在理论阶段。我们正在招聘一位Staff Platform Engineer来正确地构建它。
我们正在寻找的转变
大多数平台团队为人类开发者构建。而我们的平台则为人类和代理构建,覆盖多个业务线和受监管实体。这改变了工作的本质:
- 黄金路径变成可由机器执行的合同,而不是维基页面。代理必须能够发现它们、调用它们并验证结果。
- 每个概率性步骤都需要一个确定性的门禁。测试、策略检查和模式验证不再只是卫生条件,而是让自主性安全扩展的控制系统。
- 代理是身份。它们需要有限权限的凭证、有限影响范围以及在借贷、支付和证券监管下依然有效的审计追踪。
- 可观测性必须回答一个新问题。不只是“服务是否健康”,而是“代理是否做了正确的事情,我们如何知道”。
如果你已经这样思考这个问题,继续阅读。
你将要做的事情
构建代理工程的基础架构
- 设计并运行代理工作的执行环境:沙盒化、可重现、权限控制、可丢弃。
- 搭建并负责我们的MCP和工具网关层,使代理能够访问基础设施
查看英文原文
DevOps Staff Platform Engineer
Location: Remote (Canada)
Compensation: $135,000 - $170,000 CAD + bonus + meaningful equity
About Orion Digital
Orion Digital Corp (NASDAQ/TSX: ORIO) is a publicly listed Canadian fintech operating a multi-engine financial platform spanning lending (Mogo), payments (Carta), and wealth (Intelligent Investing).
We have moved past the traditional fintech model. Capital allocation, decisioning, and risk management sit at the core of how the business creates value, and AI runs as an enabling layer across the platform rather than as a standalone product narrative. That layer is real and in use, and the work now is making it faster, safer, and harder to break.
Why this role exists
This role builds that layer.
Platform engineering here is not a support function. Multiple business lines, regulated entities, and a public company's reporting obligations all run on infrastructure this team owns. Every deployment path, permission boundary, and audit trail either makes faster decisioning possible or quietly caps how fast the business can move.
Which means the platform is where the strategy succeeds or stalls. If deploying takes a day, if an agent cannot safely touch production, if nobody can tell what changed and why, the enabling layer stays theoretical. We are hiring a Staff Platform Engineer to build it properly.
The shift we are hiring for
Most platform teams build for human developers. Ours builds for humans and agents, across multiple business lines and regulated entities. That changes the job:
- Golden paths become machine-executable contracts, not wiki pages. An agent has to be able to discover them, invoke them, and verify the result.
- Every probabilistic step needs a deterministic gate behind it. Tests, policy checks, and schema validation stop being hygiene and become the control system that lets autonomy scale safely.
- Agents are identities. They need scoped credentials, bounded blast radius, and audit trails that hold up across lending, payments, and securities regulation.
- Observability has to answer a new question. Not just "is the service healthy" but "did the agent do the right thing, and how do we know."
If that is already how you think about the problem, keep reading.
What you will do
Build the substrate for agentic engineering
- Design and run the execution environments agents work in: sandboxed, reproducible, permissioned, disposable.
- Stand up and own our MCP and tool-gateway layer so agents reach infrastructure through governed interfaces instead of ad hoc credentials.
- Turn CI/CD into a validation loop: fast deterministic gates an agent can retry against until it passes, with humans on the loop rather than in it.
Own the platform
- Multi-cloud Kubernetes, infrastructure as code, and GitOps (ArgoCD) delivery. Hands on keyboard, not diagrams.
- Drive infrastructure modernization across our cloud environments, including the unglamorous parts: dependency mapping, and documentation that survives your absence.
- Serve every product line without building a separate platform for each one. Where they genuinely differ, respect it. Where they have drifted for no reason, converge them.
- Treat security and compliance as design constraints from the first commit rather than a review at the end.
Take toil off the company
- Push autonomous remediation into the paths that page people today. The target is alerts that resolve themselves and incidents where the first responder is an agent arriving with a hypothesis already tested.
- Own the numbers: deploy frequency, lead time, change failure rate, MTTR, and the share of operational work running without a human in the middle. Move them quarter over quarter.
- Work outside engineering too. Operations, compliance, and support are full of repetitive work a well-built platform can absorb.
What you need
Non-negotiable
- Deep, current production experience with Kubernetes, Terraform, and a modern GitOps deployment stack. You have owned what you built, including running it in production rather than handing it off once it shipped.
- Real cloud depth. We run across AWS and OCI, so experience in both is an advantage, but what matters is genuine production depth in at least one and the appetite to get fluent in the other quickly.
- You run coding and ops agents as part of your daily work, and you can show us. Sustained real use on real systems (Claude Code, Cursor, Codex, or equivalent, plus whatever you have wired together yourself), not a trial and an opinion. Be ready to walk through something you shipped this month where an agent did most of the typing and you did the judging.
- Strong opinions about where autonomy belongs and where it does not, informed by having been burned at least once.
- Writing clear enough to be read by a junior engineer and a model.
Advantage
- You have built internal developer platforms, MCP servers, or agent tooling that other engineers actually adopted.
- Non-human identity, secrets architecture, or policy-as-code in a regulated environment.
- Re-platforming at scale, especially where the old system could not go down.
- Opinions about how to evaluate infrastructure agents. Almost nobody has these yet.
We are not screening on years of experience as a proxy for capability. We will look at what you have built and how fast you build.
How we work
- Small teams, high ownership, short distance between decision and deployment.
- We would rather collaborate than dictate, and we know when to do it well versus when to do it quickly.
- Output is the measure.
Originally posted on Himalayas