站点可靠性工程师 - 低延迟交易系统
Site Reliability Engineer - Low-Latency Trading Systems
关于公司
Hi, 我们是 Ondo Finance。我们的使命是提供机构级、基于区块链的投资产品和服务。我们既有开发去中心化金融技术的技术团队,也有创建和管理代币化基金的资产管理团队。我们在代币化现金储备、代币化股票和ETF领域处于全球领先地位,并正在链上构建机构级金融服务的未来。
由高盛数字资产团队成员创立,我们获得了包括 Founders Fund、Coinbase Ventures、Pantera Capital、Tiger Global 等世界顶级投资者的支持。目前我们在该领域管理资产规模(AUM)位居前列,资金充足,能够持续发展公司。我们是完全远程办公,团队成员遍布美国。
关于职位
Ondo 运行实时交易系统,在传统和加密货币市场全天候运行。该平台包括低延迟的 Rust 引擎、用于交易、执行和损益核算的 Go 服务集群,以及在 AWS 上的多区域 Kubernetes 部署。
我们正在寻找一位具备强大系统编程能力的 SRE 来负责该平台的可靠性、可观测性和性能。这是一项需要亲力亲为的工作:你将阅读和修改 Go 和 Rust 代码,调试延迟问题直至数据源处理模块,市场交易时段进行事件响应,并构建自动化工具,以小团队保障 24/7 交易系统的健康运行。
目标成果
- 负责实时交易服务的生产环境可靠性:交易引擎、执行网关、市场数据摄入和损益/对账流水线
- 运营和演进我们的多区域 AWS(EKS)Kubernetes 集群,通过 GitOps(Flux)部署,使用 SOPS 加密的 secrets
- 构建和优化可观测性:Prometheus 指标和告警、Datadog 日志和仪表盘,以及能在损失发生前检测性能下降的 SLO
- 提升部署安全性:渐进式发布、配置重载行为,以及防止错误推送影响交易的防护机制
职责
- 全流程调试生产环境事件:过时的市场数据源、交易所汇率限制、WebSocket 断开、订单生命周期不同步,以及交易路径中的延迟退化
- 增强来自 Databento 等数据提供商及交易所原生数据源(REST 和 WebSocket)的市场数据摄入,包括过时检测、故障转移和重放功能
- 在实时网关中构建对账和数据完整性工具
查看英文原文
About the company
Hi, we're Ondo Finance. Our mission is to provide institutional-grade, blockchain-enabled investment products and services. We have both a technology arm that develops decentralized finance technology, and an asset management arm that creates and manages tokenized funds. We are the global leader in tokenized treasuries, tokenized stocks and ETFs, and are building the future of institutional-grade financial services onchain.
Founded by folks from Goldman Sachs Digital Assets Team, we’re backed by some of the best investors in the world including Founders Fund, Coinbase Ventures, Pantera Capital, Tiger Global, and more. We are currently the leaders in the space in terms of AUM and are well capitalized to continue growing the firm. We're fully remote, with team members across the U.S.
About the role
Ondo operates real-time trading systems that run around the clock across traditional and crypto venues. The platform spans low-latency Rust engines, a fleet of Go services for trading, execution, and PnL accounting, and a multi-region Kubernetes footprint on AWS.
We are looking for an SRE with strong systems programming skills to own the reliability, observability, and performance of this platform. This is a hands-on role: you will read and modify Go and Rust code, debug latency regressions down to the feed handler, run incident response during market hours, and build the automation that keeps a 24/7 trading system healthy with a small team.
Target outcomes
- Own production reliability for real-time trading services: trading engines, execution gateways, market data ingestion, and PnL/reconciliation pipelines
- Operate and evolve our multi-region Kubernetes clusters on AWS (EKS), deployed via GitOps (Flux) with SOPS-encrypted secrets
- Build and refine observability: Prometheus metrics and alerting, Datadog logs and dashboards, and the SLOs that catch degradation before it costs money
- Improve deploy safety: progressive rollouts, config-reload behavior, and guardrails that prevent a bad push from touching live trading
Responsibilities
- Debug production incidents end to end: stale market data feeds, exchange rate limits, WebSocket disconnects, order-lifecycle desyncs, and latency regressions in the trading path
- Harden market data ingestion from providers such as Databento and venue-native feeds (REST and WebSocket), including staleness detection, failover, and replay
- Build reconciliation and data-integrity tooling across live gauges, Postgres, and our S3 parquet data lake, so positions, fills, and PnL always agree
- Participate in an on-call rotation covering US equity market hours and 24/7 crypto venues
Requirements
- 5+ years in SRE, production engineering, or infrastructure roles, with meaningful time supporting real-time or latency-sensitive systems
- Strong programming ability in Go or Rust, and willingness to work in both; this role changes application code, not just infrastructure
- Deep, hands-on Kubernetes and AWS experience: you have run stateful, latency-sensitive workloads in production, not just stateless web services
- Strong observability instincts: fluent PromQL, structured-log analysis, and experience designing alerts with high signal and low noise
- Solid Linux internals and networking fundamentals: you can chase a p99 regression through the kernel, the NIC, or the GC
- Sound judgment under pressure and clear written communication during and after incidents
Nice to haves
- Experience operating trading systems, execution infrastructure, or market data infrastructure at a trading firm, exchange, or broker
- Familiarity with market microstructure and order lifecycle (order books, order types, fills and reconciliation)
- Experience with market data providers and protocols (Databento, SIP/prop equity feeds, venue WebSocket APIs)
- Exposure to crypto venues and on-chain trading
- Python for operational tooling and data analysis (pandas, parquet, BigQuery)
- Experience with GitOps workflows, infrastructure as code, and secrets management at scale
Tech stack
Go, Rust, Python | Kubernetes (EKS), Flux, SOPS | AWS (multi-region), S3 parquet lake | Prometheus, Grafana, Datadog | Postgres, CockroachDB, BigQuery | Databento, venue WebSocket/REST feeds
What we offer
- Competitive compensation including but not limited to salary, future token rights, and/or equity (according to your preferences) — We are well-funded and believe that great talent deserves great compensation.
- Full benefits (medical, vision, and dental) and flexible vacation policy (PTO).
- Remote-first team across many countries — You will be an early team member helping shape our vision, culture, and design practices.
- A+ colleagues — Our team includes alumni from: Goldman Sachs, Blackrock, Two Sigma, Bridgewater, SpaceX, AWS, Meta, Google, McKinsey, Coinbase, Circle, Uniswap.
- Best-in-class investors — We are proud to be backed by leading crypto experts and VCs, including Pantera Capital, Founders Fund and Coinbase Ventures.
Originally posted on Himalayas