高级DevOps工程师
Senior DevOps Engineer
我们是谁:
Alpaca 是一家总部位于美国的全球领先代理优先经纪基础设施公司,提供股票、ETF、期权、加密货币、固定收益、24/5 小时交易等服务。
在我们的子公司中,Alpaca 是一家持牌金融服务公司,通过我们机构级 API 为全球 40 个国家的数百家金融机构提供服务。这包括经纪自营商、投资顾问、财富管理公司、对冲基金和加密货币交易所,总计超过 1000 万笔经纪账户。
我们的全球团队是一支由经验丰富的工程师、交易员和经纪专业人士组成的多元化团队,致力于实现我们让全球每个人都能获得金融服务的使命。我们高度重视开源贡献,积极培育活跃的社区,持续提升我们获奖的、开发者友好的 API 以及其背后的强大基础设施。
Alpaca 获得了来自顶级全球投资者的 4 亿美元资金支持,包括 Portage Ventures、Spark Capital、Tribe Capital、Social Leverage、Horizons Ventures、Opera Tech Ventures、SBI Group、Derayah Financial、Unbound、Peak XV、Elefund 和 Y Combinator。
我们的团队成员:
我们是一个由 400 多名分布在世界各地的成员组成的充满活力的团队,大家喜欢在世界上的最爱地点工作,团队成员遍布美国、加拿大、日本、匈牙利、尼日利亚、巴西、英国等地!
我们正在寻找渴望为 Alpaca 快速发展做出贡献的热情人士。如果你认同我们的核心价值观——保持好奇、富有同理心、承担责任,并准备好产生重大影响,我们鼓励你申请。
职位
作为高级 DevOps 工程师,你将设计、构建和运营使 Alpaca 能够在全球范围内扩展并自信运行交易关键系统的基础设施。你将有自主权针对明确的目标设计和实施解决方案,并与团队一起对这些目标有真正的发言权。
我们不招聘任何单一工具的专家。我们寻找的是具备云架构和基础设施即代码思维的全面基础设施工程师,具有真实的平台即产品理念:一个以其他工程团队能够快速安全地交付为成功标准的人,将手动劳动视为需要被工程解决的缺陷。你能够在操作层面管理我们的数据存储(PostgreSQL、消息代理),并与我们的 SRE 和数据库专家合作完成更深入的工作。
你将有机会
查看英文原文
Who We Are:
Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more.
Amongst our subsidiaries, Alpaca is a licensed financial services company, serving hundreds of financial institutions across 40 countries with our institutional-grade APIs. This includes broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, totalling over 10 million brokerage accounts.
Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to achieve our mission of opening financial services to everyone on the planet. We're deeply committed to open-source contributions and fostering a vibrant community, continuously enhancing our award-winning, developer-friendly API and the robust infrastructure behind it.
Alpaca is proudly backed by $400 million in funding from top-tier global investors including Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, Horizons Ventures, Opera Tech Ventures, SBI Group, Derayah Financial, Unbound, Peak XV, Elefund, and Y Combinator.
Our Team Members:
We're a dynamic team of 400+ globally distributed members who thrive working from our favorite places around the world, with teammates spanning the USA, Canada, Japan, Hungary, Nigeria, Brazil, the UK, and beyond!
We're searching for passionate individuals eager to contribute to Alpaca's rapid growth. If you align with our core values—Stay Curious, Have Empathy, and Be Accountable—and are ready to make a significant impact, we encourage you to apply.
Role
As a Senior DevOps Engineer you will design, build and operate the infrastructure that lets Alpaca scale globally and run trading-critical systems with confidence. You will have the autonomy to design and implement solutions against clearly defined goals - and a real voice in shaping those goals with the team.
We are not hiring a specialist in any single tool. We are looking for a well-rounded infrastructure engineer who thinks in cloud architecture and Infrastructure-as-Code, with a genuine Platform-as-a-Product mindset: someone who measures success by how quickly and safely the rest of engineering can ship, and who treats manual toil as a bug to be engineered away. You are comfortable operating our data stores (PostgreSQL, Message Brokers) at an operator level, partnering with our SRE and database specialists on the deeper work.
Things You Get To Do
- Design and evolve our cloud architecture on GCP - networking, interconnects, IAM and high-availability topology - and express it entirely as code with Terraform, following GitOps as a first principle.
- Build and own the CI/CD pipelines that plan, review, test and safely apply IaC changes - Policy-as-Code guardrails, drift detection and progressive rollout so infrastructure changes ship as confidently as application code.
- Advance Platform-as-a-Product: build self-serve capabilities and paved paths so engineers can provision what they need, through a golden path rather than a hand-off.
- Strengthen our observability stack - metrics, logs, traces and alerting across Prometheus, Thanos, Grafana, Loki, Tempo and Alertmanager - so the platform is easy to run and reason about.
- Operate our GKE clusters and the infrastructure services that run on them - Helm-packaged workloads, message brokers (RabbitMQ, IBM MQ) and data stores.
- Participate in our Follow-The-Sun on-call model: watch and triage alerts, join and declare incidents, lead structured debugging and escalation, and drive blameless post-mortems and the post-actions that actually close the loop.
- Embed SRE practices - SLIs/SLOs and error budgets, capacity planning - into how Core Infrastructure builds and operates, working closely with our SRE function.
Who You Are (Must-Haves)
- 5+ years in a DevOps, Platform/Infrastructure, or SRE role, with a proven track record operating large-scale, high-availability, high-performance systems in production.
- Deep hands-on experience designing cloud architecture on Google Cloud Platform (GCP) as the primary cloud - landing zones, networking, IAM and high-availability topology.
- Strong Infrastructure-as-Code skills with Terraform, structuring large codebases across multiple environments, with GitOps as a first principle and least-privilege as a default mindset.
- Proven experience building CI/CD pipelines for IaC - automated plan/apply, code review, Policy-as-Code, drift detection and safe rollout.
- Significant production experience with Kubernetes (ideally GKE) and packaging/deploying workloads with Helm.
- Solid cloud and L3/L4-L7 networking fundamentals (VPCs, routing, load balancing, DNS, TLS, interconnects) and comfort debugging cross-service connectivity.
- Hands-on experience with a modern observability stack - Prometheus, Thanos, Grafana, Loki, Tempo and Alertmanager - across metrics, logs, traces and alerting.
- Operator-level familiarity with data stores such as PostgreSQL and Message Brokers (e.g. RabbitMQ, RedPanda) - able to run and troubleshoot them in production.
- A good understanding of SRE practices - SLOs/error budgets, capacity planning - and a Platform-as-a-Product mindset.
- Strong grasp of incident management end to end: joining and declaring incidents, structured debugging under pressure, escalation, clear documentation, and post-mortems that drive real change.
- Able and willing to take part in a Follow-The-Sun on-call rotation from APAC hours, and to work effectively in a distributed, async-first team with strong written communication.
Who You Might Be (Bonus Points)
You can succeed in this role without all of the below, but any of these will help you ramp faster:
- Policy-as-code and IaC quality tooling (OPA/Conftest, Checkov, tflint, Atlantis, or similar).
- Experience managing Terraform state, module registries and versioning at scale across many teams.
- Experience building self-serve developer platforms and internal golden paths (e.g. with Backstage, Tilt, or similar).
- Experience with the Alloy collector and with incident tooling such as Rootly.
- Working proficiency in Go for automation and tooling.
- Strong Linux (Debian/Ubuntu) and container (Docker/containerd) fundamentals.
- Security and compliance experience in a regulated environment (SOC 2, secrets management, audit logging).
- Familiarity with trading, brokerage, or other regulated fintech domains, and with low-latency systems.
How We Take Care of You:
- Competitive Salary & Stock Options
- Health Benefits
- New Hire Home-Office Setup: One-time USD $500
- Monthly Stipend: USD $150 per month via a Brex Card
Alpaca is proud to be an equal opportunity workplace dedicated to pursuing and hiring a diverse workforce.
Recruitment Privacy Policy