资深DevOps工程师
Staff DevOps Engineer
作为Alchemy的节点运维和基础设施工程师,你将在一个快节奏的工程团队中工作,参与设计、部署和持续改进支撑我们全球开发者平台的区块链基础设施。你将负责在多个链和多个地理区域上运行区块链RPC节点的先进系统,并利用你的经验确保为我们的客户保持高可用的实时系统。
职责:
- 在多个链和多个地理区域部署、运营和维护区块链RPC节点
- 管理代表所有区块链节点底层平台的Kubernetes集群
- 为EVM和非EVM链上的区块链客户端执行滚动升级和硬分叉迁移
- 参与值班轮班,通过PagerDuty处理现场事件,并协调跨团队解决节点中断、延迟峰值和SLO违规问题
- 开发和维护用于健康检查、自动修复和硬分叉通知的AI代理/自动化工具
- 通过ArgoCD和GitOps工作流(Helm图表)部署和管理服务
- 管理裸金属和云基础设施,包括配置、基准测试和硬件更换
- 及时响应安全通告,并在最小停机时间内协调升级
- 参与事后分析和异步评审流程;跟踪行动项并跟进解决
- 与产品、客户成功和其他工程团队跨职能协作,处理链淘汰、容量规划和SLO报告
我们寻找的人选:
- 具备设计和运营大规模、多区域、多云生产系统的经验
- 具备Kubernetes(k3s或类似)经验,包括StatefulSets、存储管理、Secrets和服务网格(Istio)
- 具备多集群环境中的密钥管理和访问控制经验
- 熟悉用于节点健康检查、升级和修复流程的自动化框架
- 具备基础设施即代码(如Terraform、Ansible、Pulumi、CloudFormation、Chef、Puppet等)经验
- 具备GitOps工具使用经验 - ArgoCD、Helm和部署管理
- (优先)具备服务网格部署经验,如Istio
- 精通云基础设施和裸金属管理,包括存储配置和快照管理
- 对可观测性工具有深入理解 - Grafana、Prometheus、Alertman
查看英文原文
As an engineer focused on node operations and infrastructure at Alchemy, you'll work within a fast-paced engineering team on the design, deployment, and continuous improvement of the blockchain infrastructure that powers our developer platform used globally. You'll operate state-of-the-art systems for running blockchain RPC nodes at scale across many chains and regions, and leverage your experience to keep a highly available live system running for our customers.
Responsibilities:
- Deploy, operate, and maintain blockchain RPC nodes across multiple chains and multiple geographic regions
- Manage Kubernetes clusters that represent the underlying platform for all blockchain nodes
- Perform rolling upgrades and hard fork migrations for blockchain clients across EVM and non-EVM chains
- Operate on-call rotations, triage live incidents via PagerDuty, and coordinate resolution across teams for node outages, latency spikes, and SLO breaches
- Develop and maintain AI agents / automation tooling for health checks, auto-heal, hard fork notifications
- Deploy and manage services via ArgoCD and GitOps workflows (Helm charts)
- Manage bare-metal and cloud infrastructure including provisioning, benchmarking, and hardware replacement
- Respond to security advisories promptly and coordinate upgrades with minimal downtime
- Contribute to postmortems and async review processes; track action items and follow up on resolutions
- Collaborate cross-functionally with product, customer success, and other engineering teams on chain deprecations, capacity planning, and SLO reporting
What We're Looking For:
- Experience designing and operating large-scale, multi-region, multi-cloud production systems
- Experience with Kubernetes (k3s or similar), including StatefulSets, storage management, Secrets and service mesh (Istio)
- Experience with secrets management and access control in multi-cluster environments
- Familiarity with automation frameworks for node health checks, upgrades, and remediation workflows
- Experience with Infrastructure-as-Code (e.g. Terraform, Ansible, Pulumi, CloudFormation, Chef, Puppet, etc)
- Experience with GitOps tooling - ArgoCD, Helm, and managing deployments
- (Preferred) Experience with service mesh deployments such as Istio
- Proficiency with cloud infrastructure and bare-metal management, including storage provisioning and snapshot management
- Strong grasp of observability tooling - Grafana, Prometheus, Alertmanager - and experience building or tuning dashboards and alert rules
- Comfort working in an on-call environment, triaging production incidents quickly and calmly using PagerDuty and structured runbooks
- Ability to write clear technical documentation and postmortems, and contribute to async-first team communication
- Experience with networking and configuring / managing VPC networks
- A basic understanding of security best practices
- (Preferred) Good understanding of web applications, microservice architecture
- (Preferred) Experience working with startups
- Passion for blockchain technologies and Web3
Perks:
- Attractive salary package
- Opportunity to work with the latest cloud and blockchain technologies
- Flexible time away
- Private Medical Insurance
- Start-up environment: internal off-site hackathons, access to company-rented hacker house during summer
- Opportunity to travel across offices