资深站点可靠性工程师
Sr. Site Reliability Engineer
Veeam 是数据与人工智能信任公司,专注于帮助组织确保其数据和人工智能得到充分理解、安全保护,并具备弹性,以推动可扩展的安全人工智能发展。作为数据弹性和数据安全态势管理领域的市场领导者,Veeam 专为身份、数据、安全和人工智能风险的融合而打造。总部位于西雅图,在超过 30 个国家设有办公室,Veeam 全球保护着超过 55 万客户,他们信任 Veeam 让业务持续运行。加入我们,一起勇敢前行,共同成长、学习,并为世界上一些最大的品牌带来真正的影响力。
##### 关于该职位
Veeam 正在构建一个全球 SRE(站点可靠性工程)职能,以支持 Veeam 数据云,我们的新 SaaS 平台。该职位专注于政府和主权云环境。
由于权限和访问要求,该团队对政府基础设施的访问受到限制。这意味着你将属于一个小团队,负责整个平台堆栈——包括所有 VDC 工作负载。你并不总能将问题转交给其他团队;你需要足够了解整个架构,以便能够承担起责任。你需要快速掌握平台,通常通过阅读代码、文档和架构资料,而不是从第一天就直接访问环境。
这是一个从零开始的职位——你将帮助定义可靠性工程在这里的工作方式,通过映射系统、编写操作手册、设定基线,并建立该团队未来将遵循的实践。
##### 你将负责
##### 发现与文档
- 快速掌握整个平台——所有 VDC 工作负载、依赖关系和风险区域。大部分工作将通过代码、文档和交流完成,而不是直接访问环境。
- 与组织内的专家合作,填补知识空白,并为团队创建入职材料。
- 编写和维护操作手册、架构文档和操作指南。
##### 可靠性与事件响应
- 在 Azure(包括 Azure 政府版)上设计高可用性和容错的基础设施。
- 在目前没有的情况下,定义 SLIs、SLOs 和错误预算。
- 运行事件响应和无责复盘。将事件转化为改进。
- 识别现代和传统工作负载中的可靠性风险,并制定符合合规要求的实际修复计划。
##### 可观测性
- 弥补可观测性差距——定义指标和日志标准
查看英文原文
Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.
##### About The Role
Veeam is building a global SRE function to support the Veeam Data Cloud, our new SaaS platform. This role focuses on our Government and Sovereign Cloud environment.
Due to clearance and access requirements, this team operates with restricted access to GOV infrastructure. That means you'll be part of a small team responsible for the full platform stack — including all VDC workloads. You won't always be able to hand off problems to other teams; you need to understand the entire architecture well enough to own it. You'll need to get up to speed on the platform quickly, often by reading code, docs, and architecture artifacts rather than getting direct access to environments from day one.
This is a ground-up role — you'll help define how reliability engineering works here by mapping systems, writing runbooks, setting baselines, and building the practices this team will run on going forward.
##### What You'll Do
##### Discovery & Documentation
- Get up to speed on the full platform — all VDC workloads, dependencies, and risk areas. Much of this will happen through code, docs, and conversations rather than direct environment access.
- Work with SMEs across the org to fill knowledge gaps and build onboarding material for the team.
- Write and maintain runbooks, architecture docs, and operational guides.
##### Reliability & Incident Response
- Design infrastructure for high availability and fault tolerance on Azure (including Azure Government).
- Define SLIs, SLOs, and error budgets where none exist today.
- Run incident response and blameless postmortems. Turn incidents into improvements.
- Identify reliability risks across modern and legacy workloads and build practical remediation plans that work within compliance constraints.
##### Observability
- Close observability gaps — define instrumentation requirements and drive implementation.
- Set alerting, telemetry, and monitoring standards with partner teams.
- Build automation to reduce toil and support fleet management.
- Participate in on-call rotations.
##### Infrastructure & Delivery
- Work with IaC, CI/CD, deployment automation, and config management — including in air-gapped or compliance-restricted environments.
- Build and maintain testing, canary deployment, and release validation pipelines.
- Integrate chaos engineering and monitoring tools, adapting choices to meet regulatory requirements.
##### Collaboration
- Work across product, platform, security, legal, compliance, and operations teams.
- Own problems end-to-end — identify gaps, drive solutions, don't wait for direction.
- Mentor other engineers and help spread SRE practices across the org.
##### Technologies we work with
- Microsoft TFS, Azure DevOps, Git, BitBucket
- Azure (Entra ID, API Management, Cosmos Db, Storage services, Azure Functions, static website hosting, Azure security, etc.)
- IaC tools (Azure ARM templates, AWS CloudFormation, Terraform, the Serverless Framework, etc.)
- Observability (Azure Monitor, AppInsights, Elastic Stack)
##### What You'll Bring
- 7+ years in Software Engineering, with 3+ years in SRE, Platform Engineering, or similar — across multi-service platforms, not just single-service environments.
- Experience with Government or Sovereign Cloud (e.g., Azure Government, AWS GovCloud).
- Experience in regulated compliance environments — government (FedRAMP, CMMC, IL2/IL4/IL5), financial (PCI-DSS, SOX), or healthcare (HIPAA, HITRUST). You understand how compliance shapes architecture and operations.
- Strong experience building and running production services on cloud infrastructure (Azure preferred, including Azure Government).
- Able to learn large, complex platforms quickly with limited guidance — comfortable building understanding from code, docs, and architecture artifacts when direct environment access is restricted.
- Can investigate systems independently and produce clear docs, risk assessments, and improvement plans.
- Comfortable working across teams — engineering, product, security, compliance, operations.
- Programming skills in one or more of: TypeScript/JS, Go, Java, C#, or similar.
- Experience with monitoring and observability tools (e.g., Prometheus, Grafana, OpenTelemetry, ELK stack).
- Experience with IaC (Terraform, Terragrunt, Pulumi) and container orchestration (Kubernetes).
- Experience with CI/CD and GitOps tooling — GitHub Actions, Azure DevOps, GitLab CI, ArgoCD, FluxCD, or [Dagger](http://dagger.io/).
- Solid grasp of distributed systems, networking, and cloud-native architecture.
- Clear written and verbal communication skills
##### Bonus Skills
- Experience on B2B SaaS platforms in regulated or government markets.
- Background in chaos engineering, resilience testing, or performance/load testing.
- Have built an SRE or reliability function from scratch before.
- Experience across mixed environments — modern cloud-native and older legacy systems.
- Familiar with AI-first development workflows — using LLM-powered tools for infrastructure automation, code generation, and documentation.
##### Why Join?
- Build the GOV reliability practice from day one — your decisions will shape how this team works.
- Help define SRE at Veeam across a globally distributed engineering org.
- Work with strong teams across product, cloud engineering, security, and compliance.
- Professional development resources including mentorship, training, and volunteer days.
- Competitive compensation and benefits.
#LI-Remote
#LI-RW1
**What you'll get**
- Unlimited paid time off, 12 paid holidays including 4 global VeeaMe Days for self-care and 24 paid volunteer hours annually through Veeam Cares
- Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents
- Medical, dental, and vision coverage starting on your first day
- Mental health support, therapy sessions, and digital wellness tools via our Employee Assistance Program
- 401(k) retirement plan with company matching contributions
- Fertility, adoption, and surrogacy support through Maven, plus paid volunteer time
- AirVet: 24/7 virtual veterinary care at no cost
- Legal services, identity protection, and supplemental health insurance options
- Tax-advantaged spending accounts for healthcare, dependent care, and commuting
- Opportunities to learn and grow through on-demand libraries (LinkedIn Learning, O’Reilly), mentoring, workshops, and learning events like our annual Global Day of Learning
**Compensation Transparency**
Veeam is committed to pay transparency and equitable compensation. For this role, the compensation range below reflects the expected total target compensation (TTC), inclusive of base pay and a competitive performance-based bonus. For roles with a commission plan, the compensation range represents On Target Earnings (OTE), which includes base salary plus variable commission. When determining compensation, Veeam takes into consideration factors such as experience, education, skills, and geographic zone. Offers are typically made below the midpoint of the range.
In addition to compensation, Veeam provides a comprehensive benefits package, including health coverage, retirement plans, and unlimited time off.
**U.S. Geographic Zones & Compensation Ranges (TTC / OTE)**
Zone 1: San Francisco Bay Area, New York City Boroughs
$151,500—$252,500 USD
Zone 2: Washington, California (excluding San Francisco Bay Area)
$138,900—$231,400 USD
Zone 3: Texas, Illinois, North Carolina, Colorado, Massachusetts, Pennsylvania, Virginia, Oregon, Nevada, Hawaii, New York (excluding NYC boroughs); Sales roles located in Georgia, Ohio, and Arizona
$126,300—$210,400 USD
Zone 4: All other US locations
$109,800—$183,000 USD
**Veeam Software is an equal opportunity employer** and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.
Personal data collected during the recruitment process will be processed in accordance with our [Recruiting Privacy Notice](https://www.veeam.com/legal/recruiting-privacy-notice.html), which explains how your information is collected, used, and handled in connection with hiring activities. By applying for this position, you consent to this processing.
By submitting your application, you confirm that the information provided, including any supporting documents, is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification may result in disqualification from consideration or, if discovered after employment begins, termination of employment.