远程工作雷达

高级云工程师,SRE

Senior Cloud Engineer, SRE

开发工程限定地区(需当地身份)与中国几乎无重叠,需长期倒时差
公司Seismic
薪资$150,000 - $175,000/年
工作地点United States
地域资格限定地区(需当地身份)
时区要求与中国几乎无重叠,需长期倒时差
用工类型permanent
发布时间昨天
数据来源4dayweek.io
前往 4dayweek.io 查看并投递 →
注意地域限制:该职位明确限定在 United States 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。
作息提示:与中国几乎无重叠,需长期倒时差。

### 概述

Seismic 正在寻找一名高级云工程师,以提升我们在 AWS、Azure、IBM Cloud 和 OCI 环境中的可靠性。你将直接参与构建自动化工具、提高可观测性并加强事件响应,作为全球分布式 SRE 组织的一部分。我们优先采用云中立的架构、自动化和工程标准,以支持速度和规模。你将参与我们的可靠性路线图,提供可靠性最佳实践的技术指导,参与重大事件响应,并与产品和工程、安全以及面向客户的团队协作。

### 你是什么样的人:

- 在面向生产的 SRE 角色中,有支持复杂 SaaS 环境的经验。
- 你以好奇、尊重和协作的心态参与现有解决方案、流程和提案的讨论。
- 在分布式系统、多云环境、Kubernetes、网络、各种基础设施技术、GitOps 和 CI/CD 方面具备扎实的技术判断力。
- 有建立和成熟 SRE 原则和实践的经验,包括 SLO、错误预算、可观测性、容量规划、事件响应和减少重复劳动。
- 能够熟练利用可观测性数据解决高严重性事件,并在事后分析中深入挖掘根本原因。
- 有在高压事件中领导的经验,并能清晰地与技术团队、高管、面向客户的利益相关者和第三方供应商沟通。

### 你会做什么:

**运营模式**

- 设计、构建和维护减少运维负担的自动化工具和系统。
- 与产品和工程负责人合作,推动可靠性实践的成熟。
- 参与建设一个包容、高责任感的文化,鼓励好奇心、协作和无责改进。

**事件管理与运营卓越**

- 积极参与事件管理生命周期的健康和持续改进,包括检测、响应、升级、缓解、利益相关者沟通、事件回顾和后续纠正措施。
- 参与全球 SRE 团队的 12 小时轮班值班制度。
- 确保事件处理以客户为中心、数据驱动、无责且跨团队一致,同时保持准确的严重性和升级决策。
- 利用警报、事件、支持和 SLO 趋势推动组织从被动应对转向主动预防。

查看英文原文

### Overview

Seismic is seeking a Senior Cloud Engineer to advance reliability across our AWS, Azure, IBM Cloud, and OCI environments, working hands-on to build automation, improve observability, and strengthen incident response as part of a globally distributed SRE org.  We prioritize a cloud agnostic approach to architecture, automation, and engineering standards to support velocity and scale. You will contribute to our reliability roadmap, provide technical guidance on reliability best practices, participate in major incident response, and work collaboratively across Product & Engineering, Security, and Customer-facing teams.

### Who you are:

- Experience in a production facing SRE role supporting a complex SaaS environment.
- You approach discussions about existing solutions, processes, and proposals with curiosity, respect, and a collaborative mindset.
- Strong technical judgment across distributed systems, multi-cloud environments, Kubernetes, networking, various infrastructure technologies, GitOps, and CI/CD.
- Experience establishing and maturing SRE principles and practices, including SLOs, error budgets, observability, capacity planning, incident response, and toil elimination.
- Proficient in using observability data to resolve high-severity incidents and dig deeper into root cause during postmortems.
- Experience leading through high-pressure incidents and communicating clearly with technical teams, executives, customer-facing stakeholders, and third-party vendors.

### What you'll be doing:

**Operating Model**

- Design, build, and maintain automation and tooling that reduces operational toil.
- Contribute to maturing reliability practices in partnership with Product & Engineering leaders.
- Contribute to an inclusive, high-accountability culture that encourages curiosity, collaboration, and blameless improvement.

**Incident Management and Operational Excellence**

- Actively participate in the health and continuous improvement of the incident-management lifecycle, including detection, engagement, escalation, mitigation, stakeholder communication, post-incident review, and corrective-action follow-through.
- Participate in a 12-hour follow-the-sun on-call rotation within the Global SRE team.
- Ensure incident practices are customer-centered, data-driven, blameless, and consistent across teams while preserving accurate severity and escalation decisions.
- Use alert, incident, support, and SLO trends to move the organization from reactive response toward proactive risk reduction.
- Build measurable feedback loops that connect incident learning to engineering standards, service maturity, product priorities, and vendor actions.
- Work closely with application engineering teams and incorporate their feedback to improve developer experience and reduce toil.

**Reliability Strategy and Service Maturity**

- Partner with service owners to ensure production readiness standards are met before each release stage.
- Provide an SRE point of view on capacity planning, resilience testing, game days, disaster-recovery readiness, and modernization of fragile or legacy workloads.
- Partner with Product and Engineering leaders to document critical customer workflows, define health expectations, surface dependencies early, and align reliability investment with business priorities.
- Participate in cross-team reliability engagements, influencing outcomes without relying on direct authority.
- Build strategic relationships with vendors in the observability, incident response, and cloud infrastructure domains.

**AI-First Reliability Engineering**

- Responsibly adopt AI-assisted and agentic workflows for alert triage, incident mitigation, postmortems, trend analysis, capacity planning, SLO analysis, and self-service knowledge.
- Keep qualified humans in the decision loop for production-impacting actions.
- Improve the context available to reliability workflows by strengthening service metadata, observability data, incident records, runbooks, architecture documentation, and corrective-action quality.

### What we have for you:

At Seismic, we’re committed to providing benefits and perks for the whole self. To explore our benefits available in each country, please visit the Global Benefits page.

If you are an individual with a disability and would like to request a reasonable accommodation as part of the application or recruiting process, please click here.

Seismic is an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to gender, age, race, religion, or any other classification which is protected by applicable law.

We are committed to fair and equitable compensation practices.Seismic's annual base salary range for this position will vary based on applicant's location, experience, job level, skills, and abilities as well as internal equity and alignment market data. _The range listed below is the minimum to the maximum of our target hiring range._ Seismic's salary range for this position is: USD $150,000.00/Yr. - USD $175,000.00/Yr.This position is also eligible to participate in Seismic's incentive plans in addition to base salary. The actual incentive amount will very and will be subject to the terms and conditions set in the applicable incentive plan.

本页面信息整理自 4dayweek.io,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

高级全球销售总监 - 金融服务

SeismicUnited States$150,000 - $170,000/年permanent今天
市场运营职能支持限定地区(需当地身份)与中国几乎无重叠,需长期倒时差

企业客户经理 - 金融服务

SeismicUnited States$130,000 - $150,000/年permanent今天
市场运营职能支持限定地区(需当地身份)与中国几乎无重叠,需长期倒时差

合作伙伴项目经理

SeismicUnited States$70,400 - $106,370/年permanent今天
市场运营职能支持限定地区(需当地身份)与中国几乎无重叠,需长期倒时差

战略联盟总监

SeismicUnited States$80,900 - $139,600/年permanent昨天
AI市场运营限定地区(需当地身份)与中国几乎无重叠,需长期倒时差

产品营销副总裁

SeismicUnited States$215,200 - $300,000/年permanent5 天前
AI市场运营限定地区(需当地身份)与中国几乎无重叠,需长期倒时差

← 返回全部职位