远程工作雷达

NOC 工程师 / SRE

NOC Engineer / SRE

开发工程职能支持限定地区(需当地身份)
公司NICE
薪资未公开
工作地点United Kingdom - Remote
地域资格限定地区(需当地身份)
时区要求无特别要求
用工类型未标注
发布时间今天
数据来源Greenhouse
前往企业招聘页投递 →
注意地域限制:该职位明确限定在 United Kingdom - Remote 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

在NiCE,我们从不给自己设限。我们不断挑战自己的极限。我们志向远大,是变革者,我们追求胜利。我们设定最高标准,并超越它们。如果你和我们一样,我们可以为你提供一个能点燃你内心热情的终极职业机会。

那么,这个职位是做什么的?

SRE-NOC职位位于传统网络操作中心(NOC)职责与工程驱动的可靠性实践的交汇点。该职位专注于24/7服务可靠性、事件响应、运维自动化和可观测性,同时通过软件和自动化主动减少运维负担。

与传统的NOC分析师不同,SRE-NOC需要通过工程手段解决问题,而不仅仅是响应警报。

你将如何产生影响?

事件响应与运维

  • 在24x7轮值值班中担任主要或升级响应人员
  • 领导或支持重大事件(MI)响应,包括初步评估、缓解和解决
  • 协调工程、基础设施、安全和产品团队
  • 执行并改进运行手册、操作指南和升级路径
  • 推动无责复盘(PIR)并跟踪纠正措施

监控、告警与可观测性

  • 负责跨基础设施、应用程序和依赖项的服务健康监控
  • 设计并维护符合SLI/SLO的告警策略
  • 通过信号与噪声优化减少告警疲劳
  • 使用以下工具构建仪表板:
  • Grafana
  • Prometheus
  • Datadog / Splunk / CloudWatch

可靠性工程与自动化

  • 自动化重复性运维任务以减少人工负担
  • 提高平均检测时间(MTTD)和平均修复时间(MTTR)
  • 开发脚本和工具(Python、Bash、Go等)以支持NOC/SRE工作流程
  • 在可能的情况下实现自愈和自动修复
  • 与工程团队合作,提升系统设计的可靠性

平台与基础设施支持

  • 支持和排查:
  • 基于Linux的系统
  • 云平台(AWS、Azure、GCP)
  • Kubernetes/容器化环境
  • 协助进行容量规划和可用性评审
  • 确保生产发布具备运营就绪状态

你是否具备所需的素质?

技术方面

  • 强大的Linux系统管理能力
  • 具有事件管理和生产支持经验
  • 熟悉:
  • 云基础设施(AWS优先)
  • 容器与编排(Docker,
查看英文原文

At NiCE, we don’t limit our challenges. We challenge our limits. Always. We’re ambitious. We’re game changers. And we play to win. We set the highest standards and execute beyond them. And if you’re like us, we can offer you the ultimate career opportunity that will light a fire within you.

So, what’s the role all about?

The SRE – NOC role sits at the intersection of traditional Network Operations Center (NOC) responsibilities and engineering‑driven reliability practices. This role focuses on 24/7 service reliability, incident response, operational automation, and observability, while actively reducing operational toil through software and automation.

Unlike a traditional NOC analyst, an SRE‑NOC is expected to engineer problems away, not just respond to alerts.

How will you make an impact?

Incident Response & Operations

  • Act as a primary or escalation responder in a 24x7 on‑call rotation
  • Lead or support Major Incident (MI) response, including triage, mitigation, and resolution
  • Coordinate across Engineering, Infrastructure, Security, and Product teams
  • Execute and improve runbooks, playbooks, and escalation paths
  • Drive blameless post‑incident reviews (PIRs) and track corrective actions

Monitoring, Alerting & Observability

  • Own service health monitoring across infrastructure, applications, and dependencies
  • Design and maintain alerting strategies that align with SLIs/SLOs
  • Reduce alert fatigue through signal‑to‑noise improvements
  • Build dashboards using tools such as:
  • Grafana
  • Prometheus
  • Datadog / Splunk / CloudWatch

Reliability Engineering & Automation

  • Automate repetitive operational tasks to reduce manual toil
  • Improve mean time to detect (MTTD) and mean time to resolve (MTTR)
  • Develop scripts and tools (Python, Bash, Go, etc.) to support NOC/SRE workflows
  • Implement self‑healing and auto‑remediation where possible
  • Partner with engineering teams to improve system design for reliability

Platform & Infrastructure Support

  • Support and troubleshoot:
  • Linux‑based systems
  • Cloud platforms (AWS, Azure, GCP)
  • Kubernetes / containerized environments
  • Assist with capacity planning and availability reviews
  • Ensure operational readiness for production releases

Have you got what it takes?

Technical

  • Strong Linux systems administration
  • Experience with incident management and production support
  • Familiarity with:
  • Cloud infrastructure (AWS preferred)
  • Containers & orchestration (Docker, Kubernetes)
  • Monitoring/alerting platforms
  • Scripting or programming experience in Python, Bash, Go, or similar
  • Understanding of networking fundamentals (DNS, TCP/IP, load balancing)

Operational

  • Experience working in 24x7 NOC or production operations environments
  • Ability to handle high‑pressure incidents calmly and effectively
  • Strong written and verbal communication for incident coordination
  • Comfort working from runbooks—but improving them when they fall short

Preferred / Differentiators

  • Experience defining or operating to SLOs / SLIs
  • Prior migration from traditional NOC → SRE model
  • Infrastructure as Code experience (Terraform, Ansible, etc.)
  • Exposure to security, compliance, or regulated environments

Requisition ID: 11707.

Reporting into: Manager, Network Operations.

Role Type: Individual Contributor.

#LI-Remote
About NiCE

NICE Ltd. (NASDAQ: NICE) software products are used by 25,000+ global businesses, including 85 of the Fortune 100 corporations, to deliver extraordinary customer experiences, fight financial crime and ensure public safety. Every day, NiCE software manages more than 120 million customer interactions and monitors 3+ billion financial transactions.

Known as an innovation powerhouse that excels in AI, cloud and digital, NiCE is consistently recognized as the market leader in its domains, with over 8,500 employees across 30+ countries.

NiCE is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, age, sex, marital status, ancestry, neurotype, physical or mental disability, veteran status, gender identity, sexual orientation or any other category protected by law.

本页面信息整理自 Greenhouse,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位