远程工作雷达

资深软件工程师,可靠性 - Command|Alert

Staff Software Engineer, Reliability - Command|Alert

开发工程限定地区(需当地身份)
公司CommandLink
薪资未公开
工作地点Philippines
地域资格限定地区(需当地身份)
时区要求日间重叠约 9 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 Philippines 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

关于Command|Link

Command|Link是一个全球SaaS平台,提供网络、语音服务和IT安全解决方案,帮助公司将其核心基础设施整合到单一供应商,并在其上构建专有的统一管理平台。Command|Link通过解决竞争对手制造的问题,彻底改变了IT行业。由于在创新和承诺方面的非凡表现,Command|Link被评为年度SD-WAN产品、ITSM愿景灯塔、年度UCaaS产品、年度NaaS产品、年度供应商以及AT&T战略增长合作伙伴。Command|Link打造了唯一一个可扩展的IT平台,解决了ISP供应商分散和IT难题。我们让客户更容易完成更多工作,最大化停机时间并改善利润。

了解更多关于我们!

这是一个100%远程职位

关于你的新角色:

Command|Alert是CommandLink的信号处理核心,是将原始安全、监控和客户定义的遥测数据转化为客户真正信任的警报的引擎。警报疲劳和噪音是该领域所有竞争对手的首要投诉,而这个职位的存在就是为了确保我们的警报是人们不会忽略的。

作为Command|Alert的高级软件工程师,你将负责整个警报管道的可靠性,并推动组织在架构方面最重要的决策。这个职位适合那些既像站点可靠性工程师又像软件工程师的人:精通SLO、错误预算和无责事故响应,具备跨安全工具、监控遥测、syslog、OpenTelemetry和L2-L4网络协议进行技术推理的能力。

主要职责:

  • 负责从OpenSearch警报评估到通过Kafka通过OpenSearch回调传递到下游通知的整个警报管道的可靠性,包括幂等性保证和在持续负载下的浸泡测试行为。
  • 为管道的可用性、延迟和交付保证定义SLO和SLI,并使用错误预算来指导可靠性工作与新功能投资之间的平衡。
  • 领导管道最严重的生产问题的事故响应,进行无责的复盘,并推动由此产生的系统性修复和自动化。
  • 构建运行该管道所需的可观测性和自动化,以满足财富1000强规模的低运维负担,并负责容量
查看英文原文

About Command|Link

Command|Link is a global SaaS Platform providing network, voice services, and IT security solutions, helping corporations consolidate their core infrastructure into a single vendor and layering on a proprietary single pane of glass platform. Command|Link has revolutionized the IT industry by tackling the problems our competitors create. In recognition for our unprecedented innovation and dedication, Command|Link was recognized as the SD-WAN Product of the Year, ITSM Visionary Spotlight, UCaaS Product of the Year, NaaS Product of the Year, Supplier of the Year, and the AT&T Strategic Growth Partner. Command|Link has built the only IT platform for scale that solves ISP vendor sprawl and IT headaches. We make it easy for our customers to get more done, maximize uptime and improve the bottom line.

Learn more about us here!

This is a 100% remote position

About your new role:

Command|Alert is CommandLink's signal-processing core, the engine that turns raw security, monitoring, and customer-defined telemetry into alerts customers actually trust. Alert fatigue and noise are the top complaint across every competitor in this space, and this role exists to make sure our alerts are the ones people don't tune out.

As a Staff Software Engineer on Command|Alert, you'll own the reliability of the alerting pipeline end to end and drive the org's most consequential decisions on how it's architected. This role suits someone who thinks like a site reliability engineer as much as a software engineer: fluent in SLOs, error budgets, and blameless incident response, with the technical range to reason across security tooling, monitoring telemetry, syslog, OpenTelemetry, and L2-L4 network protocols.

Key Responsibilities:

  • Own the reliability of the alerting pipeline end to end, from OpenSearch alert evaluation through Kafka delivery via OpenSearch callbacks to downstream notification, including idempotency guarantees and soak-tested behavior under sustained load.
  • Define SLOs and SLIs for the pipeline's availability, latency, and delivery guarantees, and use error budgets to guide how much investment goes into reliability work versus new capability.
  • Lead incident response for the pipeline's most critical production issues, running blameless post-mortems and driving the systemic fixes and automation that come out of them.
  • Build the observability and automation needed to run the pipeline at Fortune 1000 scale with low operational toil, and own capacity planning as ingestion volume and customer count grow.
  • Set the architecture for how Command|Alert evaluates rule-based thresholds, ML anomaly scores, and correlation logic, turning diverse telemetry into usable network and system topologies that power LLM-driven investigation and remediation.
  • Mentor engineers across the teams you touch and represent Command|Alert's technical direction to stakeholders outside engineering.
  • Takes on additional responsibilities and projects as needed to support the success of the team and organization.

What you'll need for success:

Required

  • Background operating as a Site Reliability Engineer, DevOps engineer, or in a similar production-ownership role, with fluency in SLOs, SLIs, error budgets, and blameless incident response.
  • Demonstrated experience designing, building, or operating high-reliability alerting or notification systems in production, including rule-based and ML-based detection at scale.
  • Strong Kafka experience, and a track record building systems where webhook reliability, idempotency, and delivery guarantees under load are non-negotiable.
  • Experience building the observability and automation that let a high-volume production system run with low operational toil.
  • A working command of telemetry and protocol data (security tooling output, syslog, OpenTelemetry, NetFlow/sFlow, SNMP, ICMP, firewall logs) and the ability to turn it into real network and system topologies.
  • Recognized mastery of Go and/or Python, with range across container orchestration and a multi-cloud footprint, and a demonstrated ability to make org-level architecture calls.

Nice to Have

  • Experience with chaos engineering or fault-injection testing (e.g., Gremlin, Chaos Mesh).
  • Familiarity with SLO and error-budget tooling and practice (e.g., Nobl9, Google's SRE workbook approach).
  • Experience with Temporal or a comparable workflow orchestration platform.
  • Experience operating in a multi-tenant, cloud-native environment with secrets management and TLS at scale.

Why you'll love life at Command|Link:

Join us at CommandLink, where you'll have the opportunity to shape the future of business communication. We value the innovative spirit and seek individuals ready to bring their unique vision and expertise to a team that values bold ideas and strategic thinking. Are you ready to make an impact?

  • Room to grow at a high-growth company
  • An environment that celebrates ideas and innovation
  • Your work will have a tangible impact
  • Flexible time off
  • Fun events at cool locations
  • Employee referral bonuses to encourage the addition of great new people to the team

At CommandLink, we’re committed to creating a fair, consistent, and efficient hiring experience. As part of our process, we use AI-assisted tools to help review and analyze applications. These tools support our recruiting team by identifying qualifications and experience that align with the requirements of each role.

AI tools are used only to assist in the evaluation process — they do not make final hiring decisions. Every application is reviewed by a member of our recruiting or hiring team before any decisions are made.

Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

资深软件工程师, 平台工程

CommandLinkColombia£120,000 - £160,000/年Full Time2 天前
开发工程限定地区(需当地身份)与中国几乎无重叠,需长期倒时差

← 返回全部职位