远程工作雷达

站点可靠性工程师(SRE)

Site Reliability Engineer (SRE)

开发工程职能支持全球可投(据职位描述推断)
公司Novacard
薪资未公开
工作地点Worldwide
地域资格全球可投(据职位描述推断)
时区要求日间重叠约 9 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
全球可投:该职位未限制候选人所在地区。仍需注意薪资可能按地区折算,以及实际签约方式(正式雇佣 / 独立合同)。

在NOVACARD,我们正在重新定义人们使用信用的方式。
我们是墨西哥首家免息且无年费的信用卡,旨在简化个人财务并让用户完全掌控——所有操作均可通过手机应用完成。使用NOVACARD,用户可获得最高200,000 MXN的信用额度,仅在使用时付款,并在5分钟内完成全部数字化管理。我们的使命是通过提供灵活性、透明度和他们实现目标所需的自由,帮助人们做出更明智的财务决策。简单财务,远大目标。
职位描述:
我们正在寻找一名站点可靠性工程师(SRE),以确保我们关键生产系统的稳定性、性能和可靠性。你将在开发与运维的交汇点工作——构建自动化工具、提升可观测性,并在问题发生前进行预防。
主要职责:

  • 确保生产系统的稳定性、性能和容错能力。
  • 开发和维护基础设施自动化和可观测性工具。
  • 监控系统健康状况,响应事件,并进行根本原因分析(RCA)。
  • 与开发团队合作,提升服务的可扩展性和可靠性。
  • 定义和管理SLI、SLO和错误预算。
  • 领导事件响应:组织恢复、记录RCA,并开展无责复盘。
  • 配置和管理Grafana和Zabbix,设计有洞察力的仪表盘,并优化告警。
  • 集成和监控外部供应商系统,在需要时与供应商技术团队协作。
  • 要求
  • 核心要求:
  • 精通俄语,英语B1+(能熟练阅读技术文档)。
  • 3年以上SRE、DevOps或基础设施工程师经验。
  • 对可观测性原则(指标、日志、追踪)有深入理解。
  • 有Grafana和Zabbix的实际操作经验(管理、配置、告警优化)。
  • 有使用AWS和CI/CD工具的经验。
  • 熟悉SLI/SLO/错误预算框架。
  • 有领导和记录事件及复盘的经验。
  • 具备自动化脚本编写能力(Python、Bash或Go)。
  • 对分布式系统和网络基础有扎实的理解。
  • 加分项:
  • 有监控和维护移动应用的经验。
  • 熟悉Terraform、Prometheus、Loki、ELK或其他类似工具。
  • 有使用Kubernetes和容器化环境的经验。
  • 我们提供的福利:
  • 全远程办公形式。
查看英文原文

At NOVACARD, we’re redefining how people use credit.
We are the first interest-free and no-annual-fee credit card in Mexico, designed to simplify personal finances and give users complete control - all from a mobile app. With NOVACARD, users can access up to $200,000 MXN in credit, only pay when they use it, and manage everything digitally in under 5 minutes. Our mission is to empower people to make smarter financial decisions by offering flexibility, transparency, and the freedom they need to reach their goals. Simple finances, big goals.About the Role:
We’re looking for a Site Reliability Engineer (SRE) to ensure the stability, performance, and reliability of our critical production systems. You’ll work at the intersection of development and operations — building automation tools, improving observability, and preventing incidents before they occur.
Key Responsibilities:

  • Ensure the stability, performance, and fault tolerance of production systems.
  • Develop and maintain infrastructure automation and observability tools.
  • Monitor system health, respond to incidents, and perform root cause analysis (RCA).
  • Collaborate with development teams to improve scalability and reliability of services.
  • Define and manage SLIs, SLOs, and Error Budgets.
  • Lead incident response: organize recovery, document RCA, and run blameless post-mortems.
  • Configure and administer Grafana and Zabbix, design insightful dashboards, and fine-tune alerting.
  • Integrate and monitor external vendor systems, collaborating with vendor technical support when needed.

Requirements
Key Requirements:

  • Fluent Russian, English B1+ (comfortable with technical documentation).
  • 3+ years of experience as an SRE, DevOps, or Infrastructure Engineer.
  • Strong understanding of observability principles (metrics, logs, traces).
  • Hands-on experience with Grafana and Zabbix (administration, configuration, alert optimization).
  • Experience working with AWS and CI/CD tools.
  • Practical knowledge of SLI/SLO/Error Budget frameworks.
  • Experience leading and documenting incidents and post-mortems.
  • Scripting skills for automation (Python, Bash, or Go).
  • Solid understanding of distributed systems and networking fundamentals.

Nice to Have:

  • Experience monitoring and supporting mobile applications.
  • Familiarity with Terraform, Prometheus, Loki, ELK, or similar tools.
  • Experience working with Kubernetes and containerized environments.

What We Offer:

  • Fully remote work format.
  • Official employment under the Russian Labor Code (for residents of Russia); contractor collaboration available for candidates from other countries.
  • Opportunity to work in an international team on a new digital product for the Mexican market.
  • A data-driven environment where your contributions have a real impact.

Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

← 返回全部职位