远程工作雷达

站点可靠性工程师

Site Reliability Engineer

开发工程职能支持未标注地域与中国几乎无重叠,需长期倒时差
公司CSC Generation
薪资未公开
工作地点Costa Rica
地域资格未标注地域
时区要求与中国几乎无重叠,需长期倒时差
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
作息提示:与中国几乎无重叠,需长期倒时差。

冒险是我们的文化。加入一个崇尚如我们热爱的地形一样大胆生活方式的团队。在Backcountry,我们扎根于冒险、认可和山内外的福祉。我们展示员工故事,庆祝里程碑,并提供专属的户外福利。无论你是在总部、零售店还是远程办公,你都将成为一支充满活力、探索精神和联系感的团队的一员。

汇报对象:Gustavo Arguedas(站点可靠性经理)
地点:远程 - 哥斯达黎加

关于该职位

Backcountry的在线平台是客户体验的核心,该职位旨在确保其可靠性、性能和可扩展性。作为站点可靠性工程师,你将与软件工程、DevOps和IT运维团队合作,在多云架构上优化系统和应用。

在6-12个月内,你将为服务的弹性和可观测性做出有意义的改进,通过自动化减少运维负担,并在开发和基础设施团队中建立信任。

这是一个精简的团队。你将承担很多责任,快速行动,并对端到端的结果负责。

你将做的事情

  • 在Backcountry平台的各个部分工作,包括服务弹性、性能调优和系统设计
  • 推动关键事件的解决,并通过事后分析确保修复措施有条不紊地实施
  • 利用AI辅助工程工具(Claude Code、GitHub Copilot、基于MCP的代理)调查、自动化并发布基础设施和应用仓库中的修复
  • 通过设计和实现自动化来减少工作量
  • 与其他站点可靠性工程师、开发人员和架构师合作,评估并实施当前和未来工作负载的最佳实践
  • 监控系统健康状况和容量,主动采取措施在问题发生前解决
  • 与工程团队合作构建、部署和支持功能
  • 为Backcountry服务构建和维护可观测性(指标、日志、追踪、分析)以及SLI/SLO仪表化
  • 参与GCP和AWS上的FinOps计划,包括容量规划和承诺使用折扣策略
  • 参与SRE团队的值班支持轮班

所需资格

  • 3年以上支持容器化生产服务的经验,最好是运行Kubernetes的经验
  • 3年以上使用基础设施即代码(Terraform、AWS CDK、A
查看英文原文

Adventure is our Culture. Join a team that celebrates a lifestyle as bold as the terrain we love. At Backcountry, we are rooted in adventure, recognition, and wellbeing on and off the mountain. We spotlight employee stories, celebrate milestones, and offer exclusive outdoor perks. Whether you are at HQ, in a retail store, or remote, you will be part of a team that thrives on energy, exploration, and connection.
Reports to: Gustavo Arguedas (Site Reliability Manager)
Location: Remote - Costa Rica
About the Role

Backcountry's online platform serves as the backbone of our customer experience, and this role exists to ensure its reliability, performance, and scalability. As Site Reliability Engineer, you will partner with software engineering, DevOps, and IT operations teams to optimize systems and applications across a multi-cloud stack.

Within 6–12 months, you will have contributed meaningful improvements to service resiliency and observability, reduced operational toil through automation, and established yourself as a trusted partner to development and infrastructure teams.

This is a lean team. You will own a lot, move fast, and make decisions with full end-to-end responsibility.

What You'll Do

  • Work on service resiliency, performance tuning, and system design across Backcountry's platform
  • Drive resolution of critical incidents and ensure fixes are methodically implemented through postmortems
  • Leverage AI-assisted engineering tools (Claude Code, GitHub Copilot, MCP-based agents) to investigate, automate, and ship fixes across infrastructure and application repositories
  • Reduce toil by designing and implementing automation
  • Partner with other Site Reliability Engineers, developers, and architects to evaluate and implement best practices for current and future workloads
  • Monitor system health and capacity, taking proactive action to fix problems before they occur
  • Collaborate with engineering teams to build, deploy, and support features
  • Build and maintain observability (metrics, logs, traces, profiles) and SLI/SLO instrumentation for Backcountry services
  • Participate in FinOps initiatives across GCP and AWS, including capacity planning and committed-use discount strategy
  • Participate in the on-call support rotation within the SRE team

Required Qualifications

  • 3+ years of experience supporting containerized production services, preferably running Kubernetes
  • 3+ years of experience with Infrastructure as Code (Terraform, AWS CDK, Ansible, etc.)
  • 3+ years of cloud experience operating in Google Cloud Platform and/or AWS (multi-cloud stack; Azure/Entra exposure is a plus)
  • Comfortable diagnosing issues and shipping bug fixes directly to application code (not just infrastructure) to keep services reliable and stable
  • Comfortable performing deep dives across both infrastructure and application/software git repositories to trace issues end-to-end
  • Proficient with AI-assisted coding tools (e.g., Claude Code, GitHub Copilot) and MCP-based agents, used to accelerate investigation, code review, and automation
  • Strong knowledge of scripting and programming languages (Bash, Python, and TypeScript/Node.js)
  • Experience managing Linux (any major distribution) in production environments
  • Excellent understanding of internet application protocols (DHCP, DNS, HTTPS, SSH, etc.)
  • Understanding of how DevOps (CI/CD) and SRE practices (SLOs, SLIs) apply to daily work
  • Hands-on experience with observability tooling (Grafana, Prometheus, Loki, OpenSearch, or equivalents) and SLI/SLO instrumentation
  • Experience with GitOps and Kubernetes packaging (ArgoCD, Helm, Kustomize)
  • Proactively track emerging technology trends and developments, evaluating which ones are worth bringing into engineering practice
  • Bachelor's degree in computer science or similar, or equivalent experience
  • Advanced-level English communication skills, both verbal and written

Preferred Qualifications

  • Experience using AI coding assistants (Claude Code, Codex, GitHub Copilot) to build fixes, write automation, and ship application and infrastructure code improvements
  • Familiarity with PCI-scoped or other regulated environments
  • Previous experience working in ecommerce environments
  • Professional certifications: GCP, CKA, or AWS

Why Join

The people who do best here are builders. They take ownership, move fast, and want to see the direct impact of their work.

  • Cross-Functional Impact: Your work directly affects platform reliability for every customer and every team that depends on Backcountry's systems.
  • Modern Tech Stack: Work across a multi-cloud environment (GCP and AWS) with modern observability tooling, AI-assisted engineering, and GitOps workflows.
  • End-to-End Ownership: Own projects from investigation through implementation — you will ship automation, improve resiliency, and see the results in production.
  • Competitive Benefits: We offer an attractive benefits package including primarily remote work, private medical and life insurance, additional paid time off, monthly allowances and reimbursements, employee discounts, and opportunities for professional growth.

Interview Process

  • Recruiter Screen - A 30-minute conversation with our recruiting team to align on the role, your background, and what you are looking for.
  • Hiring Manager Interview - Conversation with the Site Reliability Manager focused on your SRE experience, approach to incident management, and team fit.
  • Technical/Case Discussion - A deeper dive into infrastructure, observability, and problem-solving scenarios relevant to the role.
  • Reference Checks - Conducted in parallel with the final stages where possible.
  • Offer - We move quickly for the right candidate.

Interview process is subject to change. Any updates will be communicated promptly and clearly.

CSC Generation is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other characteristic protected by law.

The CSC Generation family of brands is committed to providing reasonable accommodations for qualified individuals with disabilities in our job application procedures. If you need assistance or accommodation due to a disability, please contact .

Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

高级全栈工程师

CSC GenerationIndiaFull Time3 天前
开发工程限定地区(需当地身份)

← 返回全部职位