资深站点可靠性工程师
Staff Site Reliability Engineer
关于公司
Visa 是支付技术领域的全球领导者,致力于在超过 200 个国家和地区之间为消费者、商户、金融机构和政府机构提供交易服务,致力于通过最佳的支付和收款方式提升每个人的生活。
在 Visa,你将有机会大规模地产生影响——解决有意义的挑战,提升你的技能,并看到你的贡献如何改变世界各地人们的生活。
加入 Visa,做真正重要的工作——对你、对你的社区、对世界都有意义。进步从你开始。
职位描述
资深站点可靠性工程师负责设计、实施和维护确保 Visa 应用服务高可用性和可靠性的解决方案。该角色通过帮助定义服务必须满足的可靠性期望,指导应用团队进行弹性架构和设计选择,并创建可重复的证据证明关键服务能够承受现实中的故障模式来做出贡献。由于团队规模较小,候选人需要能够独立工作,能在有限指导下领导复杂的工程技术工作,并在没有直接管理权的情况下影响应用拥有团队。
所有职位都需要具备数字素养,包括能够使用生成式 AI 工具(如 ChatGPT、Microsoft Copilot)等新兴技术来支持日常工作的能力。
可靠性与弹性工程小组致力于提升面向客户的服务的可靠性和弹性。该团队在混沌工程、可靠性标准、弹性设计、可观测性以及持续可靠性改进方面提供专业知识。该顾问级职位期望在该职能内作为高自主性的技术负责人运作,从模糊中创造清晰度,并将业务可靠性目标转化为实际的工程标准、实验、技术指导和改进计划。
主要职责:
- 负责优先服务中的复杂弹性工程工作。
- 设计、规划、执行并报告受控混沌实验和演练日。
- 建立可重复的混沌测试模式。
- 定义可靠性和弹性标准。
- 将事件、观察到的故障模式、SLO 未达标情况和实验结果转化为可操作的工程改进。
- 为应用团队提供架构和系统设计指导,特别是在电路板相关方面
查看英文原文
About Us
Visa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid.
At Visa, you'll have the opportunity to create impact at scale — tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world.
Join Visa and do work that matters – to you, to your community, and to the world. Progress starts with you.
Job Description
The Staff Site Reliability Engineer is responsible for architecting, implementing, and maintaining solutions that ensure Visa’s application services operate with high availability and reliability. This role contributes by helping define the reliability expectations that services must meet, guiding application teams through resilient architecture and design choices, and creating repeatable evidence that critical services can tolerate realistic failure modes. Because the team is still small, the candidate must be able to work independently, lead complex technical work with limited guidance, and influence application-owning teams without direct authority.
All roles require digital fluency, including the ability to work with emerging technologies such as Generative AI tools (e.g. ChatGPT, Microsoft Copilot) to support everyday work.
The Reliability & Resilience Engineering squad works to improve the reliability and resilience of the services sold to customers. The team provides subject matter expertise in chaos engineering, reliability standards, resilience design, observability, and continuous reliability improvement. This consultant-level role is expected to operate as a high-autonomy technical owner within the function, creating clarity from ambiguity and translating business reliability goals into practical engineering standards, experiments, technical guidance, and improvement plans.
Key Responsibilities:
- Own complex resilience engineering work across priority services.
- Design, plan, conduct, and report on controlled chaos experiments and game days.
- Establish repeatable chaos testing patterns.
- Define reliability and resilience standards.
- Translate incidents, observed failure modes, SLO misses, and experiment findings into actionable engineering improvements.
- Provide architecture and system-design guidance to application teams, especially around circuit breakers, load shedding, timeouts, retries, dependency isolation, graceful degradation, observability, alerting, and failure containment.
- Author technical design documents, experiment plans, standards, post-experiment reports, and improvement proposals.
- Mentor engineers through design reviews, technical guidance, and hands-on support to raise the reliability capability of the teams they work with.
- Deliver documented resilience standards, reusable experiment templates, clear production readiness criteria, completed pilot experiments, prioritised remediation backlogs, and measurable improvements in the resilience posture of selected customer-critical services.
This is a remoteposition. A remote position does not require job dutiesbeperformed within proximity of a Visa office location. Remotepositions maybe requiredto be present at a Visa office with scheduled notice.
Qualifications
Basic Qualifications:
- 5+ years of relevant work experience with a Bachelor’s Degree or at least 2 years of work experience with an Advanced degree (e.g. Masters, MBA, JD, MD) or 0 years of work experience with a PhD, OR 8+ years of relevant work experience.
- Experience in Kubernetes and related technologies such as Helm, ArgoCD, and Terraform.
- Experience in automating complex tasks and processes using programming languages such as Python, Go, and Java.
- Experience in planning and performing chaos experiments using tooling such as Gremlin, Chaos Toolkit, Litmus, or equivalents
Preferred Qualifications:
- 6 or more years of work experience with a Bachelor's Degree or 4 or more years of relevant experience with an Advanced Degree (e.g. Masters, MBA, JD, MD) or up to 3 years of relevant experience with a PhD.
- Experience working with Agile teams.
- Experience in fast-paced 24x7 environments.
- Experience in distributed systems design and maintenance.
- Experience in cloud and system architecture across hybrid platforms (AWS, GCP, Azure).
- Experience in observability and performance optimization techniques.
- Experience working as a software developer
Visa is an EEO Employer
Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability or protected veteran status. Visa will also consider for employment qualified applicants with criminal histories in a manner consistent with EEOC guidelines and applicable local law.
Originally posted on Himalayas