高级IT工程师(监管可靠性) - 墨西哥 - 2026
Senior IT Engineer (Regulatory Reliability) - MX - 2026
Nu 是拉丁美洲领先的数字银行,为巴西、墨西哥和哥伦比亚的 1.4 亿客户提供服务。公司通过利用数据和专有技术,开发创新产品和服务,引领行业变革。
秉承其“对抗复杂性,赋能人们”的使命,Nu 为客户提供完整的金融旅程,通过负责任的贷款和透明度促进金融准入和进步。公司由一种高效且可扩展的商业模式驱动,结合低成本服务与不断增长的回报。
Nu 的影响力已获得多项奖项的认可,包括《时代》100 家最具影响力公司、《快公司》最具创新力公司以及《福布斯》全球最佳银行。
访问我们的机构页面 https://www.nu.com/2026-en
关于该职位
我们正在寻找一名高级可靠性工程师,负责 Nubank 的监管基础设施的可靠性、弹性和运营卓越性。在此职位中,您将负责支持实时跨行转账、监管集成和关键监管平台的系统的关键可靠性结果。您将在基础设施、软件、网络、安全和运营之间协作,确保在严格的监管和操作限制下,关键服务保持可用、可观测和具有弹性。
您将负责的工作
- 负责 IT 监管基础设施服务的端到端可靠性,包括连接性、交易处理流程和与结算相关的操作工作流。
- 定义、衡量并持续改进关键交易和平台的 SLI、SLO 和运营健康指标。
- 领导不同严重程度生产事件的事故响应,协调恢复工作,并推动高质量的事故回顾,对纠正措施进行具体跟进。
- 设计和演进平台的可观测性,重点在于追踪、队列健康状况、签名验证、基础设施信号以及监管链接退化或交易瓶颈的早期检测。
- 规划和执行灾难恢复和业务连续性演练,包括故障转移验证、应急准备和 RTO/RPO 验证。
- 通过自动化、工具和改进的运行手册减少操作负担,应对重复性故障和操作流程。
- 推动容量规划和性能工程,确保平台能够安全地吸收高峰交易量和 ev
查看英文原文
ABOUT NU
Nu is the leading digital bank in Latin America, serving 140 million customers across Brazil, Mexico, and Colombia. The company has been leading an industry transformation by leveraging data and proprietary technology to develop innovative products and services.
Guided by its mission to fight complexity and empower people, Nu caters to customers’ complete financial journey, promoting financial access and advancement with responsible lending and transparency. The company is powered by an efficient and scalable business model that combines low cost to serve with growing returns.
Nu’s impact has been recognized in multiple awards, including Time 100 Most Influential Companies, Fast Company’s Most Innovative Companies, and Forbes World’s Best Banks.
Visit our Institutional Page https://www.nu.com/2026-en
About the role
We are looking for a Senior Reliability Engineer to drive reliability, resilience, and operational excellence for Nubank's regulatory infrastructure. In this role, you will own critical reliability outcomes for systems that support real-time interbank transfers, regulatory integrations and key regulatory platforms. You will work across infrastructure, software, networking, security, and operations to keep mission-critical services available, observable, and resilient under demanding regulatory and operational constraints.
What you'll do
- Own end-to-end reliability for IT regulatory infrastructure services, including connectivity, transaction processing flows, and settlement-related operational workflows.
- Define, measure, and continuously improve SLIs, SLOs, and operational health indicators for critical transactions and platforms.
- Lead incident response for different severity production events, coordinate recovery, and drive high-quality postmortems with concrete follow-through on corrective actions.
- Design and evolve observability for the platform, with emphasis on tracing, queue health, signature validation, infrastructure signals, and early detection of degraded regulatory links or transaction bottlenecks.
- Plan and execute disaster recovery and business continuity exercises, including failover validation, contingency readiness, and RTO/RPO verification.
- Reduce operational toil through automation, tooling, and improved runbooks for recurring failure and operational procedures.
- Drive capacity planning and performance engineering to ensure the platforms can safely absorb peak transaction volumes and evolving business demand.
- Partner closely with security, networking, middleware, and software engineering teams to improve platform hardening, resilience, and change management safety.
- Participate in an on-call rotation.
What we're looking for
- 5+ years of SRE, DevOps, or production engineering experience operating mission-critical, high-availability systems.
- Hands-on experience with IT infrastructure platforms, virtualized and hyperconverged environments such as VxRail, and physical production infrastructure.
- Hands-on experience with Linux systems, including performance tuning, kernel parameters, and security hardening.
- Track record leading incident response and postmortem processes for customer-impacting services.
- Solid knowledge of networking fundamentals: TCP/IP and routing.
- Proficiency in at least one scripting or programming language (Python, Shell scripting) for automation and tooling.
- Experience with observability stacks (Prometheus, Grafana) and distributed tracing.
- AWS infrastructure experience across services like EC2, S3, CloudWatch, KMS, EKS/ECS, VPCs, RDS/DynamoDB, and SQS/SNS.
- IT Regulatory audit experience, including evidence gathering, control validation, audit prep, and remediation follow-through.
- Experience with infrastructure-as-code tools (like Puppet) and Git-based workflows.
- Working knowledge of AI-assisted engineering tools (Claude, Cursor, Claude Code).
- Strong written and verbal communication skills in Spanish and English.
Nice to have
- Experience with Mexican payment rails (SPEI)
- SQL knowledge for operational analysis and troubleshooting.
Relevant cloud (AWS Certified Cloud Practitioner certification or higher) or IT infrastructure certifications.
Location for this opportunity
- Mexico City, Mexico
OUR BENEFITS
- Chance of earning equity at Nu
- Extended maternity and paternity leaves
- Health and life insurance
- Dental and Vision Insurance
- NuCare - Our mental health and wellness assistance program
- Nucleo - Our learning platform of courses
- NuLanguage - Our language learning program
- Holiday Bonus ("Aguinaldo") of 30 days of pay per year
- 17 days of paid vacation with 25% vacation bonus
- Gym partnership
- Food card
- Work-from-home Allowance
- Parental Consultancy
- Relocation Assistance Package, if applicable
WORK MODEL FOR THIS ROLE
- Hybrid 2-3 times/week: Our hybrid work model brings us to the office at least twice a week, on strategic days designed to maximize team connection and collaboration. For more details, visit https://building.nubank.com/nu-hybrid-work-model/
Our recruitment process may involve the use of artificial intelligence–enabled tools, such as automated interview transcription and analysis, to support the evaluation process. Artificial intelligence is not used to make final hiring decisions; all decisions are made by human reviewers.