软件工程师负责人 - SRE
Lead Software Engineer - SRE
Nu 是拉丁美洲领先的数字银行,为巴西、墨西哥和哥伦比亚的 1.4 亿客户提供服务。公司通过利用数据和专有技术开发创新产品和服务,引领行业变革。
以“对抗复杂性,赋能人们”为使命,Nu 为客户完整的金融旅程提供服务,通过负责任的贷款和透明度促进金融准入和进步。公司由一种高效且可扩展的商业模式驱动,结合低成本服务与不断增长的回报。
Nu 的影响力已获得多项奖项的认可,包括《时代》100 家最具影响力公司、《快公司》最具创新力公司以及《福布斯》世界最佳银行。
访问我们的机构页面 https://www.nu.com/2026-en
关于该职位
美国市场团队正在推出一款差异化金融产品,进入全球最大且最复杂的金融市场。我们根据真实客户反馈快速迭代,同时构建最终能够服务于 Nubank 规模客户的系统。这种组合——早期阶段的速度、监管的重量以及对高可靠性的期望,需要一位以可靠性、规模和运营卓越为主要职责的工程师。
该职位的存在是为了确保我们今天构建的系统明天能够在生产环境中被信任,并为这个团队的“生产就绪”设定标准。该工程师通过编写生产代码、塑造架构和构建系统本身来履行其职责——而不是承担运维负载。
你将负责:
定义并执行 SLO。与产品和工程合作伙伴建立有意义的 SLI 和 SLO,管理错误预算,并将其作为实际的优先级输入,而非无人问津的仪表盘。
- 构建可观测层。改进指标、日志、追踪和告警,以便尽早发现故障,精准归因,并以代码级别的信心进行调试。将监控工具向上游集成到我们所拥有的服务中。
- 领导事故响应。在需要时担任事故负责人,推动无责复盘,并将发现转化为具体的工程工作并落地。在团队中建立这种能力,使它不集中于任何一个人。
- 通过工程减少重复劳动。识别重复的运维工作,并通过软件消除它——自动化、自愈行为、更好的默认设置等
查看英文原文
ABOUT NU
Nu is the leading digital bank in Latin America, serving 140 million customers across Brazil, Mexico, and Colombia. The company has been leading an industry transformation by leveraging data and proprietary technology to develop innovative products and services.
Guided by its mission to fight complexity and empower people, Nu caters to customers’ complete financial journey, promoting financial access and advancement with responsible lending and transparency. The company is powered by an efficient and scalable business model that combines low cost to serve with growing returns.
Nu’s impact has been recognized in multiple awards, including Time 100 Most Influential Companies, Fast Company’s Most Innovative Companies, and Forbes World’s Best Banks.
Visit our Institutional Page https://www.nu.com/2026-en
ABOUT THE ROLE
The U.S. Market team is launching a differentiated financial product in the largest and most demanding financial market in the world. We’re iterating quickly on real customer signals while building systems that will eventually serve customers at Nubank scale. That combination — early-stage velocity, regulatory weight, and high reliability expectations, requires an engineer whose primary mandate is reliability, scale, and operational excellence.
This role exists to make sure the systems we’re building today can be trusted in production tomorrow, and to set the bar for what “production-ready” means on this team. The engineer in this role delivers their mandate by writing production code, shaping architecture, and engineering the systems themselves — not by absorbing operational load.
YOU'LL BE RESPONSIBLE FOR
Define and operate against SLOs. Establish meaningful SLIs and SLOs with product and engineering partners, manage error budgets, and use them as real inputs to prioritization rather than dashboards no one reads.
- Build the observability layer. Improve metrics, logs, traces, and alerting so issues are detected early, attributed precisely, and debugged with code-level confidence. Push instrumentation upstream into the services we own.
- Lead incident response. Act as incident commander when needed, drive blameless postmortems, and turn findings into concrete engineering work that lands. Build the muscle in the team so this isn’t centralized in any one person.
- Reduce toil through engineering. Identify repetitive operational work and eliminate it with software — automation, self-healing behavior, better defaults, better tooling — rather than absorbing it as ongoing overhead.
- Production Hardening. Stress-test designs for partial failure, dependency degradation, traffic spikes, and adversarial inputs. Run capacity and performance work before incidents arise. Ensure resiliency primitives are tuned and working correctly.
- Make change safe and fast. Improve release safety through progressive delivery, feature flags, canaries, rollbacks, and tested migrations. Help the squad ship faster and with lower blast radius.
- Improve developer experience especially where it removes operational friction or improves change safety. Where internal tooling or platform gaps slow the team down, build or contribute the fix. Prefer leverage over heroics.
- Partner across disciplines. Work closely with product, platform, security, compliance, and other engineering teams. Translate reliability and risk tradeoffs into language each audience can act on.
- Raise the engineering bar. Mentor engineers, review hard designs and PRs, and shape technical standards across the squad. Lead through clarity and judgment, not authority.
WE ARE LOOKING FOR A PERSON WHO HAS
Track record of owning services in production — not just shipping them, but being the engineer responsible for how they behave under real load and real failure.
- Experience defining and operating against SLOs/SLIs, and using error budgets to influence engineering and product decisions.
- Experience leading incident response and writing postmortems that produced durable improvements.
- Hands-on experience with observability tooling (metrics, structured logging, distributed tracing) and using it to diagnose nontrivial production issues.
- Deep system design experience: distributed services, asynchronous messaging, storage tradeoffs, API design, idempotency, consistency, backpressure, and graceful degradation.
- Significant industry experience building and operating production software systems in a high-ownership engineering environment.
- Comfort operating in modern cloud environments (e.g., AWS/GCP), containerized workloads, and CI/CD pipelines, and reasoning about their failure modes.
- Demonstrated technical leadership: influencing architecture across teams, mentoring strong engineers, and making the people around you better.
- Pragmatism. You can hold a high reliability bar while still helping a fast-moving squad ship.
Location for this opportunity (City, Country)
- Miami, United States
OUR BENEFITS
- Opportunity of earning equity at Nu
- Medical Insurance
- Dental and Vision Insurance
- Life Insurance and AD&D
- Extended maternity and paternity leaves
- Nucleo - Our learning platform of courses
- NuLanguage - Our language learning program
- NuCare - Our mental health and wellness assistance program
- 401K
- Saving Plans - Health Saving Account and Flexible Spending Account
- Work-from-home Allowance
- Relocation Assistance Package, if applicable.
WORK MODEL FOR THIS ROLE
Hybrid 2-3 times/week: Our hybrid work model brings us to the office at least twice a week, on strategic days designed to maximize team connection and collaboration. For more details, visit https://building.nubank.com/nu-hybrid-work-model/
Explore how we build technology at Nubank:
🔗 https://building.nubank.com.br/building.nubank.com.br http://building.nubank.com.br ↗
🎥 https://www.youtube.com/@building.nubankyoutube.com/@building.nubank http://youtube.com/@building.nubank ↗
🎧 Listen to our stories on Spotify https://open.spotify.com/show/4hAJnAkYKXl42hLXa4LmVQ ↗
Our recruitment process may involve the use of artificial intelligence–enabled tools, such as automated interview transcription and analysis, to support the evaluation process. Artificial intelligence is not used to make final hiring decisions; all decisions are made by human reviewers.