系统工程师负责人 (Kafka)
Lead Systems Engineer (Kafka)
Nu 是拉丁美洲领先的数字银行,为巴西、墨西哥和哥伦比亚的 14000 万客户提供服务。公司通过利用数据和专有技术推动行业变革,开发创新产品和服务。
以对抗复杂性并赋能人们为使命,Nu 为客户完整的金融旅程提供服务,通过负责任的贷款和透明度促进金融准入和进步。公司采用高效且可扩展的商业模式,结合低成本服务与不断增长的回报。
Nu 的影响力已获得多项奖项的认可,包括《时代》100 家最具影响力公司、《快公司》最具创新力公司以及《福布斯》全球最佳银行。
访问我们的机构页面 https://www.nu.com/2026-en
职位描述
我们正在寻找一位经验丰富的软件工程师,帮助演进和运营 Nubank 的消息传递平台以及支持大规模异步通信的基础架构。
该职位所在的团队负责一系列关键平台功能,这些功能支持多个业务领域和国家的内部系统。平台运行在一个大型且复杂的环境中,包含数百个集群、数千个代理、数十万个主题,并在多个 AWS 账户中处理非常大的每日数据量。
在高级别职位上,我们希望找到能够独立负责重要技术问题、提升可靠性和可操作性,并与团队合作推动工程决策的人选。具备 Kafka 经验是加分项,但不是必需条件。对分布式系统基础设施有深入理解,特别是 Kubernetes、网络和 AWS,是基本要求。
您将负责以下工作
- 运营和改进基于 Kafka 的大规模消息传递和平台基础架构,这些架构被 Nubank 的关键系统使用
- 为异步通信平台的可靠性、可扩展性和性能做出贡献
- 协助设计和实现高吞吐量、低延迟和容错系统
- 提升平台的可观测性、自动化和运营卓越性
- 支持生产环境中的事件分析、故障排查和根本原因修复
- 优化基础架构使用情况,帮助在基于 AWS 的环境中提高效率和成本意识
- 开发平台功能,以支持消息的安全增长
查看英文原文
ABOUT NU
Nu is the leading digital bank in Latin America, serving 140 million customers across Brazil, Mexico, and Colombia. The company has been leading an industry transformation by leveraging data and proprietary technology to develop innovative products and services.
Guided by its mission to fight complexity and empower people, Nu caters to customers’ complete financial journey, promoting financial access and advancement with responsible lending and transparency. The company is powered by an efficient and scalable business model that combines low cost to serve with growing returns.
Nu’s impact has been recognized in multiple awards, including Time 100 Most Influential Companies, Fast Company’s Most Innovative Companies, and Forbes World’s Best Banks.
Visit our Institutional Page https://www.nu.com/2026-en
About the Role
We are looking for an experienced software engineer to help evolve and operate Nubank’s messaging platform and the infrastructure that supports asynchronous communication at scale.
This role sits in a team responsible for highly critical platform capabilities that support a wide range of internal systems across multiple business domains and countries. The platform operates in a large and complex environment, with hundreds of clusters, thousands of brokers, hundreds of thousands of topics, and very large daily data volumes across multiple AWS accounts.
At the Lead level, we are looking for someone who can independently own important technical problems, improve reliability and operability, and drive engineering decisions in partnership with the team. Kafka experience is desirable, but not required. Strong knowledge of distributed systems infrastructure, especially Kubernetes, networking, and AWS, is essential.
WHAT YOU’LL BE RESPONSIBLE FOR
- Operate and improve large-scale messaging and platform infrastructure based on kafka used by critical systems across Nubank
- Contribute to the reliability, scalability, and performance of asynchronous communication platforms
- Help design and implement solutions for high-throughput, low-latency, and fault-tolerant systems
- Improve observability, automation, and operational excellence across the platform
- Support incident analysis, troubleshooting, and root cause remediation in production environments
- Optimize infrastructure usage and help drive efficiency and cost awareness across AWS-based environments
- Work on platform capabilities that enable safe growth in message volume, topic count, and cluster footprint
- Partner with other engineers and teams to evolve platform standards, tooling, and best practices
- Contribute to architectural discussions involving messaging, traffic patterns, service communication, and platform reliability
WE ARE LOOKING FOR A PERSON WHO HAS
Must-have
- Strong software engineering fundamentals and experience working with distributed systems in production.
- Solid experience with Kubernetes, networking, and AWS in large-scale or business-critical environments.
- Experience operating infrastructure-heavy platforms with high reliability and availability requirements.
- Ability to troubleshoot complex production issues across application, infrastructure, and network layers.
- Experience improving observability, automation, and operational tooling.
- Good understanding of scalability, resilience, performance, and failure isolation patterns.
- Ability to work autonomously on ambiguous technical problems and drive them to execution.
- Strong collaboration skills and ability to work across team boundaries.
Nice-to-have
- Experience with Apache Kafka or other messaging and streaming technologies.
- Experience with platform engineering, SRE, or infrastructure-focused backend engineering.
- Familiarity with multi-account AWS environments and large-scale cloud operations.
- Experience with high-throughput event-driven architectures.
- Experience balancing reliability, performance, and cost in production systems.
OUR BENEFITS
- Total compensation includes base salary, RSUs and benefits. Base salary range: $190.000 - $220.000
- Health Insurance
- Life Insurance
- Pension Plan
- Extended maternity and paternity leaves
- Nucleo - Our learning platform of courses
- NuLanguage - Our language learning program
- NuCare - Our mental health and wellness assistance program
- Vacations
Our recruitment process may involve the use of artificial intelligence–enabled tools, such as automated interview transcription and analysis, to support the evaluation process. Artificial intelligence is not used to make final hiring decisions; all decisions are made by human reviewers.