站点可靠性工程师
Site Reliability Engineer
Moniepoint Inc. 是非洲的一站式金融平台,每月帮助 2000 万企业和个人使用无缝支付、银行、信贷、跨境和企业管理工具。
作为尼日利亚最大的商户收单机构,我们支持该国大部分的 POS(销售点)交易。通过我们的子公司,Moniepoint Inc. 每年处理超过 2500 亿美元的数字支付交易价值。
我们做什么
在 Moniepoint,我们是一个以客户为中心的团队,致力于打造重新定义行业的解决方案。我们有多种产品为 businesses 提供关键服务,例如信贷、透支等。我们利用人工智能和数据来做决策,同时也采用技术和数据驱动的最佳实践来支持我们的业务。
对 Moniepoint 是一个令人难以置信的工作场所的原因感到好奇吗?查看我们关于如何培养创新、团队合作和成长文化的文章。
职位概要
我们正在寻找一名站点可靠性工程师(SRE),负责确保我们的系统平稳高效运行,同时设计解决方案以提高可见性、消除重复任务并增强系统弹性。理想的候选人能够在实时值班责任与战略工程工作之间取得平衡,以实现可持续且可扩展的服务可靠性。
职责
- 参与值班轮班,检测并分类所有环境中的服务和可靠性问题。在重大事件中担任事件指挥官:启动战情室或会议呼叫,协调跨职能团队,向所有利益相关者提供及时清晰的状态更新。
- 创建和维护有意义的仪表盘和警报。与开发团队合作,对其代码进行监控,以确保可见性。
- 开发自动化工具,消除与可靠性相关的手动和重复操作任务(toil),涵盖应用程序和基础设施。
- 实施并跟踪由工程领导层定义的服务级别指标(SLIs)和服务级别目标(SLOs)。
- 调查并解决超出 L1 和 L2 支持范围的客户投诉,特别是涉及性能、可靠性或复杂系统行为的问题。
- 要求
- 至少 5 年的 SRE 或类似角色经验,支持企业应用,熟练掌握 Java、Go 或 Python 编程语言
- 对分布式系统概念和微服务有良好的理解
查看英文原文
Who We Are
Moniepoint Inc. is Africa’s all-in-one financial platform, helping 20 million businesses and individuals access seamless payments, banking, credit, cross-border, and business management tools each month.
As Nigeria’s largest merchant acquirer, we power most of the country’s point-of-sale (POS) transactions. Through our subsidiaries, Moniepoint Inc. processes over $250 billion in digital payment transaction value annually.
What We Do
At Moniepoint, we are a customer-focused community, dedicated to crafting solutions that redefine our industry. We have several products that provide essential services for businesses, such as credit, overdrafts, etc. We leverage artificial intelligence and data to make our decisions, but also have the technology and data-driven best practices used to support our businesses.
Curious about what makes Moniepoint an incredible place to work? Check out posts on how we cultivate a culture of innovation, teamwork, and growth.
Job Summary
We are seeking a Site Reliability Engineer (SRE) responsible for ensuring our systems run smoothly and efficiently while engineering solutions to improve visibility, eliminate repetitive tasks, and increase system resilience. The ideal candidate will balance real-time on-call responsibilities with strategic engineering work to achieve sustainable and scalable service reliability.
Responsibilities
- Participate in on-call rotations to detect and triage service and reliability issues across all environments. Act as the Incident Commander during major incidents: initiating war room or bridge calls, coordinating cross-functional teams, providing timely and clear status updates to all stakeholders.
- Create and maintain meaningful dashboards and alerts. Work with development teams to instrument their code to ensure visibility.
- Develop automation to eliminate manual and repetitive operational tasks (toil) related to reliability across both applications and infrastructure.
- Implement and track Service Level Indicators (SLIs) and Service Level Objectives (SLOs) defined by the engineering leadership.
- Investigate and resolve customer complaints escalated beyond L1 and L2 support, especially those involving performance, reliability, or complex system behavior.
Requirements
- Minimum of 5 years of experience supporting enterprise applications as an SRE or similar role with proficiency in writing code in Java, Go or Python
- Good understanding of distributed systems concepts, microservices architecture and software design patterns.
- Hands-on experience with Kubernetes. You have managed applications on a major cloud provider (GCP, AWS, or Azure), and can troubleshoot common container issues.
- Experience setting up dashboards in Grafana and using APM tools like Datadog, New Relic, Signoz.You have a Solid understanding of metrics, logs, and traces.
- Proficiency in SQL (e.g., PostgreSQL, MySQL). Ability to write complex queries to debug data issues and a basic understanding of database performance.
What we can offer you
- Culture - We put our people first and prioritize the well-being of every team member. We’ve built a company where all opinions carry weight and where all voices are heard. We value and respect each other and always look out for one another. Above all, we are human.
- Learning - We have a learning and development-focused environment with an emphasis on knowledge sharing, training, and regular internal technical talks.
- Compensation - You’ll receive an attractive salary, pension, health insurance, annual bonus, plus other benefits.
What to expect in the hiring process
- A preliminary phone call with the recruiter
- A technical interview with the Hiring Manager
- A behavioural and technical interview with a member of the Executive team.
Moniepoint is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees and candidates.