资深站点可靠性工程师
Principal Site Reliability Engineer
Deimos 是一家专注于云原生开发和安全运维技术的公司。我们帮助各类企业采用云服务以提升对客户的交付能力。我们是一支完全远程办公的非洲团队,由热爱践行工程最佳实践的工程师组成。我们利用最新技术,为客户打造具有全球竞争力的解决方案。由于 Deimos 是火星的两颗卫星之一,我们自称为“火星人”,正携手前往火星执行任务
我们的团队重视学习和适应技术变化的能力,同时欣赏扎实的基础设计和软件工程的技艺。因此,我们的工程师喜欢与有不同问题需要解决的客户合作。如果你符合这些条件,那么你将非常适合我们的环境。但请注意,你必须位于我们目前招聘的国家之一:肯尼亚、加纳、尼日利亚、南非和塞内加尔
职位概述
我们正在寻找一位经验丰富的首席站点可靠性工程师加入我们的专业服务团队,负责交付软件和 DevSecOps 项目。你将向站点可靠性工程师经理汇报工作。作为首席站点可靠性工程师,你将同时担任多个项目的技术负责人,在公司内部代表高级技术领导力
SRE / DevOps 是我们的核心能力之一。你将加入一个技术精湛的团队,持续创新并为各行业的客户提供在所有公有云(AWS、Azure、GCP 等)上的高价值解决方案。我们日常使用的工具包括 Kubernetes、Helm、Terraform、GitOps、OPA、Calico、Linkerd 等
你将负责的工作
- 设计和构建先进的云原生基础设施
- 与客户进行技术讨论并制定技术路线图
- 与工程总监合作(重新)设计架构
- 协助站点可靠性经理进行资源规划
- 协助工程经理为希望晋升为首席工程师的个人制定职业发展路径
- 教授、指导、培养其他领域专家、个人贡献者以及多个团队的成员
- 文档化流程并监控性能指标
- 引导对话以消除障碍并促进跨团队协作
- 持续提高系统的稳定性、可扩展性、安全性
查看英文原文
Deimos is a Cloud-native Developer and Security Operations technology services company. We help companies of all sizes adopt the Cloud for improved service delivery to their clients. We’re a fully remote African-based team of engineers who are passionate about implementing engineering best practices. We leverage the latest technologies while building globally competitive solutions for our clients. With Deimos being one of the two moons of Mars, we refer to ourselves as “Martians” who are on a mission to Mars, together.
Our teams value the ability to learn and adapt to technology changes while appreciating solid foundational design and the craft of software engineering. As such our engineers enjoy working with various clients who have different problems to solve. If this sounds like you then you would be an ideal fit for our environment. However, you must be based in one of the countries we currently hire in which are as follows: Kenya, Ghana, Nigeria, South Africa, and Senegal.
Role Overview
We are looking for an experienced Principal Site Reliability Engineer to join our Professional Services team and deliver Software and DevSecOps projects. You will report to a Site Reliability Engineering Manager. As a Principal Site Reliability Engineer you will be expected to fill the role of a technical lead on multiple projects simultaneously, representing the senior technical leadership within our organisation.
SRE / DevOps is one of our core competencies. You will be part of a highly-skilled team that continuously innovates and delivers high value solutions to clients across various industries on all public clouds (AWS, Azure, GCP, etc). Technologies we work with daily include Kuberenetes, Helm, Terraform, GitOps, OPA, Calico, Linkerd, just to name a few.
What you will be doing
- Design and build advanced cloud-native infrastructure
- Guide technical discussions with clients and build technical roadmaps
- Collaborate with the Engineering Director(s) to (re)design architecture
- Assist the Site Reliability Manager with resource planning
- Assist engineering managers with building career paths for individuals wishing to be promoted to Principal Engineers
- Teach, mentor, grow, and provide advice to other domain experts, individual contributors, and across several teams.
- Document processes and monitor performance metrics
- Guide conversations to remove blockers and encourage collaboration across teams.
- Constantly improve the stability, scalability, security, cost-effectiveness, and operational excellence of our clients' systems.
- Continuously discover, evaluate, and implement new technologies to maximize development efficiency and security.
- Conduct infrastructure planning, testing, and development
- Provide technical leadership on multiple projects.
What you must have
- At least 7 or more years experience working in a DevOps/SRE team
- Extensive experience in DevOps/SRE, team management and collaboration
- Advanced knowledge of best practices related to data encryption and cybersecurity
- Advanced knowledge of the general DevOps/SRE landscape, architectures, and emerging technologies
- Cloud experience, preferably GCP, Azure and AWS
- Experience in Observability Practices and Incident Management
- Extensive experience with Prometheus, Grafana, the Elastic Stack and all versions of Beats, especially within Kubernetes
- Experience with Infrastructure as Code, preferably Terraform
- Experience with general automation and config management, preferably Ansible
- Extensive experience building and maintaining Kubernetes clusters and workloads
- Strong foundation of basic network and security concepts
- Ability to build robust CICD pipelines
- Familiarity with relational and non-relational databases
- Solid understanding of Linux operating systems
Qualities & Behaviours
- Exceptional interpersonal and communication skills
- A zest for automation
- Comfortable working as a remote team member and leader
- Ability to keep up to date with DevOps/SRE best practices, trends and innovation
- Passionate about mentoring and growing technical skills within the team
About you
For us to achieve our ambitious vision together as a team, It is important for our Martians to lead at all levels, be self starters who take initiative and put their hands up for challenging tasks. A growth mindset is important to us and we encourage all our Martians to openly share knowledge, support and help each other, ask questions, get creative with new technologies and learn from setbacks.
Becoming a Martian means:
- Comfortably working and learning from a fully remote, culturally diverse team based predominantly in South Africa, Kenya, Nigeria and Ghana.
- Being an open, honest and respectful communicator.
- You enjoy asking questions, identifying areas of improvement and proposing solutions, no matter your job title or whether you have been with us for a day, a month or years!
- You are comfortable taking initiative and operating independently.
- You thrive in a fast paced environment, where change is constant.
- You find it exciting to work with various clients, from different industries, each with a different problem for you and your team to solve.
- Intentionally sharing tech and industry trends that excite you with your peers.
- Seeking continuous feedback and actively taking steps to continuously grow personally and professionally.
Want to know what you get by joining us?
- Become a member of a team where we value each individual's contribution from day 1 and empower you to make suggestions, get involved and do what you love most!
- Flexibility and the freedom to work remotely.
- Work-life balance where you are not expected to work over weekends or after hours.
- A forward thinking remote company that knows how important it is to stay connected as one team, by providing virtual social platforms for employee engagement.
- A monthly work from home allowance which you can use to set yourself up to work comfortably from home. Whether that is pens, notebooks, new headphones or work snacks!
- A MacBook or Windows laptop for you to do your best work on.
- Become part of a team of exceptionally clever and talented people who like to share their knowledge and learnings.
- We support your career growth and love to celebrate your successes and advancement!
Originally posted on Himalayas