L3 Linux/Kubernetes支持/专家
L3 Linux/Kubernetes support/specialist
负责我们基于云的服务和底层基础设施的性能、可用性和可靠性,作为关键的技术专家。
主要职责包括但不限于:
· 管理、排查和优化部署在 Kubernetes、Red Hat OpenShift 和 OpenStack 平台上的容器化应用和基础设施。
- 作为核心云基础设施技术的专家(SME),包括高级 Linux(CentOS)系统管理、Docker/容器和复杂网络配置。
- 领导调查和解决复杂的高严重性客户问题,运用扎实的分析知识快速诊断整个云栈中的问题。
- 利用专业知识快速识别根本原因,并为客户事件实施有效且持久的解决方案。
- 准备并执行关键事件的彻底根本原因分析(RCA),以识别系统性问题并防止再次发生。
- 使用 Python 和 Ansible 开发、测试和维护强大的自动化脚本,以简化日常运维任务并提高整体服务效率。
- 识别并实现自动化机会,减少维护和部署活动中的手动工作量。
- 提供端到端的升级、监控和紧急(EME)支持,作为最终升级点以确保服务可用性并满足 SLA。
- 直接与客户团队和内部团队沟通,了解需求并提供定制化的技术解决方案。
- 跟踪云计算和容器化领域的行业最佳实践和新兴技术。
排班:
· 我们的团队采用跟随太阳的模式,确保全天候覆盖和快速响应时间。虽然这可能偶尔需要在非典型本地工作时间进行轮班,但总工作时长将保持与标准工作周一致。
- 轮班值班是该职位的强制性部分。值班专家需负责为紧急情况提供即时支持,包括直接与客户联系(通过电话/Teams)以及及时参与指定的作战室进行事件调查和解决。
- 所有干预操作均在生产环境中进行,必须严格遵循标准网络接触政策,在紧急事件处理期间需全程保持意识和合规性。
所需技能包括但不限于:
·
Linux
查看英文原文
Responsible for the performance, availability, and reliability of our cloud-based services and underlying infrastructure, acting as a critical technical subject matter expert.
Main tasks include, but are not limited to:
· Manage, troubleshoot, and optimise containerised applications and infrastructure deployed on Kubernetes, Red Hat OpenShift, and OpenStack platforms.
- Serve as the Subject Matter Expert (SME) for core cloud infrastructure technologies, including advanced Linux (CentOS) system administration, Docker/Containers, and complex networking configurations.
- Lead the investigation and resolution of complex, high-severity customer issues, applying strong analytical knowledge to quickly diagnose problems across the entire cloud stack.
- Utilise your expertise to quickly identify root causes and implement effective, durable solutions for customer incidents.
- Prepare and conduct rigorous Root Cause Analysis (RCA) for critical incidents to identify systemic issues and prevent recurrence.
- Develop, test, and maintain robust automation scripts using Python and Ansible to streamline daily operational tasks and improve overall service efficiency.
- Identify and implement automation opportunities to reduce manual effort in maintenance and deployment activities.
- Provide end-to-end Escalation, Monitoring, and Emergency (EME) support, acting as a final escalation point to ensure service availability and meet SLAs.
- Liaise directly with the customer team and internal teams to understand requirements and deliver tailored technical solutions.
- Stay current with industry best practices and emerging technologies in cloud and containerization.
Scheduling:
- Our team operates on a follow-the-sun model, ensuring 24/7 coverage and rapid response times. While this may occasionally require shifts outside typical local business hours, the total work hours will remain consistent with a standard work week.
- A rotational on-call schedule is a mandatory part of this position. On-call experts are responsible for providing immediate support for urgent cases, including direct customer contact (via call/Teams) and prompt engagement in designated war rooms for case investigation and resolution.
- All interventions are performed in live environments and must strictly follow the Standard Network Touch Policy, requiring full awareness and compliance during urgent case resolution.
Skill set required to cover these services includes, but is not limited to:
- Linux Expertise: Strong knowledge and proven hands-on experience with Linux administration.
- Networking Foundations: Strong knowledge of core networking principles (TCP/IP, routing, load balancing, firewalls) in a cloud environment.
- Containerization & Virtualisation: Strong knowledge of Kubernetes orchestration, OpenStack platforms, and Docker/Containerization.
- Problem-Solving Mindset: Possess sharp troubleshooting skills combined with an analytical mindset to dissect and address complex challenges.
- Scripting and Automation: Solid Python scripting skills for task automation and system management.
- Configuration Management:Hands-on experience with Ansible for configuration management.’
- Root Cause Analysis (RCA):Expertise in preparation and implementation of RCAs.
- Escalation and Monitoring:Proven experience with EME (Escalation, Monitoring, and Emergency) management processes.
One or more certifications from the list below will be considered an added advantage:
Red Hat Certified Specialist in Cloud Infrastructure (EX210), Red Hat Certified Engineer (RHCE) in Red Hat OpenStack (EX310), RHCSA, RHCE, CKA, EX280 (Red Hat Certified Specialist in OpenShift Administration), EX380 (Red Hat Certified Specialist in OpenShift Automation and API Management).
Communication & Attitude:
· Excellent communication skills (written and verbal) and the ability to articulate complex issues clearly.
· Demonstrated ability to handle customer pressure and manage expectations with a positive, professional attitude.
Strong team player, self-motivated, and capable of excelling in a high-pressure, dynamic support environment.
Originally posted on Himalayas