高级站点可靠性工程师(远程)
Senior Site Reliability Engineer (Remote)
职位描述
地点:完全远程,欧洲时区(CET ±2小时)
开始日期:尽快
语言:流利的英语是必须的
行业:云计算
我们在Pragmatike招聘,以扩展团队并推动我们内部项目的增长。
我们的重点是开发云计算领域的前沿解决方案,同时培养协作和创新的文化。加入我们意味着成为一支充满热情的团队的一员,你的想法和技能将直接有助于塑造未来的技术。
如果你对在一个动态且灵活的环境中参与有雄心的项目感到兴奋,我们很期待收到你的消息!
职责
- 运行和维护基于Linux的基础设施(Debian/Ubuntu)。
- 在裸机、虚拟化和本地环境中部署、管理和扩展Kubernetes集群。
- 管理整个集群生命周期:升级、节点池、网络、存储和安全加固。
- 使用Ansible、Bash/Python和GitOps工作流程实现自动化部署和运维。
- 设计和维护网络架构,包括VLAN、L2/L3路由、VPN和多站点连接。
- 构建自动化部署流程(PXE引导、Preseed、cloud-init)。
- 部署和维护可观测性堆栈(Prometheus/Grafana、Loki、ELK、Graylog)。
- 领导跨平台的事件响应和升级活动。
- 提高系统的可用性并降低各层级的延迟。
- 在多个基础设施层级(物理网络/硬件、平台虚拟化、软件服务)定义和实施SLO/SLI。
- 优化警报和监控流程,提供可操作的见解。
- 建立并维护轮班排班,确保覆盖不同时区。
- 制定标准操作程序(SOP),用于重复性的运维任务。
- 协调Policlouds的物理维护工作(定期维护、硬件问题、数据中心运维)。
- 管理虚拟化和编排层(OpenStack、Proxmox、VMware)。
- 协助开发和维护所有产品的整体架构。
- 为未来计划规划资源,考虑需求和增长预测。
- 与开发团队合作,提高整体质量和优化资源利用率。
- 与跨职能利益相关者(Hivenet、Policloud、客户成功团队)合作。
要求
- 在生产环境中具有高级的、实际操作的Kubernetes经验。
- 强大的网络知识,熟悉TCP/IP、路由协议、防火墙和网络安全。
查看英文原文
Job Description
Location: Fully remote EU timezone (CET ±2h)
Start date: ASAP
Languages: Fluent English is mandatory
Industry: Cloud Computing
We are hiring at Pragmatike to expand our team and drive the growth of our internal projects.
Our focus is on developing cutting-edge solutions in Cloud Computing, while fostering a culture of collaboration and innovation. Joining us means being part of a passionate team where your ideas and skills directly contribute to shaping tomorrows technologies.
If you're excited about working on ambitious projects in a dynamic and flexible environment, we'd love to hear from you!
RESPONSIBILITIES
- Operate and maintain Linux-based infrastructure (Debian/Ubuntu).
- Deploy, manage, and scale Kubernetes clusters across bare-metal, virtualized, and on-prem environments.
- Oversee full cluster lifecycle: upgrades, node pools, networking, storage, and security hardening.
- Implement automation for provisioning and operations using Ansible, Bash/Python, and GitOps workflows.
- Design and maintain networking architecture including VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
- Build automated deployment workflows (PXE boot, Preseed, cloud-init).
- Deploy and maintain observability stacks (Prometheus/Grafana, Loki, ELK, Graylog).
- Lead incident response and escalation activities across the platform.
- Improve system availability and reduce latency at all levels.
- Define and implement SLOs/SLIs at multiple infrastructure levels (physical network/hardware, platform virtualization, software services).
- Optimize alerting and monitoring pipelines to provide actionable insights.
- Establish and maintain on-call schedules to ensure coverage across timezones.
- Develop Standard Operating Procedures (SOPs) for repeatable operations and maintenance tasks.
- Coordinate physical maintenance for Policlouds (periodic maintenance, hardware issues, DC-Ops).
- Manage virtualization and orchestration layers (OpenStack, Proxmox, VMware).
- Help develop and maintain overall architecture across all products.
- Plan resources for future initiatives, accounting for demand and growth projections.
- Work with development teams to improve overall quality and optimize resource utilization.
- Collaborate with cross-functional stakeholders (Hivenet, Policloud, Customer Success teams).
REQUIREMENTS
- Expert-level, hands-on experience operating Kubernetes in production environments.
- Strong network engineering skills (VLANs, L2/L3 routing, VPNs, multi-site connectivity) - this is essential for the role.
- Strong proficiency with Linux systems administration (Debian/Ubuntu).
- Solid understanding of networking fundamentals and ability to design complex network architectures.
- Experience building and maintaining automation workflows (Ansible, Bash/Python, Git-based).
- Experience with observability stacks such as Prometheus, Grafana, ELK, Loki, or Graylog.
- Background with virtualization technologies (OpenStack, Proxmox, VMware).
- Experience with bare-metal provisioning and MAAS (Metal as a Service).
- Strong understanding of distributed systems and container orchestration.
- Process-oriented mindset with ability to develop SOPs and operational procedures from scratch.
- Experience with incident response, escalation procedures, and on-call rotations.
- Ability to work autonomously in a fast-paced, engineering-driven environment.
- Strong technical skills combined with alignment to team values.
NICE TO HAVE
- Experience with service mesh (Istio, Linkerd) or advanced CNI implementations.
- Knowledge of Cloudflare APIs, DNS automation, or tunnel configurations.
- Experience with GPU infrastructure, node preparation, or resource scheduling.
- Familiarity with security best practices (RBAC, firewalls, network policies).
- Exposure to IT asset management or license tracking workflows.
- Experience working in multi-timezone environments and coordinating across distributed teams.
- Background establishing reliability practices and SRE frameworks in growing organizations.
Why Join Us:
- 100% remote work with flexible hours
- High-impact role with autonomy and ownership
- Collaborative and international engineering team
- Cutting-edge tech stack with strong focus on reliability and automation.