远程工作雷达

高级站点可靠性工程师(远程)

Senior Site Reliability Engineer (Remote)

开发工程全球可投(据职位描述推断)
公司pragmatike
薪资未公开
工作地点Armenia / Latvia / Spain / Albania / Bosnia & Herzegovina / Estonia / Poland / Portugal / Italy / Romania
地域资格全球可投(据职位描述推断)
时区要求无特别要求
用工类型FullTime
发布时间9 天前
数据来源Ashby
前往企业招聘页投递 →
全球可投:该职位未限制候选人所在地区。仍需注意薪资可能按地区折算,以及实际签约方式(正式雇佣 / 独立合同)。

职位描述

地点:完全远程,欧洲时区(CET ±2小时)
开始日期:尽快
语言:流利的英语是必须的
行业:云计算

我们在Pragmatike招聘,以扩展团队并推动我们内部项目的增长。

我们的重点是开发云计算领域的前沿解决方案,同时培养协作和创新的文化。加入我们意味着成为一支充满热情的团队的一员,你的想法和技能将直接有助于塑造未来的技术。

如果你对在一个动态且灵活的环境中参与有雄心的项目感到兴奋,我们很期待收到你的消息!

职责

- 运行和维护基于Linux的基础设施(Debian/Ubuntu)。

- 在裸机、虚拟化和本地环境中部署、管理和扩展Kubernetes集群。

- 管理整个集群生命周期:升级、节点池、网络、存储和安全加固。

- 使用Ansible、Bash/Python和GitOps工作流程实现自动化部署和运维。

- 设计和维护网络架构,包括VLAN、L2/L3路由、VPN和多站点连接。

- 构建自动化部署流程(PXE引导、Preseed、cloud-init)。

- 部署和维护可观测性堆栈(Prometheus/Grafana、Loki、ELK、Graylog)。

- 领导跨平台的事件响应和升级活动。

- 提高系统的可用性并降低各层级的延迟。

- 在多个基础设施层级(物理网络/硬件、平台虚拟化、软件服务)定义和实施SLO/SLI。

- 优化警报和监控流程,提供可操作的见解。

- 建立并维护轮班排班,确保覆盖不同时区。

- 制定标准操作程序(SOP),用于重复性的运维任务。

- 协调Policlouds的物理维护工作(定期维护、硬件问题、数据中心运维)。

- 管理虚拟化和编排层(OpenStack、Proxmox、VMware)。

- 协助开发和维护所有产品的整体架构。

- 为未来计划规划资源,考虑需求和增长预测。

- 与开发团队合作,提高整体质量和优化资源利用率。

- 与跨职能利益相关者(Hivenet、Policloud、客户成功团队)合作。

要求

- 在生产环境中具有高级的、实际操作的Kubernetes经验。

- 强大的网络知识,熟悉TCP/IP、路由协议、防火墙和网络安全。

查看英文原文

Job Description

Location: Fully remote EU timezone (CET ±2h)
Start date: ASAP
Languages: Fluent English is mandatory
Industry: Cloud Computing

We are hiring at Pragmatike to expand our team and drive the growth of our internal projects.

Our focus is on developing cutting-edge solutions in Cloud Computing, while fostering a culture of collaboration and innovation. Joining us means being part of a passionate team where your ideas and skills directly contribute to shaping tomorrows technologies.

If you're excited about working on ambitious projects in a dynamic and flexible environment, we'd love to hear from you!

RESPONSIBILITIES

- Operate and maintain Linux-based infrastructure (Debian/Ubuntu).

- Deploy, manage, and scale Kubernetes clusters across bare-metal, virtualized, and on-prem environments.

- Oversee full cluster lifecycle: upgrades, node pools, networking, storage, and security hardening.

- Implement automation for provisioning and operations using Ansible, Bash/Python, and GitOps workflows.

- Design and maintain networking architecture including VLANs, L2/L3 routing, VPNs, and multi-site connectivity.

- Build automated deployment workflows (PXE boot, Preseed, cloud-init).

- Deploy and maintain observability stacks (Prometheus/Grafana, Loki, ELK, Graylog).

- Lead incident response and escalation activities across the platform.

- Improve system availability and reduce latency at all levels.

- Define and implement SLOs/SLIs at multiple infrastructure levels (physical network/hardware, platform virtualization, software services).

- Optimize alerting and monitoring pipelines to provide actionable insights.

- Establish and maintain on-call schedules to ensure coverage across timezones.

- Develop Standard Operating Procedures (SOPs) for repeatable operations and maintenance tasks.

- Coordinate physical maintenance for Policlouds (periodic maintenance, hardware issues, DC-Ops).

- Manage virtualization and orchestration layers (OpenStack, Proxmox, VMware).

- Help develop and maintain overall architecture across all products.

- Plan resources for future initiatives, accounting for demand and growth projections.

- Work with development teams to improve overall quality and optimize resource utilization.

- Collaborate with cross-functional stakeholders (Hivenet, Policloud, Customer Success teams).

REQUIREMENTS

- Expert-level, hands-on experience operating Kubernetes in production environments.

- Strong network engineering skills (VLANs, L2/L3 routing, VPNs, multi-site connectivity) - this is essential for the role.

- Strong proficiency with Linux systems administration (Debian/Ubuntu).

- Solid understanding of networking fundamentals and ability to design complex network architectures.

- Experience building and maintaining automation workflows (Ansible, Bash/Python, Git-based).

- Experience with observability stacks such as Prometheus, Grafana, ELK, Loki, or Graylog.

- Background with virtualization technologies (OpenStack, Proxmox, VMware).

- Experience with bare-metal provisioning and MAAS (Metal as a Service).

- Strong understanding of distributed systems and container orchestration.

- Process-oriented mindset with ability to develop SOPs and operational procedures from scratch.

- Experience with incident response, escalation procedures, and on-call rotations.

- Ability to work autonomously in a fast-paced, engineering-driven environment.

- Strong technical skills combined with alignment to team values.

NICE TO HAVE

- Experience with service mesh (Istio, Linkerd) or advanced CNI implementations.

- Knowledge of Cloudflare APIs, DNS automation, or tunnel configurations.

- Experience with GPU infrastructure, node preparation, or resource scheduling.

- Familiarity with security best practices (RBAC, firewalls, network policies).

- Exposure to IT asset management or license tracking workflows.

- Experience working in multi-timezone environments and coordinating across distributed teams.

- Background establishing reliability practices and SRE frameworks in growing organizations.

Why Join Us:

- 100% remote work with flexible hours

- High-impact role with autonomy and ownership

- Collaborative and international engineering team

- Cutting-edge tech stack with strong focus on reliability and automation.

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位