平台工程师(远程工作)
Platform Engineer (remote work)
CloudLinux 构建 Linux 基础设施和安全产品。你将加入我们的自动化与管理服务小组,与 Dmitrii Petrov 密切合作,解决跨团队和服务的问题:云成本数据、基础设施库存、网络策略和容量流程。查看我们的网站获取更多信息 https://cloudlinux.com/
在基础设施部门,平台小组是一个小团队。我们运行可观测性平台、公司的 GitLab 以及其背后的 CI 运行器,一些较小的工程服务,以及部门用于配置和部署的自动化工具。
我们正在寻找一名平台工程师加入 PaaS 团队,负责一组已达成一致的平台服务:保持它们的可靠性并符合约定的服务水平,使更改和恢复变得可重复,并通过自动化和自助服务减少重复性运维工作。
你可以自由地实现各种方案,同时对结果负责:你选择方法,在评审中进行辩护,并为服务后续的行为负责。
大部分工作涉及站点可靠性工程:每天稳定运行的服务,以及产品团队请求的正确处理。其中一部分是构建工作。可观测性平台和运行器集群都是在过去一年内从零开始构建的,未来还会有更多类似的工作。期望工作内容在平台工程、事件响应和帮助其他团队使用我们提供的服务之间分配。
你将负责:
运行可观测性平台。保持其健康状态,对接团队,监控成本和容量,并维护在其之上的告警系统。
运行 GitLab 和 CI 运行器集群。包括升级、容量管理、访问权限、备份和恢复演练。
保持我们其余服务的健康状态,包括生产服务所需的监控和操作手册。
在有需求时部署新服务。研究选项,选择设计,并根据最佳实践从零开始搭建服务:代码化、监控、备份、文档齐全。
处理开发人员的请求。包括访问权限、对接、流水线问题、新的导出器和仪表盘。回应这些请求,并将重复性的转化为自助服务。
处理事件。诊断并减轻影响,安全地恢复服务,然后完成根本原因分析和事后总结。提供事件所需的预防或检测改进。
所有内容都以代码形式交付,通过合并请求进行评审。每次更改前都要计划并检查。
为工程师撰写文档
查看英文原文
CloudLinux builds Linux infrastructure and security products. You will join our Automation & Management Services cell, working closely with Dmitrii Petrov to solve problems across teams and services: cloud-cost data, infrastructure inventory, network policies and capacity workflows.Check out our website for more information https://cloudlinux.com/Inside the Infrastructure Department, the Platform cell is a small team. We run the observability platform, the company's GitLab and the CI runners behind it, a few smaller engineering services, and the automation the department relies on for provisioning and configuration.We are looking for a Platform Engineer for the PaaS team to take ownership of an agreed set of platform services: keep them reliable and within agreed service levels, make changes and recovery repeatable, and reduce recurring operational work through automation and self-service.You get real freedom in how you implement things, and you own the result: you pick the approach, defend it in review, and answer for how the service behaves afterwards.Most of the job is what site reliability engineering is about: services that run well day after day, and requests from product teams handled properly. Some of it is building. The observability platform and the runner cluster were both built from scratch within the last year, and there will be more of that. Expect the work to split between platform engineering, incident response, and helping engineers in other teams use what we run.What you'll doRun the observability platform. Keep it healthy, onboard teams, watch cost and capacity, and maintain the alerting that runs on top of it.Run GitLab and the CI runner fleet. Upgrades, capacity, access, backups and restore drills.Keep the rest of our services healthy, with the monitoring and runbooks a production service needs.Deploy new services when they are requested. Research the options, pick a design, and stand the service up from scratch according to good practice: as code, monitored, backed up, documented.Work with developers' requests. Access, onboarding, pipeline problems, new exporters and dashboards. Answer them, and turn the recurring ones into self-service.Run incidents. Diagnose and mitigate impact, restore service safely, then complete the root-cause analysis and post-mortem. Deliver the prevention or detection improvements the incident calls for.Ship everything as code, reviewed in merge requests. Plan and check before every change.Write for engineers outside the team. Runbooks, onboarding guides, maintenance notices and status updates that people can act on.Work with AI agents. Delegate collection and drafting to them, review their output as you would a colleague's merge request, and record what you learn where the team can find it.RequirementsMust haveSenior-level experience in infrastructure, platform or site reliability engineering, including at least one production service you were responsible for keeping up. We will ask you to walk us through it in detail: what broke, how you found out, and what you changed so it would not happen again.Linux systems administration and debugging on bare metal and virtual machines. Much of our infrastructure is not Kubernetes.Kubernetes in production delivered through GitOps, including cluster upgrades you performed yourself.Infrastructure as code as your delivery form: Ansible and Terraform or OpenTofu, changes reviewed in merge requests.GitLab administration and GitLab CI in production, self-hosted or SaaS. Deep experience with another CI system is acceptable if you can show the same depth.Working knowledge of the Prometheus and Grafana ecosystem: you have run it for a team, written alert rules and dashboards, and can read PromQL. Depth here is welcome, and learnable.Written technical explanation for engineers outside your team: runbooks, notices, answers to requests.Strong communication and interpersonal skills. This role deals with people at least as much as with servers: most work starts as a conversation with a product team, and you need to understand what they actually need, agree scope, priority and timing with them, push back politely when a request should not be done as asked, and keep everyone informed while the work is in progress. We are looking for someone other teams enjoy working with.Advanced use of AI engineering assistants such as Claude and Codex: providing context, breaking down tasks, designing agent loops, and delegating plans for unattended, end-to-end execution within defined scope and permissions, with clear stop conditions. You can explain, debug and test the resulting automation, and verify generated commands, scripts and conclusions before they touch production.English - upper-intermediate or higher - to ensure clear communication of progress within the teams.Nice to haveAlerting design: SLOs, burn-rate alerts, thresholds sized from data.MicroVM isolation for CI: Kata Containers, Firecracker or gVisor.S3-compatible object storage operations: Ceph RGW or similar.AWS with real cost work.Self-hosted Sentry, or another Kafka, ClickHouse and Redis-backed application you have kept alive under load.Python or Go for exporters and small internal services.You do not need to have run every system on this list. Solid fundamentals and the judgment to pick up an unfamiliar service, make it observable and hand back a runbook matter more than matching every line.What this role is not focused onNot a ticket-queue operator. Recurring requests get turned into self-service, not processed one by one forever.Not a pure cloud or Kubernetes role. Bare metal and virtual machines are a large part of our infrastructure.Not the DBA, the network engineer or the security engineer. Those teams run their own systems; we provide the platform they monitor them with.BenefitsWhat's in it for you?A focus on professional development.Interesting and challenging projects.Fully remote work with flexible working hours, which allows you to schedule your day and work from any location worldwide.Paid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leaves.Compensation for private medical insurance.Co-working and gym/sports reimbursement.Budget for education.The opportunity to receive a reward for the most innovative idea that the company can patent.By applying for this position, you consent to the processing of your personal data as described in our Privacy Policy (https://cloudlinux.com/candidate-privacy-notice), which provides detailed information on how we maintain and handle your data.