Pipeline工程师 (全球远程,可 anywhere)
Pipeline Engineer (worldwide remote, work anywhere)
CloudLinux 是一家全球远程优先的公司。我们以我们的原则为指导:做正确的事,员工第一,我们是远程优先,我们提供高吞吐量、低成本的 Linux 基础设施和安全产品,帮助公司提高运营效率。我们团队中的每个人都互相支持,尽我们所能确保大家的成功。了解更多请访问我们的网站 https://cloudlinux.com/
Imunify360 安全套件是 CloudLinux Inc. 的产品,该公司是托管提供商最安全、最稳定的 #1 操作系统的制造商。Imunify 是一款专为共享服务器和 VPS/独立服务器设计的创新安全解决方案。该自动化、易于使用的产品采用六层安全方法,提供全面且完整的攻击防护。
我们正在寻找一位经验丰富的工程师,负责将威胁情报转化为数千万网站的保护措施的自动化流水线。当出现新漏洞时,一系列自动化系统必须能够发现它,获取受影响的源代码,生成 WAF 规则,生成测试,证明规则有效且不会破坏合法流量,并在数千万网站上部署——然后足够关注生产环境,如果出现问题可以自动回滚。目前这条链已经存在并全天候运行,一个不可靠的环节可能会导致客户失去保护。我们正在寻找一位强大的工程师,快速扩展它同时保持其可靠性和透明度。这涉及解决困难的工程问题,在大规模数据中工作,面对各种意外情况。
这不是一个分析师或研究岗位,但在此领域有额外经验是欢迎的。我们需要的是一个能构建持续工作的系统的人:内部有复杂的逻辑,外部则枯燥但可靠。
该职位为完全远程,工作时间灵活,允许你规划每天的工作,并从世界任何地方工作。
你将负责:
目前投入生产的系统,需要大幅扩展:
自动化保护流水线 —— 从威胁情报到验证并部署的规则的链条。多阶段、高度自主,并且每天必须在固定的时间窗口内完成。
渐进式发布自动化 —— 在受控阶段将保护措施部署到整个集群,具备自动防护机制,可以在无人参与的情况下暂停或回滚某个阶段。
质量门禁 —— 根据实时生产信号决定是否……
查看英文原文
CloudLinux is a global remote-first company. We are driven by our principles: do the right thing, employees first, we are remote first, and we deliver high-volume, low-cost Linux infrastructure and security products that help companies to increase the efficiency of their operations. Every person on our team supports each other and does what we can to ensure we all are successful. Check out our website for more information https://cloudlinux.com/Imunify360 Security Suite is a product of CloudLinux Inc., the maker of the #1 OS in security and stability for hosting providers. Imunify is an innovative security solution designed specifically for shared and VPS/Dedicated servers. The automated, easy-to-use solution with the six-layer approach to security delivers comprehensive and complete attack prevention.We are looking for an experienced engineer to own the automated pipelines that turn threat intelligence into shipped protection for tens of millions of websites.When a new vulnerability appears, a chain of automated systems has to notice it, obtain the vulnerable source, produce WAF rules, generate tests, prove the rules work and don't break legitimate traffic, and roll it out across tens of millions of websites — then watch production hard enough to pull it back automatically if it misbehaves. Today that chain exists and runs 24/7, and a single unreliable link could cost customers their protection. We are looking for a strong Engineer to expand it rapidly while keeping it reliable and transparent. This involves solving hard engineering problems, working with data at scale, in a world with many surprises.This is not an analyst role and not a research role, but additional experience in this domain is welcome. What we need is someone who builds systems that keep working: complex logic underneath, boring and reliable on the outside.The position is fully remote with flexible hours, allowing you to plan your day and work from anywhere in the world.What you would ownSystems that are in production today and need to grow considerably:The automated protection pipeline — the chain from threat intelligence to a validated, deployed rule. Multi-stage, largely autonomous, and required to finish inside a fixed time window every day.Progressive release automation — rolling protection out across the fleet in controlled stages, with automated guardrails that hold or roll back a stage without a human in the loop.Quality gates — deciding, from live production signal, whether something we shipped is doing harm, and acting on that before it reaches the next stage. Errors are costly in both directions: miss a problem and customer sites break; over-correct and protection is silently removed.CI at scale — validation that stands up real, disposable environments across a large matrix of software versions and configurations, and returns a trustworthy verdict fast enough to stay inside the release window.LLM orchestration and cost control — several subsystems are driven by AI agents inside purpose-built harnesses, with evaluation, budgets and spend accounting as load-bearing components.Observability and alerting across all of it — the pipelines are expected to report their own condition, prove they are healthy, and escalate on their own.This runs on top of a petabyte-scale threat-intelligence store and live telemetry from more than 60 million websites, under strict end-to-end latency budgets measured in hours. The failure mode we care most about eliminating is a budget missed silently.We will go into the specifics during the interview process.Key responsibilitiesDesigning, building and operating the automated pipelines described above, end to end;Turning fragile multi-stage batch jobs into resumable, idempotent, observable systems with explicit state machines and recovery paths;Defining and enforcing latency budgets and SLOs per stage, and making violations visible and actionable rather than silent;Building the observability layer — metrics, dashboards, alerting and health gates — so the pipeline reports its own condition instead of needing someone to go and look;Designing and implementing guardrails: automatic hold and rollback, blast-radius limits, kill switches, and safe-by-default behaviour when an upstream dependency is unavailable;Making the systems low-maintenance: eliminating manual steps, removing standing human babysitting, and reducing the operational surface rather than adding to it;Writing and maintaining unit and integration tests for logic that is genuinely hard to test — concurrency, partial failure, external API flakiness, multi-stage state;Investigating and resolving complex issues across ClickHouse, GitLab CI, S3/object storage, Prometheus/Grafana and third-party APIs;Collaborating with the security analysts and the Server team on architecture, and pushing back when a proposed design will not survive contact with production.Requirements5+ years of professional backend / platform / infrastructure engineering experience;Demonstrable experience building and operating multi-stage data or automation pipelines — CI/CD systems, ETL/ELT, build and release automation, job orchestration, ML/data platforms, or similar. This is the single most important requirement. We will ask you to walk us through one in detail;Real depth in at least one of Python, Go or Rust. We use all three, and we are not hiring a language specialist — which one you bring genuinely does not matter. Depth is simply how we verify the experience behind it is real, so expect specific questions about systems you have designed and shipped: why they are shaped the way they are, how they behave under failure, and what you would build differently today;Systems design judgement, more than raw coding throughput. The difficult part of this role is deciding what to build, working out where it will break, and making it prove its own correctness — not volume of code produced;Practical experience with workflow orchestration and job scheduling (Airflow, Temporal, Prefect, Dagster, Argo, custom schedulers — whatever you have actually run in production);A working instinct for reliability engineering: idempotency, retries with backoff, exactly-once vs at-least-once, checkpointing and resumability, graceful degradation, backpressure, and safe handling of partial failure;Hands-on observability experience — Prometheus/Grafana, LGTM stack, or equivalent — including designing metrics rather than only consuming dashboards someone else built;Deep CI/CD experience, ideally GitLab CI including dynamic/child pipelines and self-hosted runners; comfort with Docker and container-based test environments;Experience with object storage (S3/Ceph or equivalent) and with large-scale analytical stores — ClickHouse or another columnar database;Comfort designing state machines and long-running processes that survive restarts, and reasoning about concurrency across multiple in-flight rollouts;Excellent debugging skills across system, network and data layers;Strong communication skills and comfort working in a distributed team;Proficiency in spoken and written English.Nice to haveExperience with progressive delivery — canary and percentage-based rollouts, feature flags, automated rollback, blast-radius control;Experience running AI/LLM systems in production, particularly cost control, token accounting, evaluation harnesses, and dealing with non-deterministic components inside a deterministic pipeline;Experience with fleet-scale telemetry and with building quality gates on top of noisy production signal;Familiarity with WordPress, PHP, or WAF/ModSecurity concepts;Experience with configuration management (Ansible, Puppet, Salt) and with Linux service operations.You do not need a cybersecurity background. Most of the hard problems here are orchestration, reliability, correctness under concurrency and observability. Domain knowledge is learnable and we have specialists to learn it from — pipeline engineering judgement is what we cannot substitute.We value engineers who areCurious and fearless problem solvers — not afraid to dig into existing systems, investigate root causes, and propose improvements;Sceptical by default — who ask what would falsify a conclusion before acting on it, and who trust measurements over plausible reasoning;Pragmatic and detail-oriented — focused on building reliable, maintainable systems, and allergic to solutions that require a human to remember something;Owners — comfortable being the person accountable for whether a pipeline ran correctly last night;Effective communicators — able to articulate ideas clearly, exchange feedback constructively, and foster collaboration across teams;Engaging and proactive — contributing energy, initiative, and a positive presence that strengthens team culture.BenefitsWhat's in it for you?A focus on professional development.Interesting and challenging projects.Fully remote work with flexible working hours, that allows you to schedule your day and work from any location worldwide.Paid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leaves.Compensation for private medical insurance.Co-working and gym/sports reimbursement.Budget for education.The opportunity to receive a reward for the most innovative idea that the company can patent.By applying for this position, you consent to the processing of your personal data as described in our Privacy Policy (https://cloudlinux.com/candidate-privacy-notice), which provides detailed information on how we maintain and handle your data.