高级自动化工程师(远程工作)
Senior Automation Engineer (remote work)
CloudLinux 构建 Linux 基础设施和安全产品。你将加入我们的自动化与管理服务团队,与 Dmitrii Petrov 密切合作,解决跨团队和跨服务的问题:云成本数据、基础设施库存、网络策略和容量流程。请访问我们的网站了解更多信息 https://cloudlinux.com/ 我们正在寻找一名高级自动化工程师加入基础设施团队。你将独立负责达成的成果——从与工程团队、财务或安全团队明确需求,到交付和运营一个有用的系统。工作内容包括新开发、改进、修复和维护。没有传统的值班轮班;你在工作时间内响应自己系统的警报。
你将做什么
将不完整的任务说明转化为达成一致的结果和可行的计划。跟踪跨团队依赖关系,解释权衡,尽早提出障碍并沟通估算变化。与用户确认完成情况。
构建可维护的软件和云/基础设施 API 集成。处理凭证、分页、重试、部分失败和重复执行,确保无人值守的工作流保持安全可靠。
将计费、使用和库存数据转化为可重复的输出。核对总数和周期,应用已定的分配规则,并使缺失或不确定的映射可见。
使用现有系统扩展库存、网络策略和容量自动化。保留必要的人员决策并验证操作结果。
通过版本控制的代码和 CI 交付变更。理解实时状态,测试恢复并安全发布;诊断代码、Linux、API、数据和网络中的故障。
负责监控、可操作的警报、文档和维护。让其他工程师能够操作系统,在使用编码代理的同时仍对设计、代码、测试和结果负责。
要求
必须具备
高级经验,通常在软件、基础设施、平台或自动化工程领域有 5 年以上经验。你能解释自己的设计决策、实现和操作结果。
强大的生产环境 Python 编程、软件设计、测试、调试和 API 集成技能。你能理解陌生代码并改进现有系统。
实际的 AWS 经验和扎实的 Linux/网络基础,包括跨主机、服务、网络和云边界的诊断能力。
实用的 SQL 和数据验证技能:连接、单位、周期、所有权映射和对账。你能区分...
查看英文原文
CloudLinux builds Linux infrastructure and security products. You will join our Automation & Management Services cell, working closely with Dmitrii Petrov to solve problems across teams and services: cloud-cost data, infrastructure inventory, network policies and capacity workflows.Check out our website for more information https://cloudlinux.com/We are looking for a Senior Automation Engineer to join the Infrastructure team. You will independently own agreed outcomes—from clarifying the need with engineering teams, Finance or Security to delivering and operating a useful system. The work includes new development, improvements, repairs and maintenance. There is no traditional on-call rotation; you respond to alerts from your own systems during working hours.What you will doTurn an incomplete brief into an agreed result and workable plan. Track cross-team dependencies, explain tradeoffs, raise blockers early and communicate changes to estimates. Confirm completion with the users.Build maintainable software and cloud/infrastructure API integrations. Handle credentials, pagination, retries, partial failures and repeated execution so unattended workflows remain safe and reliable.Transform billing, usage and inventory data into repeatable outputs. Reconcile totals and periods, apply agreed allocation rules and make missing or uncertain mappings visible.Extend inventory, network-policy and capacity automation using existing systems. Retain necessary human decisions and verify the operational result.Deliver changes through version-controlled code and CI. Understand live state, test recovery and release safely; diagnose failures across code, Linux, APIs, data and networking.Own monitoring, actionable alerts, documentation and maintenance. Leave systems another engineer can operate, and use coding agents while remaining responsible for design, code, tests and outcomes.RequirementsMust haveSenior-level experience, typically 5+ years in software, infrastructure, platform or automation engineering. You can explain your own design decisions, implementation and operational results.Strong production Python, software design, testing, debugging and API integration skills. You can understand unfamiliar code and improve existing systems.Hands-on AWS experience and solid Linux/networking fundamentals, including diagnosis across host, service, network and cloud boundaries.Practical SQL and data validation: joins, units, periods, ownership mappings and reconciliation. You distinguish evidence from assumptions.Safe infrastructure delivery through IaC and CI: state, drift, idempotency, dependencies and rollback, using tools such as Terraform/OpenTofu and Ansible.Demonstrated advanced coding-agent workflows for real engineering tasks. You plan and delegate substantial work, supply context, run agents without continuous supervision within defined permissions and stop conditions, and independently explain, debug, test and verify their output.Autonomous delivery and clear collaboration: investigate problems, obtain decisions, challenge unsafe or unnecessarily complex approaches and follow through. English B2 or higher for technical discussions, written decisions and internal customer communication.Nice to haveExperience with cloud billing or FinOps tooling (e.g. AWS CUR/FOCUS, Athena), inventory or event-driven integrations (e.g. NetBox or similar), network policies, flow monitoring, identity and access management, capacity tooling, additional cloud platforms, OpenNebula, Kubernetes, Go, or self-service systems.Our environment includes Python, Ansible, Terraform/OpenTofu, GitLab CI/CD, Grafana, and related observability tools. Prior experience with every tool or domain is not required. We value the ability to quickly learn unfamiliar technologies, products, and internal business rules.BenefitsWhat's in it for you?A focus on professional development.Interesting and challenging projects.Fully remote work with flexible working hours, which allows you to schedule your day and work from any location worldwide.Paid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leaves.Compensation for private medical insurance.Co-working and gym/sports reimbursement.Budget for education.The opportunity to receive a reward for the most innovative idea that the company can patent.By applying for this position, you consent to the processing of your personal data as described in our Privacy Policy (https://cloudlinux.com/candidate-privacy-notice), which provides detailed information on how we maintain and handle your data.