DevOps工程师
DevOps Engineer
职位描述
我们正在寻找一名DevOps工程师,为我们的AI驱动产品构建基础设施基础,使其快速、可靠、安全且成本高效。你将搭建云、部署和监控系统,使小团队能够自信地发布,并在试点阶段及以后顺利扩展。
你将负责:
· 设计、配置和管理云基础设施(AWS、GCP或Azure)
· 构建和维护CI/CD流水线,实现快速、安全的自动化部署
· 使用基础设施即代码(Terraform、Pulumi或其他类似工具)实现可重复的环境
· 随着系统增长,对服务进行容器化和编排(Docker、Kubernetes)
· 建立可观测性:跨整个技术栈的监控、日志、警报和追踪
· 管理和优化LLM和AI工作负载的成本、延迟和可靠性
· 在各环境中负责安全、密钥管理和访问控制
· 建立备份、灾难恢复和事件响应实践
要求:
· 4年以上DevOps、SRE或基础设施工程经验
· 精通至少一个主要云服务商(AWS/GCP/Azure)
· 熟练使用基础设施即代码工具(Terraform、Pulumi、CloudFormation)
· 具备容器化和编排经验(Docker、Kubernetes)
· 具备扎实的CI/CD经验(GitHub Actions、GitLab CI、CircleCI或其他类似工具)
· 强大的脚本编写能力(Bash、Python或Go)
· 具备实施监控和可观测性工具的经验(Prometheus、Grafana、Datadog等)
· 具备以安全为先的思维,有密钥和访问管理经验
加分项:
· 具备管理AI/ML或LLM密集型工作负载基础设施的经验
· 熟悉高吞吐量API使用的成本优化
· 具备GPU基础设施或推理服务经验
· 具备从零开始搭建基础设施的早期经验
查看英文原文
About the Role
We're looking for a DevOps Engineer to build the infrastructure foundation that keeps our AI-powered product fast, reliable, secure, and cost-efficient. You'll set up the cloud, deployment, and monitoring systems that let a small team ship confidently and scale smoothly through the pilot and beyond.
What You'll Do
· Design, provision, and manage cloud infrastructure (AWS, GCP, or Azure)
· Build and maintain CI/CD pipelines for fast, safe, automated deployments
· Implement infrastructure-as-code (Terraform, Pulumi, or similar) for reproducible environments
· Containerize and orchestrate services (Docker, Kubernetes) as the system grows
· Set up observability: monitoring, logging, alerting, and tracing across the stack
· Manage and optimize the cost, latency, and reliability of LLM and AI workloads
· Own security, secrets management, and access controls across environments
· Establish backup, disaster-recovery, and incident-response practices for the pilot
- 4+ years of DevOps, SRE, or infrastructure engineering experience
- Strong hands-on experience with at least one major cloud provider (AWS/GCP/Azure)
- Proficiency with infrastructure-as-code tools (Terraform, Pulumi, CloudFormation)
- Experience with containerization and orchestration (Docker, Kubernetes)
- Solid CI/CD experience (GitHub Actions, GitLab CI, CircleCI, or similar)
- Strong scripting skills (Bash, Python, or Go)
- Experience implementing monitoring and observability tooling (Prometheus, Grafana, Datadog, etc.)
- Security-first mindset and experience with secrets and access management
- Experience managing infrastructure for AI/ML or LLM-heavy workloads
- Familiarity with cost optimization for high-throughput API usage
- Experience with GPU infrastructure or inference serving
- Early-stage experience standing up infrastructure from scratch