远程工作雷达

高级平台工程师,云基础设施

Senior Platform Engineer, Cloud Infrastructure

开发工程职能支持限定地区(需当地身份)
公司Virtasant
薪资未公开
工作地点United States
地域资格限定地区(需当地身份)
时区要求日间重叠约 9 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 United States 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

高级平台工程师,云基础设施
类型:远程办公
工作时间:太平洋时区(上午8:00 – 下午5:00 PST)
职位描述:
我们正在寻找一位资深工程师来构建和运营我们的产品所依赖的云原生平台。这是一个实际操作的基础设施工程岗位:您将设计并运行生产环境的Kubernetes平台,编写和维护扩展该平台的Go服务和控制器,并负责其他工程团队所依赖系统的可靠性。
这项工作涵盖平台架构、网络、可观测性以及生产运维。您将编写实际代码:Go服务、Kubernetes控制器、自定义中间件,但您创造的价值是通过平台功能和可靠性来衡量的,而不是交付的代码行数。我们希望找到一位能够承担模糊且跨季度的项目,并将其推向生产环境的人选。
主要职责:
平台与Kubernetes工程:

  • 设计、构建和运营生产环境的Kubernetes集群,包括集群网络、工作负载隔离和多区域拓扑。
  • 与Kubernetes内部机制协作:资源配额管理、调度和集群行为、NetworkPolicy执行,以及自定义控制器或运营商。
  • 实现和运营服务网格功能:服务间的mTLS、基于服务账户的身份验证和授权,以及内部和外部请求路径的流量管理。
  • 优化容器化工作负载的性能、成本和资源效率。

Go、Python或Java软件开发:

  • 编写、重构和维护扩展平台的生产服务、控制器和中间件。
  • 阅读并为复杂的现有代码库做出贡献,包括我们定制或扩展的开源项目。
  • 构建内部工程团队使用的HTTP、REST和gRPC服务接口。
  • 编写有意义的单元测试和集成测试,并将可测试性作为设计属性,而非事后考虑。

可靠性与生产运维:

  • 领导平台级问题的事件响应;调查根本原因并撰写导致持久解决方案的复盘报告。
  • 使用日志、指标、追踪和分析工具排查生产系统。
  • 在分布式系统中诊断和解决性能和可靠性问题。
  • 定义并推动SLO,构建使其可操作的告警机制。

基础设施即代码与交付:

  • 负责整个平台的基础设施即代码。
查看英文原文

Senior Platform Engineer, Cloud Infrastructure
Type: Remote
Coverage: Pacific Hours (8:00 AM – 5:00 PM PST)
Job Description:
We are looking for a senior engineer to build and operate the cloud-native platform that our products run on. This is a hands-on infrastructure engineering role: you will design and run production Kubernetes platforms, write and maintain the Go services and controllers that extend them, and own the reliability of systems that other engineering teams depend on.
The work spans platform architecture, networking, observability, and production operations. You will write real code: Go services, Kubernetes controllers, custom middleware, but the value you create is measured in platform capability and reliability, not lines shipped. We are looking for someone who is comfortable owning ambiguous, multi-quarter initiatives and driving them to production.
Key Responsibilities:
Platform and Kubernetes Engineering:

  • Design, build, and operate production Kubernetes clusters, including cluster networking, workload isolation, and multi-region topologies.
  • Work with Kubernetes internals: resource quota management, scheduling and cluster behaviour, NetworkPolicy enforcement, and custom controllers or operators.
  • Implement and operate service mesh capabilities: mTLS between services, service-account-level authentication and authorisation, and traffic management across internal and external request paths.
  • Optimise containerised workloads for performance, cost, and resource efficiency.

Software Development in Go, Python or Java:

  • Write, refactor, and maintain production services, controllers, and middleware that extend the platform.
  • Read and contribute to complex existing codebases, including open-source projects that we customise or extend.
  • Build HTTP, REST, and gRPC service interfaces used by internal engineering teams.
  • Write meaningful unit and integration tests, and treat testability as a design property rather than an afterthought.

Reliability and Production Operations:

  • Lead incident response for platform-level issues; investigate root causes and author postmortems that result in durable fixes.
  • Troubleshoot production systems using logs, metrics, traces, and profiling tools.
  • Diagnose and resolve performance and reliability problems across distributed systems.
  • Define and drive SLOs, and build the alerting that makes them actionable.

Infrastructure as Code and Delivery:

  • Own infrastructure as code across the platform, authoring reusable modules and maintaining them as the platform evolves.
  • Build and improve CI/CD and GitOps delivery workflows so that teams can ship safely and frequently.
  • Balance developer velocity against reliability, security, and compliance requirements.
  • Plan and execute cloud migration initiatives, including moving production workloads between cloud providers or environments while maintaining reliability and minimizing downtime.

Observability:

  • Build and maintain metrics, dashboards, alerting policies, and distributed tracing.
  • Instrument services so that failures are diagnosable without a code change.

Collaboration:

  • Partner with product, security, and infrastructure teams to gather requirements and align on architecture.
  • Contribute to design reviews and help set technical direction.
  • Mentor other engineers and raise the standard of engineering practice around you.

Qualifications:
Education and Experience:

  • 6+ years of professional experience in software, platform, infrastructure, or site reliability engineering, including significant time operating production distributed systems.
  • Demonstrated experience building and operating production Kubernetes platforms, not only deploying onto them.
  • Production experience writing Go, Python or Java.
  • Experience designing systems from an ambiguous starting point and carrying them to production.
  • Experience migrating cloud services, including planning and executing moves of production workloads between providers or environments.
  • Degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

Technical Skills:

  • Strong understanding of Kubernetes internals: networking (CNI), NetworkPolicy, resource management, and cluster behaviour under load.
  • Hands-on experience with a service mesh (Istio, Envoy, Linkerd, or similar) and with mTLS and workload identity.
  • Solid Linux fundamentals, including cgroups and resource management.
  • Infrastructure as code at scale, Terraform or an equivalent.
  • Production experience with at least one major cloud platform (GCP, AWS, or Azure); multi-cloud experience is a strong plus.
  • Observability tooling: Prometheus, Grafana, OpenTelemetry, and query languages such as PromQL.
  • Production experience with relational databases, including PostgreSQL or managed Postgres-compatible services, and an understanding of their replication and failover characteristics.
  • Docker and container tooling as part of the delivery lifecycle.
  • Strong debugging and performance profiling skills.

Preferred:

  • Experience with Go testing frameworks such as Ginkgo and Gomega.
  • Experience building Kubernetes controllers, operators, or other API-server extensions.
  • Additional strength in Python.
  • Experience with identity and access management: SSO, Keycloak, OIDC, SAML, or secret management with Vault or a cloud equivalent.
  • Experience designing for high availability and disaster recovery across regions or providers.
  • Experience working in a monorepo, and with build systems such as Bazel.
  • Ability to read Java.
  • Exposure to compliance frameworks such as SOC 2 or GDPR.
  • Interest in the reliability and safety of LLM-backed systems running in cloud-native environments.
  • Experience with Alibaba Cloud (AliCloud) is highly preferred.
  • Experience with large-scale Data Platform technologies such as Apache Spark and Apache Flink is highly preferred.

Soft Skills:

  • Strong analytical and problem-solving ability.
  • Clear written and verbal communication, particularly in design documents, postmortems, and cross-team requirements gathering.
  • Able to work independently while contributing effectively within a distributed team.
  • Comfortable in a fast-moving, highly technical environment.

Please Note:
· This is a platform engineering role. We are not looking for candidates whose experience has been limited to consuming Kubernetes or cloud services. The ideal candidate understands how the underlying platforms work, has built and operated production Kubernetes infrastructure, and can explain the architecture, implementation, and operational trade-offs behind the tools they use.

Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位