远程工作雷达

[MLA] 高级站点可靠性工程师(SRE) – Kubernetes

[MLA] Senior Site Reliability Engineer (SRE) – Kubernetes

开发工程未标注地域
公司Software Mind
薪资未公开
工作地点Kraków, Lesser Poland Voivodeship, Poland
地域资格未标注地域
时区要求无特别要求
用工类型Full-time
发布时间2026-08-17
数据来源SmartRecruiters
前往企业招聘页投递 →

Software Mind 为全球的公司提供具有影响力的技术解决方案。科技巨头和独角兽企业、变革性项目、新兴技术以及无限的机会——这些是我们的日常写照。我们组建跨职能的工程团队,他们拥有责任感并追求卓越,因此我们一直在寻找那些在每个项目中都充满激情和创造力的优秀人才。我们的文化崇尚开放、尊重、坚韧与勇气,并将工作与乐趣结合在一起。

项目 – 你将承担的职责
我们是 AI Experience Framework 团队,负责构建支撑 ServiceNow AI 首先用户界面的平台——一个基于 Lit 和服务端渲染 Web 组件的 SSR 运行时(karuna),运行在多层代理/HTTP2 路由链之后,配有分片的 V8 隔离池,同时搭配 ServiceNow Glide/Java 平台层(karuna-glide),提供元数据、ACL 和服务构件。该职位负责整个堆栈的生产可靠性:Kubernetes 部署与运维、可观测性,以及对系统 Node.js 和 JVM 端的现场故障排查——不是通用的基础设施工作。
职位 – 你将如何贡献
· 支持 Kubernetes 上运行的生产服务的部署、运维和可靠性。
· 监控服务健康状况,并调查分布式应用中的生产事件。
· 参与值班支持、事件响应、根本原因分析、事后总结和可靠性改进。
· 与工程团队合作,排查应用程序运行时、网络和服务间的问题。
· 支持 CI/CD、基于 GitOps 的部署、可观测性和生产监控。
· 在客户导向的待办事项列表和既定优先级内工作。

要求 – 你需要的经验
· 5 年以上在站点可靠性工程、DevOps、平台工程、生产工程或相关领域的经验,包括近期在 Kubernetes 基础的生产服务中具备扎实的实践经验。
· 3 年以上 Kubernetes 生产环境的实践经验,包括部署、扩展、发布/回滚、资源调优和服务间故障排查。
· 强大的生产事件响应经验,包括值班、操作手册、事后总结和通知规范。
· 使用 Splunk 进行日志聚合、搜索和生产环境故障排查的经验。
· Prometheus 和 Grafana 的使用经验。

查看英文原文

Software Mind develops solutions that make an impact for companies around the globe. Tech giants & unicorns, transformative projects, emerging technologies and limitless opportunities – these are a few words that describe an average day for us. Building cross-functional engineering teams that take ownership and crave more means we’re always on the lookout for talented people who bring passion and creativity to every project. Our culture embraces openness, acts with respect, shows grit & guts and combines employment with enjoyment.

Project – the aim you'll have 
We are the AI Experience Framework team that builds the platform powering ServiceNow's AI-first user interfaces - an SSR runtime (karuna) built on Lit and server-rendered web components, running behind a multi-tier proxy/HTTP2 routing chain with sharded V8 isolate pools, paired with a ServiceNow Glide/Java platform layer (karuna-glide) that supplies metadata, ACLs, and service artifacts. This role owns production reliability for that stack end to end: Kubernetes deployment and operations, observability, and hands-on troubleshooting of both the Node.js and JVM sides of the system - not generalist infrastructure work. 
Position – how you’ll contribute
· Support the deployment, operation, and reliability of production services running on Kubernetes.
· Monitor service health and investigate production incidents across distributed applications.
· Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.
· Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams.
· Support CI/CD, GitOps-based deployments, observability, and production monitoring.
· Work within a client-directed backlog and established priorities.

Expectations – the experience you need
· 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services.
· 3+ years of hands-on production Kubernetes experience strongly preferred. Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshooting
· Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene
· Splunk experience for log aggregation, search, and production troubleshooting
· Prometheus and Grafana experience, specifically building alert rules and dashboards, not only using existing dashboards
· CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux
· Strong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networking
· Production troubleshooting experience across Node.js and JVM/Java services, with strong depth in at least one runtime environment. Experience may include Node.js heap snapshots, CPU profiling, event-loop and memory analysis, as well as JVM GC log analysis, thread dumps, JVM tuning, and Java service latency investigation.
· Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication
· Very good spoken and written English. 
Additional skills – the edge you have
· Web Components / Lit experience, to perform first-level debugging of UI-related issues
· Server-side rendering or isomorphic runtime experience
· Canary rollout / multi-version production operations
· Distributed tracing and request-context correlation
· KEDA or event-driven autoscaling
· Experience with enterprise platform integration layers

Our offer – professional development, personal growth:
· Flexible employment and remote work  
· International projects with leading global clients 
· International business trips  
· Non-corporate atmosphere 
· Language classes 
· Internal & external training 
· Private healthcare and insurance  
· Multisport card 
· Well-being initiatives 
Position at: Software Mind

本页面信息整理自 SmartRecruiters,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位