高级SRE/DevOps工程师(Kubernetes/A2P消息)
Senior SRE / DevOps Engineer (Kubernetes / A2P Messaging)
这个环境适合你如果喜欢从零开始构建基础设施,解决复杂的可靠性问题,并紧密关注生产环境中的系统
你将负责一个新高吞吐量消息平台的基础设施建设与运维,从Kubernetes集群和网络层到可观测性、部署和事件响应全程负责
该平台每天处理约100万条消息,支持约175个客户和120个供应商,并维护超过300个长期运行的消息连接
这不是典型的无状态Web环境。该平台运行长期的TCP会话,需要在部署、故障和流量高峰期间保持稳定的网络端点和可预测的行为
你将参与两个基础设施站点的平台建设,并伴随其经历生产就绪、迁移、超期支持和稳定运行阶段
薪资:每小时150–190 PLN + VAT(B2B)
地点:完全远程(波兰)🌎 或波兰罗兹(Łódź)
日常挑战:
- 在能提升交付效率的地方使用AI辅助工程——包括基础设施工具、自动化、文档和操作手册
- 使用基础设施即代码和自动化交付,在两个站点构建和运维Kubernetes基础设施
- 设计部署机制,使长期连接能够优雅地断开,而不是在发布过程中被中断
- 负责平台网络,包括稳定的入口和出口IP、第4层负载均衡、TLS以及与外部基础设施提供商协调的连接性
- 构建和运维可观测性堆栈:指标、日志、仪表板和告警,提供有用信号而非噪音
- 围绕真实生产行为设计监控,包括对单个连接和故障模式的可见性
- 管理证书生命周期、秘密管理和平台访问控制
- 在高可用环境中运维PostgreSQL,包括复制、故障转移、备份和验证恢复
- 设计并执行跨两个站点的备份和灾难恢复流程
- 与工程师和QA密切合作进行性能和生产规模的负载测试,调查最先失败的部分及其原因
- 在迁移和生产切换期间支持平台
- 参与生产事件响应,并在稳定运行阶段参与24×7的值班轮班
我们寻找的人选:
- 有电信、运营商相关经验
查看英文原文
This environment will suit you if you enjoy building infrastructure from the ground up, solving hard reliability problems and staying close to the systems you operate in production.
You’ll build and operate the infrastructure behind a new high-volume messaging platform, taking ownership from the Kubernetes cluster and networking layer through observability, deployment and incident response.
The platform processes around 1 million messages every day, supports approximately 175 customers and 120 suppliers, and maintains more than 300 long-lived messaging connections.
This is not a typical stateless web environment. The platform runs long-lived TCP sessions that need stable network endpoints and predictable behaviour during deployments, failures and traffic spikes.
You’ll help build the platform across two infrastructure sites and stay with it through production readiness, migration, hypercare and steady-state operation.
Rate: 150–190 PLN + VAT (B2B) per hour.
Location: Fully remote (Poland) 🌎 or Łódź (Poland)
Your everyday challenges:
- Use AI-assisted engineering where it improves delivery - including infrastructure tooling, automation, documentation and runbooks.
- Build and operate Kubernetes infrastructure across two sites using infrastructure-as-code and automated delivery.
- Design deployment mechanisms that allow long-lived connections to drain gracefullyinstead of being dropped during releases.
- Take ownership of platform networking, including stable ingress and egress IPs, Layer 4 load balancing, TLS and connectivity coordinated with an external infrastructure provider.
- Build and operate the observability stack: metrics, logs, dashboards and alerting that provide useful signal rather than noise.
- Design monitoring around real production behaviour, including visibility into individual connections and failure modes.
- Own certificate lifecycle, secrets management and platform access control.
- Operate PostgreSQL in a highly available environment, including replication, failover, backups and verified restores.
- Design and exercise backup and disaster-recovery procedures across two sites.
- Work closely with engineers and QA on performance and production-scale load testing, investigating what fails first and why.
- Support the platform during migration and production cutovers.
- Take part in production incident response and from the steady-state phase, a 24×7 on-call rotation.
What are we looking for?
- Experience with telecom, carrier, messaging or other environments built around persistent network connections.
- Knowledge of SMPP, SIP, SS7 or similar telecom protocols.
- Strong production experience operating Kubernetes on self-managed, on-premise or similarly infrastructure-heavy environments - not only managed cloud services.
- Experience with stateful, long-lived TCP workloads on Kubernetes, including connection draining, stable ingress/egress, Layer 4 load balancing and deployment behaviour.
- Strong Linux and networking fundamentals. You are comfortable diagnosing problems involving routing, NAT, firewalls, MTU, TLS or packet-level behaviour.
- Practical troubleshooting experience with tools such as tcpdump and production network diagnostics.
- Experience designing monitoring and alerting, not only maintaining dashboards somebody else created. You understand what deserves to wake a human up, and what does not.
- Strong infrastructure-as-code and CI/CD experience for containerised systems.
- Practical PostgreSQL operations experience, including replication, failover, backup and importantly - verified restore.
- Ability to work directly with engineers from external infrastructure and network providers.
- Professional English.
- Real production on-call and incident-response experience.
What will strengthen your candidacy?
- Experience configuring or troubleshooting IPsec connectivity with external parties.
- Experience designing or operating multi-site active-active or active-passive environments.
- Hands-on disaster-recovery exercises rather than DR plans that existed only on paper.
- Performance engineering experience, including Linux kernel or network tuning for high connection counts.
- Security hardening experience, including CIS-style benchmarks, vulnerability management or software supply-chain practices.
- Experience being the first SRE or Platform Engineer on a system and defining how it should be operated.
Why join us?
- Build it and run it: you won’t inherit an infrastructure somebody else designed or throw your work over the wall after launch. You’ll help build the platform and remain close to it in production.
- A genuinely difficult reliability problem: long-lived protocol traffic, stateful connections, two-site infrastructure and production traffic at telecom scale.
- Influence architecture from day one: deployment, observability and operability are design constraints here, not tasks postponed until after development.
- Greenfield infrastructure: you’ll have real influence over how the Kubernetes platform, delivery pipelines, monitoring and operational practices are created.
- Small senior team: short decision paths and direct collaboration with engineers and the solution architect.
- Production ownership: you’ll follow the system through build, migration, hypercare and steady-state operation instead of disappearing after implementation.
- Modern engineering environment: automation and AI-assisted tooling are used where they genuinely improve engineering and operational work.
Sounds like a fit? Let’s talk! 🚀
Originally posted on Himalayas