周末后端工程师
Weekend Backend Engineer
职位描述
这个职位结合了平台工程和结构化的第一线运维职责。
你将在中央工程团队工作,构建和改进我们的内部开发者平台——重点在于可观测性、负载测试、弹性及系统稳健性——同时参与定义好的一级事件值班轮班。
我们的系统每天处理数百万请求,分布在多个微服务中。稳定性、可扩展性和性能至关重要。此职位直接有助于提升大规模下的可靠性,而不仅仅是功能交付。
你将按照结构化的区域轮班模式工作,以帮助提供7x24小时的覆盖,作为我们第一线响应模型的一部分。
我们的技术栈(部分经验是预期的)
- 语言:Java 21
- 框架:Spring Boot、Spring Data、Spring Cloud
- 架构:微服务、REST APIs、事件驱动系统
- 数据库:MySQL、MyBatis、ShardingSphere、MongoDB
- 缓存:Redis(AWS ElastiCache)、Elasticsearch
- 消息:RocketMQ
- 云/基础设施:Docker、Kubernetes、AWS
- 可观测性:Grafana、Prometheus、Loki、Tempo、CloudWatch
你将负责的工作内容
平台工程(值班时间外为主)
- 提升服务间的可观测性(指标、追踪、日志)。
- 设计并增强负载测试框架和弹性工具。
- 构建可复用的平台能力和内部库。
- 识别重复出现的运维痛点并永久解决。
- 为提升可扩展性和可维护性的标准做出贡献。
- 参与架构讨论和技术评审。
一级运维值班(在分配的轮班时间内)
- 参与定义好的区域轮班模式。
- 在值班时间内提供第一线事件响应。
- 通过Rootly和结构化运行手册处理警报。
- 执行已批准的缓解步骤。
- 必要时升级到产品团队。
- 确保轮班之间的清晰文档和交接。
此职位不负责产品事件的最终解决,但确保结构化和快速的初步排查。
你将带来的能力
- 3年以上后端工程经验(Java/Spring优先)。
- 具有高吞吐量生产系统的运维经验。
- 对分布式系统和故障模式有深入理解。
- 具备可观测性工具和生产诊断的经验。
- 理解负载测试和性能调优。
- 强大的SQL技能。
- 具备容器化/云环境的经验
查看英文原文
About the Role
This role combines platform engineering and structured first-line operational responsibility.
You will work within Central Engineering, building and improving our internal developer platform - focusing on observability, load testing, resiliency, and system robustness - while also participating in a defined Level 1 incident duty rotation.
Our systems handle millions of requests daily across distributed microservices. Stability, scalability, and performance are critical. This role directly contributes to improving reliability at scale, not just feature delivery.
You will work in a structured regional shift pattern to help provide 24/7 coverage as part of our first-line response model.
Our Stack (experience in some is expected)
- Language: Java 21
- Frameworks: Spring Boot, Spring Data, Spring Cloud
- Architecture: Microservices, REST APIs, Event-Driven Systems
- Databases: MySQL, MyBatis, ShardingSphere, MongoDB
- Caching: Redis (AWS ElastiCache), Elasticsearch
- Messaging: RocketMQ
- Cloud/Infra: Docker, Kubernetes, AWS
- Observability: Grafana, Prometheus, Loki, Tempo, CloudWatch
What You’ll Be Doing
Platform Engineering (Primary Focus Outside Duty Window)
- Improve observability across services (metrics, tracing, logging).
- Design and enhance load testing frameworks and resiliency tooling.
- Build reusable platform capabilities and internal libraries.
- Identify recurring operational pain points and eliminate them permanently.
- Contribute to standards that improve scalability and maintainability.
- Participate in architecture discussions and technical reviews.
Level 1 Operational Duty (During Assigned Shift Window)
- Participate in a defined regional shift pattern.
- Provide first-line incident response within your duty window.
- Triage alerts via Rootly and structured runbooks.
- Execute pre-approved mitigation steps.
- Escalate to product teams when required.
- Ensure clear documentation and handover between shifts.
This role does not own product incident resolution, but ensures structured and rapid triage.
What You’ll Bring
- 3+ years experience in backend engineering (Java/Spring preferred).
- Experience operating high-throughput production systems.
- Strong understanding of distributed systems and failure modes.
- Experience with observability tools and production diagnostics.
- Understanding of load testing and performance tuning.
- Strong SQL proficiency.
- Experience with containerised/cloud environments (Docker/K8s/AWS).
- Ability to work autonomously and follow structured operational protocols.
- Clear written and verbal communication skills in English.
Preferred:
- Experience building shared internal tooling or platform libraries.
- Experience participating in on-call or incident response rotations.
Working Pattern
- This role operates within a defined regional shift structure to support 24/7 coverage. This is shift-based coverage during working hours (not overnight pager callouts).
- Engineers are expected to be available for rapid response during their assigned duty window.
- Exact shift hours and rotation details will be discussed during the interview process.