远程工作雷达

系统可靠性/生产工程师

Site Reliability / Production Engineer

开发工程职能支持未标注地域日间重叠约 2 小时,需偶尔早起或晚睡
公司Storyteller
薪资€20,000/年
工作地点Algeria
地域资格未标注地域
时区要求日间重叠约 2 小时,需偶尔早起或晚睡
用工类型Contractor
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
作息提示:日间重叠约 2 小时,需偶尔早起或晚睡。

🇩🇿 年薪最高20,000欧元,全职合同制
🌎 可以在阿尔及利亚任何地方远程办公!
🌙 共享非工作时间的英国支持,包括晚间主动值班和工作日夜间电话响应
✨ 令人兴奋的高增长产品,被领先的全球品牌所依赖,尤其是在体育领域
💻 使用最新的硬件、AI工具和产品工作流程
我们寻找能够对现场系统问题负责的动手型生产工程师:确定客户影响,调查证据,采取安全措施并保持响应进度。
你将全程使用AI,但AI不会替代判断。你需要监督其输出,理解任何行动的风险,并验证真实的客户结果是否已恢复。

关于我们
Storyteller 是一个高增长的 B2B SaaS 平台,允许公司将其 Stories 集成到自己的应用和网站中。由 Instagram 和 Snapchat 推广的 Stories 帮助我们的客户提高参与度、留存率和收入。
我们的平台包含 Web、iOS 和 Android 的 SDK,以及发布工具、分析和广告支持——让企业几天内就能获得完整的 Stories 解决方案。我们与全球知名的体育和媒体品牌合作,你的工作将被数百万用户实时使用。
我们的生产环境涵盖 Storyteller、Storypilot 以及支持客户现场工作流的服务。因此可靠性不仅仅是基础设施的问题:我们需要了解客户是否受到影响,快速响应,协调合适的人选,并在每次重大事件后改进我们的系统。

职位简介
我们正在招聘两名站点可靠性/生产工程师。在你的值班期间,你将是现场事件的第一技术响应者。
与支持团队紧密合作,你将评估客户影响,调查系统,采取适当措施,并仅在需要他们的特定知识或判断时才引入产品开发人员。
这不是一个被动的升级角色。你将负责技术响应,合理地自行解决问题,并使升级具体且有用。根据事件情况,你可能需要重启或扩展服务、回滚部署、更改配置、修复数据、部署有限修复或进行小的代码更改。
当事件较少时,你将改进可靠性系统:减少警报噪音,加强客户结果监控,改进自动化检测和响应流程,以及优化故障排除流程。

查看英文原文

🇩🇿 Up to EUR 20,000 per year, on a full-time, contractor contract
🌎 Fully remote working from anywhere in Algeria! 
🌙 Shared out-of-hours UK coverage, including active evening shifts and weekday overnight pager duty
✨ Exciting high growth product, relied on by leading global brands, particularly within sports
💻 Working with the latest hardware, AI tools, and product workflows.
We are looking for hands-on production engineers who can take ownership when live systems need attention: establish the customer impact, investigate the evidence, take safe action and keep the response moving.
You will use AI throughout the work, but not as a substitute for judgement. You will be expected to supervise its output, understand the risk of any action and validate that the real customer outcome has recovered.
ABOUT US
Storyteller is a high-growth B2B SaaS platform that lets companies integrate Stories into their own apps and websites. Popularised by Instagram and Snapchat, Stories help our clients increase engagement, retention and revenue.
Our platform includes SDKs for Web, iOS and Android, alongside publishing tools, analytics and advertising support—giving enterprises a complete Stories solution in days. We work with globally recognised sports and media brands, and your work will be used live by millions of people.
Our production environment spans Storyteller, Storypilot and the services that support our customers’ live workflows. Reliability is therefore about more than infrastructure: we need to understand when customers are affected, respond quickly, coordinate the right people and improve our systems after every material incident.

About the Role
We are hiring two Site Reliability / Production Engineers. You will be the first technical response for live incidents during your coverage window.
Working closely with Support, you will assess customer impact, investigate the system, take proportionate action and bring in product developers only when their specific knowledge or judgement is genuinely needed.
This is not a passive escalation role. You will own the technical response, solve what you reasonably can yourself and make escalations specific and useful. Depending on the incident, you may restart or scale services, roll back deployments, change configuration, repair data, deploy a bounded fix or make a small code change.
When incidents are quiet, you will improve the reliability system: reduce alert noise, strengthen customer-outcome monitoring, improve diagnostics, create runbooks and AI Skills, automate repeated work and make our products easier to operate.
Working Pattern
This role provides out-of-hours production coverage, so the schedule is a core part of the position rather than occasional overtime.

  • The two hires will share an agreed rota that ensures one engineer is actively working from 17:00-01:00 UK time, seven days a week.
  • Neither person will work seven days a week; the active shifts will be divided between both hires, with appropriate rest days.
  • The two engineers will also share weekday pager coverage from 01:00-06:00 UK time. The detailed allocation of active and pager shifts will be explained during the hiring process.
  • Weekend daytime coverage is provided separately and is not an additional expectation for these roles.
  • When there are no live incidents, the active shift will be used for reliability-improvement work.
  • Meetings and collaboration with management and product teams will be arranged within the agreed working pattern.

The detailed rota, rest arrangements, leave cover, compensation and on-call terms will be confirmed clearly during the hiring process. Please consider the UK-time evening and overnight requirements carefully before applying.
RESPONSIBILITIES
Respond to live incidents

  • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state.
  • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code.
  • Use AI throughout triage and diagnosis while checking its conclusions against real evidence.
  • Choose and execute a proportionate mitigation, rollback, repair or bounded fix.
  • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green.
  • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication.
  • Join customer conversations occasionally when direct technical involvement is genuinely useful.

Coordinate the right response

  • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix.
  • Escalate with evidence, customer impact, actions already taken and the specific decision or help required.
  • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait.
  • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners.

Improve the reliability system

  • Remove, consolidate and tune low-value alerts, and design monitoring around real service and customer outcomes.
  • Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into better alerts, runbooks, AI Skills, automation or product improvements.
  • Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve.
  • Create safe, supervised automation for common operational actions.
  • Work with product teams to close observability, rollback, runbook and supportability gaps.
  • Detect and help contain unusual service-cost behaviour, then route wider follow-up to the appropriate cost or product owner.
  • Make reliability and on-call performance easier for the company to understand and improve over time.

QUALIFICATIONS
What we're looking for

  • Agency and ownership - You take responsibility for ambiguous live problems, gather evidence, choose a path and follow through after the immediate pressure has passed.
  • Operational judgement - You can separate customer impact, symptoms and likely causes, make practical decisions under uncertainty and recognise when an intervention is no longer safe or bounded.
  • Technical comfort and aptitude - You are comfortable exploring unfamiliar systems through code, logs, APIs, data, infrastructure and command-line tools, and can make hands-on changes with a clear validation plan.
  • AI-native execution - You use AI for substantive technical work - investigation, hypothesis generation, code, automation, incident analysis and workflow improvement - while supervising the agent and challenging its conclusions.
  • Accuracy and validation discipline - You actively look for false confidence and verify outcomes through appropriate technical and customer signals.
  • Systems thinking - You look for repeated patterns and improve the triggers, owners, runbooks, automation, metrics and feedback loops around the work.
  • Clear coordination and communication - You communicate calmly and concisely with Support, developers and non-technical stakeholders, making evidence, impact, uncertainty, ownership and next actions easy to understand.
  • Curiosity and resilience - You learn unfamiliar products and tools quickly, keep investigating when the first hypothesis fails and change your approach when the evidence demands it.

Previous responsibility for live production systems or an on-call rota is strongly preferred because it is useful evidence that you understand the realities of incident response. It is not an automatic requirement: we will also consider candidates who demonstrate exceptional ownership, judgement, technical aptitude, learning velocity and performance in the practical assessment.
You do not need experience with every technology in our stack, a previous SRE job title, people-management experience or the ability to recall every command without AI assistance. The ability to learn an unfamiliar environment, act safely and validate your work matters more than matching a long technology checklist.
Nice to have

  • Cloud platforms such as Azure or Cloudflare.
  • Distributed application and API diagnostics.
  • Databases, queues and background-processing systems.
  • Observability, alerting and incident-management platforms.
  • Infrastructure, deployment and release automation.
  • Application development and safe production debugging.
  • AI coding agents and workflow automation.

RECRUITMENT PROCESS
We keep the process straightforward, practical and respectful of your time.
1. Hiring Manager Conversation (20-30 mins)
A short call to get to know you, talk through the working pattern and answer your questions.
2. Paid Take-home Task (~60-90 mins)
A small, bounded production-incident exercise using evidence such as a Support report, alerts, logs, metrics, deployment history, code and an imperfect runbook. We compensate you for completing it regardless of the outcome.
You are encouraged to use AI. We are interested in how you establish impact, investigate and revise hypotheses, choose a safe response, validate the outcome and communicate the incident - not in your ability to reproduce commands or syntax from memory.
3. Task Review and CTO Interview (60-75 mins)
We will review your submission together, explore the decisions and trade-offs you made, and discuss how you supervised AI-generated analysis or changes. You will also meet Dave, our CTO, and talk about production judgement, escalation, validation, communication and how you improve the system after an incident.
And that's it.
Privacy Notice
We process your personal data for recruitment purposes in line with UK data protection law. AI tools may assist in reviewing applications, but decisions are made by our team. We retain data only as necessary for recruitment and compliance. You can request access or deletion of your data at any time by emailing .
Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位