远程工作雷达

站点可靠性工程师

Site Reliability Engineer

开发工程职能支持限定地区(需当地身份)日间重叠仅 1 小时,需熬夜配合
公司Intermedia Intelligent Communications
薪资未公开
工作地点Portugal
地域资格限定地区(需当地身份)
时区要求日间重叠仅 1 小时,需熬夜配合
用工类型Full Time
发布时间2 天前
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 Portugal 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。
作息提示:日间重叠仅 1 小时,需熬夜配合。

关于Intermedia

你是否在寻找一家能倾听你声音的公司?一家你能带来改变的公司?你是否能在快节奏的工作环境中茁壮成长?你是否每天早上都充满期待地与优秀的人一起创造成功?如果是这样,那么Intermedia就是你的选择。
Intermedia已成为云通信和协作技术的领先提供商,使企业能够更好地连接。我们有强劲的增长、盈利记录,并营造了一个每个人都很重要的环境。在这里,我们快节奏且有时颇具挑战性,但保证你不会感到无聊。你会在Intermedia找到一个可以尽情发挥对创建和支撑优秀云技术的热情的地方。此外,我们始终注重内部晋升,许多员工已经在我们这里工作了10年、15年甚至20年以上!

Intermedia的文化建立在团队合作和透明沟通之上。我们彼此负责,始终互相支持!
你准备好留下自己的印记了吗?

职位描述

我们正在寻找一名站点可靠性工程师(SRE),以增强客户对我们系统始终可用的信心。随着我们产品部署的全球扩展,该职位将加强我们的核心SRE职能,并支持语音/统一通信(UC)服务与人工智能驱动功能的整合。你将在生产基础设施、实时通信和依赖AI的服务之间工作,以提升可靠性、可扩展性、可观测性和客户体验,同时与全球跨职能团队协作。
你将负责:

• 通过监控可用性并从整体视角审视系统健康状况来运行和优化生产环境。
• 构建软件、自动化工具和系统以管理平台基础设施和应用程序。
• 提升我们一系列云软件解决方案的可靠性、质量、可扩展性和上市速度。
• 测量和优化系统性能,提前识别容量限制、故障模式和性能瓶颈,避免影响客户。
• 为大型分布式应用程序和服务提供运维支持和工程支持。
• 收集和分析操作系统、应用、网络和服务指标,以支持性能调优和故障隔离。
• 与开发团队合作,通过严格的测试、发布流程和生产就绪实践改进服务。

查看英文原文

About Intermedia

Are you looking for a company where YOUR VOICE is heard? Where can you MAKE A DIFFERENCE? Do you
THRIVE in a FAST-PACED work environment? Do you wake every morning EXCITED to work with GREAT
PEOPLE and create SUCCESS TOGETHER? Then Intermedia is the place for you.
Intermedia has established itself as a leading provider of cloud communications and collaboration tech that allows
companies to connect better. We have a strong track record of growth, profitability, and creating an environment
where everyone matters. Everyone. While we are fast-paced and admittedly a bit intense, we promise that you won’t
be bored. You will find Intermedia is a place where you can indulge your passion for creating and supporting great
cloud technology. What’s more, we always look to promote from within and have many employees who have been
with us 10, 15, and 20+ years!

Culture at Intermedia is built on teamwork and transparency. We hold each other accountable and always have each other’s back!
Are you ready to make your mark?

About the Role

We are looking for a Site Reliability Engineer (SRE) to provide confidence to our customers that our systems will be
available at all times. As we continue the global expansion of our product deployments, this role will help strengthen
our core SRE function and support the integration of Voice/Unified Communications (UC) services with AI-powered
capabilities. You will work across production infrastructure, real-time communications, and AI-dependent services to
improve reliability, scalability, observability, and customer experience while collaborating with global, crossfunctional teams.
What you will be doing:

• Run and improve production environments by monitoring availability and taking a holistic view of system health.
• Build software, automation, and systems to manage platform infrastructure and applications.
• Improve reliability, quality, scalability, and time-to-market across our suite of cloud software solutions.
• Measure and optimize system performance, identifying capacity constraints, failure modes, and performance
bottlenecks before they affect customers.
• Provide operational support and engineering for large, distributed applications and services.
• Gather and analyze operating system, application, network, and service metrics to support performance tuning
and fault isolation.
• Partner with development teams to improve services through rigorous testing, release procedures, and
production-readiness practices.
• Participate in system design, platform management, capacity planning, incident response, and post-incident
improvement.
• Create sustainable systems and services through automation, infrastructure improvements, and reduction of
operational toil.
• Balance feature delivery and reliability using well-defined service level indicators, service level objectives, and
error budgets.
• Improve the reliability of Voice/UC platforms and integrations, with attention to real-time signaling and media
flows, call quality, latency, jitter, packet loss, failover, and end-to-end service availability.
• Partner with Voice/UC and AI engineering teams to operationalize AI-enabled communications capabilities
such as speech recognition, transcription, summarization, intelligent routing, conversational assistance, and
text-to-speech where applicable.
• Build observability for end-to-end Voice/UC + AI service paths, correlating communications telemetry with
application, infrastructure, and AI-service metrics to accelerate detection and root-cause analysis.
• Design and test graceful degradation, dependency isolation, retry/fallback patterns, and recovery procedures
so core communications remain resilient when AI or downstream services are impaired.
• Automate validation and production-readiness checks for Voice/UC + AI integrations, including performance,
scale, reliability, and release verification.

What you will bring to the role:

• Bachelor degree in computer science or another highly technical or scientific discipline, or equivalent practical
experience.
• 4-7 years of experience in technical roles involving production operations, systems engineering, SRE/DevOps,
CI/CD implementation, software deployment, and maintenance of production systems.
• Experience with Agile methodologies, DevOps practices, CI/CD pipelines, infrastructure automation, and
production monitoring/observability.
• Experience with distributed systems, cloud infrastructure, containers, and dynamic resource management
frameworks such as Kubernetes.
• Experience with distributed storage technologies such as NFS, HDFS, or S3, or comparable cloud storage
technologies.
• Hands-on troubleshooting skills across Linux, applications, networks, APIs, and distributed service
dependencies.
• Working knowledge of Voice/UC or real-time communications concepts, including technologies such as SIP,
RTP/SRTP, WebRTC, SBCs, media services, or equivalent communications platforms.
• Experience supporting or integrating AI-enabled services, APIs, or workflows; familiarity with operational
considerations for speech/voice AI, machine-learning services, or LLM-based applications is preferred.
• Ability to use metrics, logs, traces, and service-level indicators to diagnose complex production issues and
drive measurable reliability improvements.
• A proactive approach to spotting problems, areas for improvement, and performance bottlenecks, with strong
cross-functional communication skills.

Diversity, Inclusion, and Equal Opportunity
We hire, promote, and compensate employees based on their ability to perform their job responsibilities, without
regard to race, color, creed, religion, sex, gender, marital status, national origin, ancestry, age, citizenship, physical
or mental disability, sexual orientation, or any other basis protected by applicable law (collectively referred to in our
Code of Conduct as “Protected Classes”). We do not tolerate employment discrimination in the workplace, and we
are committed to making reasonable accommodations for identified disabilities or other limitations as required by all
applicable laws. We are an equal opportunity employer and value diversity at our company. We do not discriminate
on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or
disability status.

Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位