系统可靠性开发工程师1级
Site Reliability Developer 1
**这是一个灵活的职位,可以选择在我们的多伦多办公室全职工作,每周混合办公,或远程工作。** **优先考虑位于大多伦多地区(GTA)的候选人,能够每月大约到多伦多办公室工作两天。**
**请注意,该职位需要每2-3周参与一次值班轮班,覆盖周一至周日的10点至22点。** Vena正在寻找一名SRE加入我们的SaaS技术与运营(STO)团队。如果你热爱构建高度可扩展、可靠且自动化的服务,那么这个职位适合你。我们是一个创新的团队,致力于通过利用行业领先的自动化和编排实践,为Vena的SaaS平台提供卓越的客户体验。作为一位站点可靠性开发人员1级,你将成为STO团队的第一道接触点,专注于提升Vena SaaS平台的可观测性、可扩展性、稳定性和安全性。
#### 你将带来的影响
- 支持关键的ITIL流程,包括事件管理、请求管理、问题管理和变更管理。
- 定义并记录运行手册和标准操作流程。
- 处理来自应用支持团队和其他内部利益相关者的运维请求。
- 在定义的SLA范围内进行问题分类和解决,以确保出色的客户体验,并解除其他开发和支持团队的障碍。
- 在服务上线后对其进行维护,通过测量和监控可用性、延迟和整体系统健康状况。
- 识别和排查问题,调查根本原因,并在整个组织中推动修复。
- 与基础设施即代码(IaC)合作,注重持续改进。
- 在敏捷环境中与跨职能团队成员协作,推进功能和实现。
- 作为运营职能的一部分,报告SLA和性能指标。
- 参与值班轮班。
我们使用的工具:
请注意,这仅反映了我们当前技术栈的一部分,随着我们的发展,我们不断演进和重新审视我们的技术栈:
- 通过基础设施即代码(Terraform)、配置即代码(Ansible)和CI/CD(Jenkins)管理的现代AWS云基础设施
- RDS MySQL、Redshift、Redshift Spectrum、MongoDB和Elasticsearch
- Kinesis、SQS和RabbitMQ
- 使用Python编写的DevOps工具
- 后端应用程序使用Java、Dropwizard、Spring Boot和Hibernate编写
- 前端应用
查看英文原文
**This is a flexible position and has the option of working in our Toronto office full time, hybrid throughout the week or working remotely.** **Preference will be given to candidates located in the GTA who are able to attend the Toronto office approximately two days per month.**
**Please note that this role includes participating in an on-call rotation every 2–3 weeks, covering 10am–10pm from Monday through Sunday.** Vena is looking for an SRE to join our SaaS Technology and Operations (STO) team. This role is a match for you if you love building highly scalable, resilient, and automated services. We are an innovative team which aims to provide exceptional customer experience by leveraging best-in-class automation and orchestration practices for Vena's SaaS platform. As a Site Reliability Developer 1, you will be part of the first level of contact into the STO team, focusing on improving the observability, scalability, stability and security of Vena’s SaaS platform.
#### How You'll Make an Impact
- Support key ITIL processes, including Incident management, request management, problem management and change management.
- Define and document runbooks and standard operating procedures.
- Field operational requests from our Application Support team and other internal stakeholders
- Triage and solve issues within defined SLA’s to ensure an excellent customer experience and to unblock other development and support teams
- Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
- Identify and troubleshoot problems, investigate root causes, and champion fixes across the organization.
- Work with infrastructure-as-Code (IaC) with a focus on continuous improvement.
- Collaborate with cross-functional team members on features and implementation within an agile environment.
- Report on SLAs and performance metrics as part of the Operations function.
- Participate in on-call rotation.
What we use:
Please note this reflects only a portion of our current technical stack, and we are constantly evolving and revisiting our stack as we grow:
- A modern AWS cloud infrastructure managed through infrastructure-as-code (Terraform), configuration-as-code (Ansible), and CI/CD (Jenkins)
- RDS MySQL, Redshift, Redshift Spectrum, MongoDB, and Elasticsearch
- Kinesis, SQS, and RabbitMQ
- DevOps tools written in Python
- Back-end applications written using Java, Dropwizard, Spring Boot, and Hibernate
- Front-end applications written using TypeScript, JavaScript, React (Context Api and Hooks), and Redux
- Monitoring with DataDog, and CloudWatch
#### We'd Love to See
- Bachelor’s degree in computer science, Software engineering or equivalent experience
- 2+ years of experience in an IT Operational, DevOps, SRE, or Software Engineering role.
- Experience with cloud computing (AWS and Azure) services and a developing-level of knowledge with the management and setup of cloud infrastructure.
- You can write code - in any language. You have implemented your work in a production environment and can back it up with examples.
- Experience with tools and platforms such as: Ansible, Build/Release Pipelines, Docker, Github, Terraform etc.
- Developing-level of knowledge with distributed systems in the cloud using observability and telemetry for oversight of code deployments and service level objectives (SLOs).
- Developing experience with the operational aspects of software systems using telemetry, centralized logging, and alerting with tools such as: CloudWatch, Datadog, Prometheus, etc.
_The base salary range for this role is 85,000 - 115,000 CAD._
_*Our salaries are tailored to roles, levels and locations. Your individual pay within this range is influenced by factors like work location, skills, experience and education. As you progress in your role, your compensation may adapt, offering flexibility for growth beyond initial levels. For specifics, your recruiter will provide details and address any questions during the hiring process._