站点可靠性工程师(SRE)– 白班
Site Reliability Engineer (SRE) – Day Shift
职责
Peraton 正在寻找一名站点可靠性工程师(SRE)加入一个负责在 AWS 商业和 AWS GovCloud 环境中运行的生产系统操作可靠性的团队。理想的候选人应具备在 Azure 和 GCP 中部署和管理系统组件的经验。主要的生产工作负载运行在 Red Hat OpenShift Service on AWS(ROSA)上
SRE 与平台工程师、安全团队和应用开发人员紧密合作,确保基础设施服务的可靠性、可用性,并以符合开发人员和安全要求的方式进行部署
工作地点:远程
班次安排:这是白班职位,工作时间为美国东部时间(EST)上午7点至下午3点
你将负责:
- 运营和维护生产基础设施服务和应用程序,确保其可用性、可靠性、性能、安全性和运营健康状况
- 使用定义的 SLI、SLO、仪表板、警报和其他可观测性工具监控服务和应用程序;持续改进操作问题的检测、诊断和解决
- 与应用团队合作,定义应用可观测性需求,并在组织的可观测性工具中实现适当的指标、日志、追踪、仪表板和警报
- 管理生产事件和服务中断,包括值班响应、故障排查、服务恢复、根本原因分析和事件后纠正措施
- 通过已建立的部署流水线执行应用和基础设施发布,包括通过测试环境和生产环境的发布、验证、回滚和与发布相关的故障排除
- 管理已部署基础设施的运营生命周期,包括升级、补丁、配置更改、维护和技术更新
- 通过容量规划、性能测试、故障模式分析、灾难恢复、备份、故障转移和恢复测试来评估和提高服务的弹性
- 通过使用可靠性指标、事件趋势、容量数据和服务健康指标来识别和解决可靠性风险和操作技术债务,以优先改进
- 使用一切皆代码的方法自动化操作活动,以提高一致性、可重复性、测试、部署、恢复和操作效率
- 与平台工程和应用团队协作,识别操作性问题
查看英文原文
Responsibilities
Peraton is seeking a Site Reliability Engineer (SRE) to join a team respobsible for the operational reliability of production systems running in AWS Commercial and AWS GovCloud environments. The ideal candidate has working knowledge of deploying and managing system components in Azure and GCP. The primary production workload runs on Red Hat OpenShift Service on AWS (ROSA)
The SRE partners closely with platform engineers, the security team, and application developers to ensure the infrastructure services are reliable, available, and deployed in a way that meets both developer and security requirements.
Work Location: Remote
Shift Schedule: This is a Day Shift position with working hours from 7am – 3pm Eastern Standard Time (EST)
What you will do:
- Operate and maintain production infrastructure services and applications, ensuring availability and reliability, performance, security, and operational health.
- Monitor services and applications using defined SLIs, SLOs, dashboards, alerts, and other observability tools; continuously improve the detection, diagnosis, and resolution of operational issues.
- Partner with application teams to define application observability requirements and implement appropriate metrics, logs, traces, dashboards, and alerts into the organization's observability tooling.
- Manage production incidents and service disruptions, including on-call response, troubleshooting, service restoration, root-cause analysis, and post-incident corrective actions.
- Execute application and infrastructure releases through established deployment pipelines, including promotion through staging and production, validation, rollback, and release-related troubleshooting.
- Manage the operational lifecycle of deployed infrastructure, including upgrades, patching, configuration changes, maintenance, and technology refreshes.
- Assess and improve service resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing.
- Identify and address reliability risks and operational technical debt by using reliability metrics, incident trends, capacity data, and service health indicators to prioritize improvements.
- Automate operational activities using an everything-as-code approach to improve consistency, repeatability, testing, deployment, recovery, and operational efficiency.
- Collaborate with platform engineering and application teams to identify operational requirements, provide feedback on reusable infrastructure building blocks, and continuously improve the reliability and operability of the environment.
Qualifications
Basic Qualifications:
- Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level clearance.
- Bachelor's Degree and 8 years of experience, or a High School diploma/equivalent and 12 years of experience.
- 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering.
- Hands-on experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or comparable Kubernetes-based platforms
- Strong infrastructure-as-code experience with Terraform and Ansible/Ansible Tower.
- Experience with CI/CD platforms GitLab and Jenkins, including reliability gating and deployment automation.
- Proficient in Linux and Windows Server administration
- Experience with enterprise observability tools such as Dynatrace, Datadog, Splunk and Open Telemetry.
- Demonstrated ownership of an SLI/SLO and alerting program, including error budgets, alert rationalization, and noise reduction.
- Scripting/automation proficiency in Python, Bash, PowerShell, or Go.
- Experience operating in federal or regulated environments (FISMA, FedRAMP, NIST 800-53).
Preferred Qualifications:
- AWS Solutions Architect, AWS DevOps Engineer, or AWS SysOps certification
- Red Hat Certified Specialist in ROSA, Red Hat Certified System Administrator in OpenShift
- Azure Administrator Associate, GCP Associate Cloud Engineer certification
- Dynatrace Associate, Datadog Log Management Fundamentals certification
- GitLab CI/CD Associate certification, Certified Jenkins Engineer (CJE)
- Terraform Associate certification
Peraton Overview
Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees do the can’t be done by solving the most daunting challenges facing our customers. Visit peraton.com to learn how we’re keeping people around the world safe and secure.
Target Salary Range
$104,000 - $166,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual’s experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.EEO
EEO: Equal opportunity employer, including disability and protected veterans, or other characteristics protected by law.Originally posted on Himalayas