站点可靠性工程师
Site Reliability Engineer
网站可靠性工程师(SRE)/专业领域专家(SME)- 计算机系统工程师/架构师将为GEOMAP平台在安全云环境中的可靠性、可扩展性、性能和运营弹性提供高级支持。该职位专注于提升云托管和容器化系统的服务可用性、监控、事件响应、自动化和生产稳定性,以支持美国空军关键任务的地理空间能力。
网站可靠性工程师将与开发、DevSecOps、云、数据库、测试和支持团队合作,识别系统性问题,降低运营风险,并实施工程解决方案以提高平台的长期可靠性。
*此职位取决于合同授予。*
职责
- 为GEOMAP云托管系统和服务提供高级工程支持,以提高可靠性、可用性、性能和可维护性。
- 分析生产问题、重复事件和运营趋势,识别根本原因并推荐持久的纠正措施。
- 支持跨应用程序、基础设施和容器化服务的监控、警报、日志记录和可观测性解决方案的设计与实现。
- 开发并推荐自动化方法,以减少人工操作,提高部署一致性,并增强系统弹性。
- 与软件工程师、DevSecOps工程师、Kubernetes工程师、数据库工程师和生产支持人员合作,提高服务健康状况和发布准备度。
- 支持高优先级运营问题的事件响应、问题管理、服务恢复和事后审查。
- 评估系统性能、容量和可扩展性需求,并提供优化和降低运营风险的建议。
- 协助定义服务可靠性目标、运营指标和支持模型,以确保任务持续运行。
- 参与涉及云环境、CI/CD流水线、容器编排和安全部署模式的基础设施和平台工程工作。
- 支持与可靠性、恢复能力和生产运营相关的架构评审、技术评估和工程分析。
- 开发或完善操作手册、标准操作程序、可靠性工程实践和技术文档。
- 提供
查看英文原文
The Site Reliability Engineer (SRE) / Subject Matter Expert (SME) – Computer Systems Engineer/Architect will provide senior-level reach-back expertise to support the reliability, scalability, performance, and operational resilience of the GEOMAP platform in secure cloud environments. This role focuses on improving service availability, monitoring, incident response, automation, and production stability across cloud-hosted and containerized systems supporting mission-critical geospatial capabilities for the U.S. Air Force.
The Site Reliability Engineer will collaborate across development, DevSecOps, cloud, database, testing, and support teams to identify systemic issues, reduce operational risk, and implement engineering solutions that improve long-term platform reliability.
*This position is contingent upon contract award.*
Responsibilities
- Provide senior-level engineering support to improve reliability, availability, performance, and maintainability of GEOMAP cloud-hosted systems and services.
- Analyze production issues, recurring incidents, and operational trends to identify root causes and recommend durable corrective actions.
- Support the design and implementation of monitoring, alerting, logging, and observability solutions across applications, infrastructure, and containerized services.
- Develop and recommend automation approaches that reduce manual effort, improve deployment consistency, and increase system resilience.
- Partner with software engineers, DevSecOps engineers, Kubernetes engineers, database engineers, and production support personnel to improve service health and release readiness.
- Support incident response, problem management, service restoration, and post-incident reviews for high-priority operational issues.
- Evaluate system performance, capacity, and scalability needs and provide recommendations for optimization and operational risk reduction.
- Assist in defining service reliability objectives, operational metrics, and support models for sustained mission operations.
- Contribute to infrastructure and platform engineering efforts involving cloud environments, CI/CD pipelines, container orchestration, and secure deployment patterns.
- Support architecture reviews, technical assessments, and engineering analyses related to reliability, recoverability, and production operations.
- Develop or refine runbooks, standard operating procedures, reliability engineering practices, and technical documentation.
- Provide reach-back support for surge requirements, complex production investigations, and priority modernization or stabilization efforts as directed.
- Performs other related duties as assigned.
Qualifications
- Active Secret clearance required.
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field; Master’s degree preferred.
- Minimum of 8 years of experience supporting enterprise systems, cloud platforms, site reliability engineering, production engineering, systems engineering, or related technical roles.
- Experience supporting AWS environments, including monitoring, performance tuning, troubleshooting, incident response, and operational sustainment.
- Experience with Linux administration, scripting, and troubleshooting distributed applications in production environments.
- Experience with containerized systems and orchestration platforms such as Kubernetes.
- Experience supporting CI/CD pipelines, release automation, infrastructure-as-code, and operational reliability in Agile or DevSecOps environments.
- Experience with monitoring, logging, and alerting tools used to support enterprise application performance and infrastructure visibility.
- Strong analytical, troubleshooting, documentation, and communication skills, with the ability to translate operational issues into engineering improvements.
- Ability to work effectively across cross-functional teams in a mission-focused DoD environment.
Preferred
- Experience supporting AWS Cloud One or other secure federal cloud environments.
- Experience supporting geospatial or Esri-based platforms, including ArcGIS Enterprise or related technologies.
- Familiarity with service reliability practices such as SLIs, SLOs, error budgets, incident postmortems, and capacity planning.
- Experience with Risk Management Framework (RMF), STIG compliance, vulnerability remediation, and secure system hardening practices.
- AWS, Kubernetes, or other relevant cloud or reliability engineering certifications.
- Experience supporting technical refresh, platform modernization, or high-availability design initiatives in enterprise environments.
About Us
Diné Development Corporation (DDC) is a Navajo Nation owned family of companies that provides government agencies and commercial organizations with high-quality IT, professional, environmental, and research and development services. DDC is dedicated to empowering the Navajo Nation and communities we serve.
Benefits
Eligible full-time employees receive a comprehensive benefits package, including medical, dental, vision, life and disability coverage, retirement savings with company match, paid time off, voluntary supplemental benefits, and access to an employee assistance program. The package also includes educational assistance, with tuition reimbursement.
EEO Statement
This contractor and subcontractor shall abide by the requirements of 41 CFR 60-1.4(a), 60-300.5(a), and 60-741.5(a). These regulations prohibit discrimination against qualified individuals based on their status as protected veterans or individuals with disabilities, and prohibit discrimination against all individuals based on their race, color, religion, sex, sexual orientation, gender identity, national origin, or for inquiring about, discussing, or disclosing information about compensation, or any other basis prohibited by law. We participate in E-Verify.
Originally posted on Himalayas