高级站点可靠性工程师 (Fedramp)
Senior Site Reliability Engineer (Fedramp)
关于Coalfire
Coalfire致力于通过解决客户最困难的网络安全挑战,让世界变得更安全。我们处于技术的最前沿,为客户提供咨询、评估、自动化服务,最终帮助公司应对不断变化的网络安全环境。我们总部位于伊利诺伊州芝加哥,美国和英国各地都有办公室,并为全球客户提供支持。
但这只是我们的工作内容——这并不是我们是谁。
我们是思想领袖、顾问和网络安全专家,但最重要的是,我们是一支充满热情的问题解决者团队,他们渴望学习、成长并带来改变。
为什么加入我们
FedRAMP现在将合规性定义为一组关键安全指标(KSIs);可量化的安全结果,由自动化持续验证。每天保持在该标准上是一项运维工作。Coalfire将交付工程组织成以能力为导向的团队,而运行团队在建设团队离开后,仍能确保授权的客户环境可用且可证明符合规范。作为高级站点可靠性工程师,您将负责一个运营能力,例如修复、监控与警报或备份与恢复,并设计Coalfire如何观察客户环境并使其保持合规状态。您将负责自动化和每个受管理环境所继承的流程,您将处理最困难的运维问题,并为周围的工程师设定并持续提高工程标准。如果您渴望创新,擅长追求卓越的运维,并在协作环境中茁壮成长,欢迎加入一个致力于让世界更安全、更合规的团队。
您将负责
- 负责受管理环境的一个运营能力,例如监控与警报或备份与恢复,包括其自动化、运行手册以及每个受管理环境所继承的服务标准。
- 设计Coalfire如何观察受监管的云环境:遥测和日志管道、服务级别目标、警报质量以及背后的升级路径。使警报变得有用且可操作。
- 构建和维护持续监控证据管道,以便每天以机器可读的形式证明授权状态,在框架允许的情况下。
- 负责您所负责能力的备份与恢复工程:经过测试的恢复流程、可衡量的恢复目标以及运行它们的自动化。
查看英文原文
About Coalfire
Coalfire is on a mission to make the world a safer place by solving our clients’ hardest cybersecurity challenges. We work at the cutting edge of technology to advise, assess, automate, and ultimately help companies navigate the ever-changing cybersecurity landscape. We are headquartered in Chicago, Illinois with offices across the U.S. and U.K., and we support clients around the world.
But that’s not who we are – that’s just what we do.
We are thought leaders, consultants, and cybersecurity experts, but above all else, we are a team of passionate problem-solvers who are hungry to learn, grow, and make a difference.
Why Join Us
FedRAMP now defines compliance as a set of Key Security Indicators (KSIs); measurable security outcomes that automation validates continuously. Holding an estate at that standard day after day is an operations job. Coalfire organizes its delivery engineering into capability-focused teams, and the Run teams keep authorized client environments available and provably compliant long after the build team leaves. As a Senior Site Reliability Engineer you own one operational capability, such as remidation, monitoring and alerting or backup and recovery, and you design how Coalfire observes a client estate and holds it in a compliant state. You own the automation and the procedure that each managed environment inherits, you carry escalation for the hardest operational problems, and you set and continually raise the engineering bar for the engineers around you. If you are driven by a desire to innovate, excel at operational excellence, and thrive in a collaborative environment, come be part of a team committed to making the world a more secure and complaint place.
What You'll Do
- Own one operational capability for the managed estate, such as monitoring and alerting or backup and recovery, including its automation, its runbooks, and the service standard each managed environment inherits.
- Design how Coalfire observes a regulated cloud environment: telemetry and log pipelines, service-level objectives, alert quality, and the escalation paths behind them. Make alerts useful and actionable.
- Build and maintain the continuous-monitoring evidence pipeline so you can prove the authorized state each day, in machine-readable form where the framework allows it.
- Own backup and recovery engineering for your capability: recovery procedure you have tested, recovery objectives you can measure, and the automation that runs them during an outage.
- Serve as the senior escalation point in client environments. Diagnose and troubleshoot beyond the runbook, resolve the incident, then fix the runbook so the next engineer on call does not repeat the diagnosis.
- Lead incident and problem management for your domain, including incident command on major events, blameless post-incident review, and corrective action that closes the problem out.
- Automate operational toil out of the estate with infrastructure-as-code, pipelines, and scripting. Decide what Coalfire automates once for the whole estate and what stays specific to one client.
- Partner with Engagement Architects and the Build teams on transition into managed operations: operational readiness review, monitoring and runbook coverage, and the service commitments Coalfire can meet at go-live.
- Hold on-call for your capability and improve the rotation: coverage, alert actionability, and the load it places on the team.
- Represent operational posture in front of clients, and support renewals and expansions with a credible account of what Coalfire operates for them and at what cost.
- Mentor and lead Site Reliability Engineers and junior staff: review designs as well as changes, set operational standards, and grow named individuals in your domain.
- Author and peer review code, runbooks, operational design documentation, and the compliance artifacts that evidence continuous monitoring, inclusive of vendor best practices.
What You'll Bring
- Automation-first mindset with deep Infrastructure-as-Code (Terraform or equivalent), CI/CD, and scripting (Python, Go, or similar); working use of policy-as-code
- Deep operational command of at least one major cloud platform (AWS, Azure, or GCP) and its native monitoring, logging, and recovery services, with working knowledge of a second
- Observability engineering depth: metrics, logging and log pipelines, distributed tracing, SLI and SLO definition, and alert design that produces action
- Demonstrated incident response and incident command capability, and the discipline to convert a resolved incident into automation or procedure
- Backup, recovery, and resilience engineering, including tested recovery procedure against stated recovery objectives
- Working command of NIST 800-53, FedRAMP, or comparable security control frameworks, and the judgment to map continuous-monitoring obligations to real operational design
- Ability to lead technical conversations with clients about operational posture, risk, and trade-offs with both engineers and executives
- The instinct to solve a problem once and package the solution so it works across the managed estate
- Demonstrated ability to mentor engineers and improve the output of a team
- Excellent communication, organizational, and problem-solving skills
- Effective documentation skills, including technical diagrams, runbooks, and written descriptions
- Ability to work independently and as part of a team with a professional attitude and demeanor
- Critical thinking, and the ability to balance security and availability requirements against mission needs
- BS or above in a related Information Technology field or equivalent combination of education and experience
- 5+ years in site reliability engineering, cloud operations, platform engineering, or managed services
- 5+ years operating production cloud environments in AWS, Azure, or GCP, including monitoring, incident response, and automation
- Demonstrated experience owning an operational capability, monitoring platform, or reusable automation used by multiple teams or clients
- Experience as the senior operational escalation point on client-facing managed services, including incident command on major events
- Advanced experience with Infrastructure-as-Code and orchestration/automation tools such as Terraform and Ansible
- Experience transitioning environments from build into steady-state operations
REQUIRED CERTIFICATIONS:
- Professional- or specialty-level certification in AWS, Azure, or GCP (associate-level considered with equivalent demonstrated depth)
EDUCATION:
Bachelor’s degree (four-year college or university) or equivalent combination of education and work experience.
Bonus Points
- Direct familiarity with FedRAMP Rev 5 and 20x, Key Security Indicators (KSIs), and continuous monitoring obligations
- Experience with OSCAL or JSON machine-readable compliance formats
- Recognized depth in a specialty area: SIEM and log pipelines, observability platforms, vulnerability and patch management at scale, or disaster-recovery engineering
- Experience operating a monitored compliance environment against contractual service levels
- Relevant certifications such as cloud security or DevOps specialty certifications, CISSP, or GIAC
- Previous experience supporting clients from within a professional services or managed services organization
- Experience contributing to renewals, expansions, and pre-sales technical solutioning for managed services
- Familiarity with configuration baseline standards such as CIS Benchmarks and DISA STIG
- Familiarity with frameworks such as FedRAMP, FISMA, HIPAA, HITRUST, or PCI
Why You’ll Want to Join Us
At Coalfire, you’ll find the support you need to thrive personally and professionally. In many cases, we provide a flexible work model that empowers you to choose when and where you’ll work most effectively – whether you’re at home or an office.
Regardless of location, you’ll experience a company that prioritizes connection and wellbeing and be part of a team where people care about each other and our communities. You’ll have opportunities to join employee resource groups, participate in in-person and virtual events, and more. And you’ll enjoy competitive perks and benefits to support you and your family, like paid parental leave, flexible time off, certification and training reimbursement, digital mental health and wellbeing support membership, and comprehensive insurance options.
At Coalfire, equal opportunity and pay equity is integral to the way we do business. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran. Coalfire is committed to providing access, equal opportunity, and reasonable accommodation for individuals with disabilities in employment, its services, programs, and activities. To request reasonable accommodation to participate in the job application or interview process, contact our Human Resources team at .
Originally posted on Himalayas