高级软件工程师,站点可靠性
Sr. Software Engineer, Site Reliability
在Bloomerang,我们相信改变是有所目的的。我们支持非营利组织的力量和潜力,通过为使命而生的团队和技术,激发更高级别的影响力。我们强大的捐赠平台和卓越的支持服务,使数以万计的非营利组织能够筹集更多资金、招募更多人员并留住更多成员,推动最大化的影响力,并提升非营利领域的可能性。这就是为什么,即使非营利领域捐赠减少,Bloomerang的客户仍能逐年增加筹款。
我们也在打造蓬勃发展的员工。加入一个以使命驱动的文化,建立在我们核心价值观——简化、关怀与行动之上。我们知道我们的员工是成功的关键,我们自豪地拥有当今最创新和最有技能的员工之一。与我们一同感受充满活力和不可阻挡的力量吧!
职位描述
我们正在将三级支持工程团队转型为专注于提升产品可靠性、可观测性和运营效率的站点可靠性工程(SRE)组织。我们正在寻找一位经验丰富的SRE,具备扎实的软件工程基础,并热衷于参与这一转型,推动SRE实践的成熟。
这是一项高度协作、动手能力强的工作,涉及应用代码、遥测、数据库、API和基础设施,以诊断复杂的生产问题并提高可靠性。生产支持和产品缺陷仍是当前工作的一部分,你将通过可观测性、SLO、自动化、永久性修复和主动可靠性工程来减少被动工作量。
成功的标准
成功不仅仅通过解决的问题数量来衡量,而是通过不再需要人工干预的问题来衡量。你将帮助更早地发现问题,减少重复性问题和繁琐工作,加强事件响应,并为积极的可靠性工程创造更多空间。
你将负责
- 负责复杂的生产支持升级和工单分类,与可靠性工作一起提供现场排查和解决。
- 与软件工程团队合作,调查复杂的生产问题,识别根本原因和可靠性风险,并推动解决重复性问题和缺陷的永久性方案。
- 为团队带来经过验证的SRE实践,促进主动可靠性、持续改进、自动化和共享责任。
- 从初步分类和缓解到恢复、根本原因分析,领导事件响应。
查看英文原文
At Bloomerang, we believe change happens on purpose. We champion the power and potential of nonprofits, igniting next-level impact with the team and technology built for purpose. Our powerful giving platform and stellar support enable tens of thousands of nonprofits to raise more, recruit more, and retain more, fueling maximum impact and raising the bar on what’s possible for the nonprofit sector. That's why, even as the nonprofit sector sees declines in giving, Bloomerang customers raise more year over year.
We're also in the business of creating thriving employees. Join a mission-driven culture built on our core values of Simplify, Care and Act. We know our people are the key to our success, and we're proud to be home to some of the most innovative and skilled individuals in the workforce today. Come feel invigorated and unstoppable with us!
The Role
We are evolving our Tier 3 Support Engineering team into a Site Reliability Engineering (SRE) organization focused on improving product reliability, observability, and operational efficiency. We're looking for an experienced SRE who brings strong software engineering fundamentals and is excited to help shape this transformation and mature our SRE practices.
This is a highly collaborative, hands-on role working across application code, telemetry, databases, APIs, and infrastructure to diagnose complex production issues and improve reliability. Production support and product defects remain part of today's work as you help reduce reactive effort through observability, SLOs, automation, permanent fixes, and proactive reliability engineering.
What Success Looks Like
Success isn't measured solely by issues resolved, but by issues that no longer require manual intervention. You'll help detect problems earlier, reduce recurring issues and toil, strengthen incident response, and create more capacity for proactive reliability engineering.
What You Will Do
- Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
- Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions to recurring problems and defects.
- Bring proven SRE practices to the team and foster proactive reliability, continuous improvement, automation, and shared ownership.
- Lead incident response from triage and mitigation through recovery, root cause analysis, and blameless post-incident reviews, turning lessons learned into reliability improvements.
- Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.
- Define and mature SLIs and SLOs that measure system reliability and customer experience.
- Develop synthetic monitoring for critical customer journeys to detect failures before they impact customers.
- Identify sources of recurring operational toil and drive automation, tooling, process improvements, or permanent fixes that reduce manual effort.
- Use AI-assisted tools and source code repositories to accelerate triage, troubleshooting, code analysis, automation, and technical investigation.
- Participate in a rotating on-call schedule, primarily during business hours, with limited after-hours and weekend support.
What You Need to Succeed
SRE Experience & Transformation
- Hands-on Site Reliability Engineering experience applying software engineering practices to production reliability and helping establish or mature SRE practices.
- Strong knowledge of SLIs, SLOs, error budgets, observability, automation, and toil reduction.
Observability & Incident Management
- Experience building monitoring, dashboards, alerts, and telemetry using tools such as Honeycomb, New Relic, Grafana, CloudWatch, Kibana, or similar.
- Experience with production incident management, root cause analysis, blameless post-incident reviews, and corrective-action follow-through.
Technical Depth
- Strong programming and scripting skills to navigate and troubleshoot application code and build automation and operational tooling.
- Strong SQL and relational database skills for production troubleshooting and safe data correction; PostgreSQL experience preferred.
- Strong code literacy and debugging skills, including navigating unfamiliar codebases, understanding application flow, reviewing code and change history, and identifying potential reliability issues.
- Experience troubleshooting cloud-hosted applications using source code, logs, APIs, telemetry, event streams, and databases. Comfort navigating application stacks across technologies such as PHP, .NET, and Node.js; deep expertise in each is not required.
AI, Ownership & Collaboration
- Demonstrated experience using and embracing AI-assisted tools in day-to-day engineering workflows, including troubleshooting, code analysis, scripting, and automation.
- A collaborative self-starter who tackles difficult problems, adapts to changing priorities, and challenges the status quo.
- Strong communication and collaboration skills across Software Engineering, Product, Support, DevOps, and other technical teams.
Benefits
Health + Wellness
You’ll have access to generous health, vision, and dental insurance options as well as HealthiestYou, a healthcare service that offers convenient, confidential access to quality doctors 24/7, anytime, anywhere.
Time Off
You'll get a competitive PTO package that includes 20 PTO days, 3 flex days, 4 optional volunteer days, 12 paid holidays, as well as paid parental leave. More is more!
401k
You'll receive a 401k match to help invest in your future.
Equipment
Everything you need to be successful, shipped right to your door. You got this. We got you.
Compensation
The salary range for this position is $114,800 - $150,000. You may also be eligible for a discretionary bonus. Actual compensation within the range will be dependent on your skills, experience, qualifications, and location, as well as applicable employment laws
Location
This is a permanent, full-time, fully remote position (within the U.S. and select Canadian Provinces only). Employees living in Indianapolis, IN are welcome to work from our company headquarters. We do not offer Visa sponsorship or relocation assistance at this time.
Accommodations
Applicants who require accommodations may contact careers@bloomerang.com to request an accommodation in completing an application.
Bloomerang is an Equal Opportunity Employer. Individuals seeking employment at Bloomerang are considered without regard to race, color, religion, national origin, age, sex, marital status, ancestry, physical or mental disability, veteran status, gender identity, or sexual orientation.