资深站点可靠性工程师
Staff Site Reliability Engineer
Filevine 是一家法律人工智能公司,为法律工作的未来提供法律运营智能。基于一个统一的真相系统,Filevine 将数据、文档、工作流程和团队整合到一个统一的平台中——现代法律工作在这里以清晰和一致的方式进行。
由 LOIS(法律运营智能系统)驱动,Filevine 在每个案件中连接上下文,将法律运营从被动转为主动。LOIS 能读取、理解并推理您的数据,以揭示洞察力、自动化复杂性,并为专业人士提供清晰度和信心,让他们看到更多、了解更多、做得更多。由一群卓越的协作者和创新者推动,Filevine 的快速增长使其获得了德勤和 Inc. 的 AI 奖项和认可,成为国内最具创新性和增长最快的科技公司之一。
薪酬信息:240,000 - 280,000 美元
基本薪资范围代表该职位薪资范围的低值和高值。该职位的总薪酬包将根据每位员工的所在地、资格、教育背景、工作经验、技能和绩效来确定。我们重视薪酬公平——所列范围只是 Filevine 员工总薪酬包的一个组成部分。该职位还享有股票期权、带薪休假政策以及全面的福利包。
工作地点期望:远程办公,或在以下地点之一选择混合办公/在办公室办公:旧金山、纽约、芝加哥或盐湖城
酷炫的公司福利:
- 一家充满活力、快速发展的公司,专注于帮助组织蓬勃发展
- 医疗、牙科和视力保险(适用于全职员工)
- 具有竞争力且公平的薪酬
- 产假和陪产假(适用于全职员工)
- 短期和长期残疾保障
- 有机会向专注的领导团队学习
- 顶级的公司周边产品
隐私政策通知
Filevine 将根据我们的隐私政策中所述处理您的个人信息。
关于此机会或其他任何开放职位的信息,只会由使用 "filevine.com" 域名的代表发送。其他地址联系的都不是 Filevine 的成员,不应予以回应。
职位概要
作为 Filevine 的高级站点可靠性工程师,您是 SRE 团队的技术权威,并是推动业务发展的战略合作伙伴。
查看英文原文
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. Grounded in a singular system of truth, Filevine brings together data, documents, workflows, and teams into one unified platform—where modern legal work happens with clarity and consistency.
Powered by LOIS, the Legal Operating Intelligence System, Filevine connects context across every matter to transform legal operations from reactive to proactive. LOIS reads, understands, and reasons across your data to surface insight, automate complexity, and give professionals the clarity and confidence to see more, know more, and do more. Fueled by a team of exceptional collaborators and innovators, Filevine’s rapid growth has earned AI awards and recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Compensation Information: $240,000 - $280,000
The base salary range represents the low and high end of the salary range for this position. The total compensation package for this position will be determined by each individual’s location, qualifications, education, work experience, skills and performance. We believe in the importance of pay equity - the range listed is just one component of Filevine’s total compensation package for employees. This position is also eligible for stock options, a paid time off policy, as well as a comprehensive benefits package.
Work Location Expectation: Remote or option to be hybrid/in-office in one the following locations: San Francisco, New York, Chicago, or Salt Lake City
Cool Company Benefits:
- A dynamic, rapidly growing company, focused on helping organizations thrive
- Medical, Dental, & Vision Insurance (for full-time employees)
- Competitive & Fair Pay
- Maternity & paternity leave (for full-time employees)
- Short & long-term disability
- Opportunity to learn from a dedicated leadership team
- Top-of-the-line company swag
Privacy Policy Notice
Filevine will handle your personal information according to what’s outlined in our Privacy Policy.
Communication about this opportunity, or any open role at Filevine, will only come from representatives with email addresses using "filevine.com". Other addresses reaching out are not affiliated with Filevine and should not be responded to.
Role Summary
As a Staff Site Reliability Engineer at Filevine, you are the senior technical authority on the SRE team
and a strategic partner to engineering leadership. You don’t just maintain systems — you shape
engineering culture, define the technical standard for how Filevine runs in production, and bridge the
gap between high-level business goals and robust, internet-scale technical execution. You bring a
forward-looking perspective — actively shaping how AI and machine learning drive the future of
reliability practice.
You own the roadmap across two critical SRE domains — Observability & Alerting and Platform
Infrastructure — and are accountable for ensuring the team solves reliability problems permanently
rather than absorbing them as toil. You operate as the senior IC counterpart to the Engineering
Manager: technical correctness lives with you. You partner with the Reliability Architect and engineering
leadership on significant technical decisions, mentor engineers across experience levels, and influence
reliability strategy across the broader organization. Reliability at Filevine protects revenue. You are the
senior technical voice responsible for ensuring that uptime, incident response, and every production
change meet the operational standard the business demands.
This role does not participate in on-call rotation, but you are deeply invested in the engineers who do —
shaping the on-call strategy, tooling, and culture that make production support sustainable and
effective.
Who You Are
The Technical Authority
• Master of the Craft: You bring deep expertise in distributed systems, cloud infrastructure,
observability, and reliability engineering. You raise the technical standard for every engineer
around you and thrive where the challenges are complex and the stakes are real.
• Technical Leader and Mentor: You are passionate about mentoring engineers and investing in
their growth. You influence technical direction and communicate production risk clearly across
engineering, product, and executive audiences.
• Forward-Thinking & AI/ML Fluent: You bring deep knowledge of AIOps and drive the use of
AI and machine learning in observability, anomaly detection, incident response, automated
remediation, and resource optimization.
• Production-Scale Problem Solver: You turn ambiguous, complex reliability challenges into
durable solutions for systems where availability, performance, and production changes carry
meaningful business impact.
• Software-Minded Builder: You use software, automation, Infrastructure as Code, and platform
capabilities to eliminate toil and make systems safer, more scalable, and easier to operate.
What you will do
- Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure,
- and operational excellence.
- Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed
- systems.
- Champion SLIs, SLOs, error budgets, capacity planning, operational readiness, and automation
- across the service lifecycle.
- Lead the organization through complex production incidents and turn post-incident learning into
- permanent engineering improvements.
- Build self-service platform capabilities that reduce toil, improve engineering safety and velocity,
- and make every team more capable of owning their own reliability.
- Mentor engineers and serve as a trusted technical authority for long-term reliability and platform
- direction.
Qualifications
- 12+ years of experience in software engineering, infrastructure, platform engineering, or SRE,
- including 6+ years in SRE and 3+ years leading complex, cross-functional technical initiatives
- for distributed production systems.
- Expert-level depth in observability and platform infrastructure, with broad expertise in incident
- response, capacity planning, automation, and reliability engineering.
- Advanced experience with a major container-orchestration platform, preferably Kubernetes, and
- an observability platform such as New Relic, Datadog, or equivalent.
- Strong software-engineering ability in Python, Go, Bash, or another general-purpose language,
- with experience building production tooling, automation, or platform capabilities.
- Proven ability to mentor engineers and communicate technical risk clearly to engineering,
- product, and executive audiences.
- Experience in a regulated environment such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is
- strongly preferred.