远程工作雷达

平台支持工程师

Platform Support Engineer

开发工程全球可投(远程优先企业)
公司Braintrust
薪资未公开
工作地点San Francisco / New York City / Seattle
地域资格全球可投(远程优先企业)
时区要求无特别要求
用工类型FullTime
发布时间6 天前
数据来源Ashby
前往企业招聘页投递 →
全球可投:该职位未限制候选人所在地区。仍需注意薪资可能按地区折算,以及实际签约方式(正式雇佣 / 独立合同)。

关于公司

Braintrust 是一个代理可观测性平台。通过主动将智能应用于代理追踪并自动呈现最关键的趋势,Braintrust 使团队能够了解代理在生产环境中的行为,并提供改进它们的工具。

Notion、Stripe、Box、OpenAI 和 Cloudflare 的团队使用 Braintrust 来追踪他们的代理,发现可观测性数据中的问题,并运行评估以了解如何改进。

关于职位

我们最大的客户不仅使用 Braintrust —— 他们运行它。他们将我们的系统部署在他们自己的 AWS、Azure 和 GCP 账户中,位于他们自己的 VPC 后面,符合他们自己的合规要求,在他们自己的规模下运行。当混合部署停滞时,当数据摄入积压时,当上周还很快的查询现在变慢时,他们会找到我们。平台支持团队负责这些事情。我们是基础设施、性能和可靠性方面的技术前线。

我们正在招聘中高级平台支持工程师加入一个小型、高自主权的团队。你将与我们的云基础设施和工程团队并肩工作,并与我们的开发支持工程师合作,后者负责客户体验的 SDK 和 API 部分。如果你喜欢硬核的基础设施问题,并且更喜欢当真实客户在另一端时的问题,这就是你的职位。

你将负责

- 负责跨 AWS、Azure 和 GCP 的混合和自托管 Braintrust 部署的客户支持 —— 从首次安装到稳定运行状态。

- 调试真实的基础设施问题:Kubernetes 工作负载、Terraform 状态、网络和 VPC 配置、IAM 和权限、TLS 以及云供应商的特殊之处。

- 诊断后端的性能和可靠性问题 —— 数据摄入吞吐量、查询延迟、数据库和对象存储行为 —— 使用日志、指标和追踪来找到原因而不是症状。

- 主导影响客户的事件响应:分类、在情况紧急时清晰沟通,并推动解决问题。

- 提交修复代码。向我们的后端服务、Terraform 模块和部署工具提交 PR,而不是把每个问题都交给工程团队。

- 构建让下一个问题更容易解决的工具 —— 诊断工具、健康检查、预检验证,以及让客户自行解除阻塞的自助路径。

- 编写和维护运行手册和部署文档,将一次艰难获得的答案变成永久解决方案。

- 提供反馈

查看英文原文

ABOUT THE COMPANY

Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them.

Teams at Notion, Stripe, Box, OpenAI, and Cloudflare use Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve.

ABOUT THE ROLE

Our largest customers don't just use Braintrust — they run it. They deploy our stack inside their own AWS, Azure, and GCP accounts, behind their own VPCs, under their own compliance requirements, at their own scale. When a hybrid deployment stalls, when ingest backs up, when a query that was fast last week isn't, they come to us. Platform Support is the team that owns that. We're the technical front line for infrastructure, performance, and reliability.

We're hiring Platform Support Engineers at both mid and senior levels to join a small, high-ownership team. You'll work shoulder to shoulder with our Cloud Infrastructure and Engineering teams, and alongside our Developer Support Engineers, who own the SDK and API side of the customer experience. If you like hard infrastructure problems, and you like them more when a real customer is on the other end, this is the role.

WHAT YOU'LL DO

- Own customer-facing support for hybrid and self-hosted Braintrust deployments across AWS, Azure, and GCP — from first install through steady-state operation.

- Debug real infrastructure problems: Kubernetes workloads, Terraform state, networking and VPC configuration, IAM and permissions, TLS, and cloud-provider quirks.

- Diagnose performance and reliability issues in the backend — ingest throughput, query latency, database and object-store behavior — using logs, metrics, and traces to get to cause rather than symptom.

- Lead incident response for customer-impacting issues: triage, communicate clearly while it's still on fire, and drive it to resolution.

- Ship fixes. Submit PRs to our backend services, Terraform modules, and deployment tooling rather than handing every problem to Engineering.

- Build the tooling that makes the next one easier — diagnostics, health checks, preflight validation, and self-service paths that let customers unblock themselves.

- Write and maintain the runbooks and deployment documentation that turn one hard-won answer into a permanent one.

- Feed patterns back to Engineering and Product, so the recurring failure modes stop recurring.

- Participate in an on-call rotation for critical customer issues.

WHAT WE'RE LOOKING FOR

- Experience in a customer-facing technical role — Support Engineering, SRE, DevOps, Solutions Architecture, or Infrastructure Engineering — or backend/infra engineering experience with real appetite for customer work.

- Strong Kubernetes fundamentals: you can deploy, debug, and scale actual workloads, and read a failing pod's story from its events and logs.

- Hands-on Terraform, and depth in at least one major cloud (AWS strongly preferred).

- Comfort in a backend codebase — Python, TypeScript, or Go — enough to reproduce a bug, trace it to its source, and fix it.

- Fluency with observability tooling, and the instinct to reach for data before opinion.

- Clear, calm, direct communication under pressure, especially when the customer is technical, blocked, and losing time.

- Ownership. You take a problem personally and follow it until the customer is running again.

BONUS POINTS FOR

- Supporting self-hosted or on-prem enterprise software, especially in regulated environments.

- Multi-cloud experience, particularly Azure or GCP alongside AWS.

- Database and data-infrastructure depth — Postgres, ClickHouse, or similar analytical stores.

- Experience with observability, ML infrastructure, or developer platforms.

- Familiarity with LLM APIs and how teams are building and evaluating agents in production.

- Having built support or diagnostic tooling that measurably reduced ticket volume.

WHY JOIN BRAINTRUST

- Work on genuinely hard infrastructure problems, at the scale and pace of the teams building the best AI products in the world.

- Join a team early enough to shape how it operates — its standards, its tooling, and its bar.

- Sit close to both the customer and the code, with the mandate to fix things in either direction.

BENEFITS INCLUDE

- Medical, dental, and vision insurance

- Daily lunch, snacks, and beverages

- Flexible time off

- Competitive salary and equity

- Wifi & cellphone stipend

EQUAL OPPORTUNITY

Braintrust is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

本页面信息整理自 Ashby,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位