站点可观测性工程师
Site Observability Engineer
Site Observability Engineer- 远程
Bright Vision Technologies 是一家技术咨询和软件开发公司,为美国各地的客户提供云计算、人工智能、数据和企业解决方案。
这是一个加入一家成熟且备受尊敬的组织的绝佳机会,提供巨大的职业发展潜力。
职位名称:Site Observability Engineer
地点:100% 远程(美国)
职位类型:全职,直接W2
薪资范围:每年10万至15万美元
所需经验:6年以上
赞助:美国公民、绿卡持有者、EAD持有者以及H-1B转签候选人欢迎申请。我们无法为该职位的新H-1B签证申请提供赞助。
职位简介
我们正在寻找一名Site Observability Engineer,负责设计和运营指标、日志、追踪和警报平台,使工程团队对其运行的系统充满信心。该角色涵盖完整的可观测性堆栈——从采集代理和管道到长期存储、仪表板和警报工作流,重点关注可用性、信号质量和运营投资回报率。理想的候选人具备在大规模环境中构建和运营可观测性平台的经验,了解开源与SaaS方法之间的权衡,并能将嘈杂的遥测数据转化为工程师和业务利益相关者的可操作见解。
主要职责
· 设计和运营企业级可观测性平台,涵盖指标、日志、追踪、事件和合成监控。
- 架构 Prometheus / Thanos / Mimir、Grafana、Loki、Tempo、OpenTelemetry 和 Datadog 部署,确保高可用性和扩展性。
- 制定服务仪器化的标准,包括 OpenTelemetry 的采用、指标命名、标签基数和结构化日志规范。
- 定义并执行 SLO、SLI 和错误预算,并构建将它们落地的仪表板和警报。
- 构建警报策略,减少噪音,突出可操作信号,并与 PagerDuty、Opsgenie 或类似工具中的值班工作流程无缝集成。
- 运营大规模时间序列和日志存储平台,平衡保留期、查询性能和成本。
- 设计分布式追踪管道,并帮助团队使用追踪来诊断延迟和可靠性问题。
- 开发自助工具、预设道路库和模板,使产品团队更容易采用可观测性标准。
- 推动成本管理和标签基数规范
查看英文原文
Site Observability Engineer- Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Site Observability Engineer
Location: 100% Remote (U.S.)
PositionType: Full-time, Direct W2
Salary Range: $100,000–$150,000 Annually
ExperienceRequired: 6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary
We are looking for a Site Observability Engineer to design and operate the metrics, logging, tracing, and alerting platforms that give engineering teams confidence in the systems they run. The role spans the full observability stack — from collection agents and pipelines to long-term storage, dashboards, and alerting workflows — with a strong focus on usability, signal quality, and operational ROI. The ideal candidate has built and operated observability platforms at scale, understands the trade-offs between open-source and SaaS approaches, and can translate noisy telemetry into actionable insight for both engineers and business stakeholders.
Key Responsibilities
· Design and operate enterprise-grade observability platforms covering metrics, logs, traces, events, and synthetic monitoring.
- Architect Prometheus / Thanos / Mimir, Grafana, Loki, Tempo, OpenTelemetry, and Datadog deployments for high availability and scale.
- Develop standards for service instrumentation, including OpenTelemetry adoption, metric naming, label cardinality, and structured logging conventions.
- Define and enforce SLOs, SLIs, and error budgets, and build the dashboards and alerts that operationalize them.
- Build alerting strategies that minimize noise, surface actionable signals, and integrate cleanly with on-call workflows in PagerDuty, Opsgenie, or similar tools.
- Operate large-scale time-series and log storage platforms, balancing retention, query performance, and cost.
- Design distributed tracing pipelines and help teams use traces to diagnose latency and reliability issues.
- Develop self-service tooling, paved-road libraries, and templates that make adoption of observability standards easy for product teams.
- Drive cost management and label-cardinality discipline across the observability estate.
- Lead incident response readiness improvements through better dashboards, alerting hygiene, and post-incident analysis tooling.
- Partner with SRE and platform teams to integrate observability into deployment pipelines, canary analysis, and progressive delivery workflows.
- Evaluate and recommend observability vendors and open-source tools based on cost, capability, and operational maturity.
- Mentor engineering teams on observability fundamentals, debugging techniques, and SLO-driven operations.
- Maintain documentation, onboarding guides, and runbooks for the observability platform.
Required Qualifications
· Bachelor’s degree in Computer Science or a related field.
- Five or more years of experience in SRE, platform engineering, or observability roles.
- Deep hands-on experience with Prometheus, Grafana, and at least one major commercial observability platform such as Datadog, New Relic, or Splunk.
- Strong understanding of OpenTelemetry, distributed tracing, and structured logging.
- Proficiency in at least one general-purpose language such as Go, Python, or Java.
- Experience operating high-cardinality, high-throughput metrics and log pipelines.
- Strong understanding of SLOs, error budgets, and SRE principles.
- Experience integrating observability with CI/CD and incident management tooling.
- Solid grasp of Linux internals, networking, and container platforms.
- Excellent communication and collaboration skills.
Preferred Qualifications
· Experience with Thanos, Mimir, Cortex, Loki, or Tempo at scale.
- Contributions to OpenTelemetry or observability open-source projects.
- Familiarity with eBPF-based observability tooling.
- Experience driving observability cost optimization initiatives.
- Exposure to regulated environments with audit-grade logging requirements.
How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to or contact us at (908) 505-3544. Learn more about Bright Vision Technologies at .
Bright Vision Technologies is an Equal Opportunity Employer.
Equal Employment Opportunity (EEO) Statement
Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.
BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
Originally posted on Himalayas