远程工作雷达

分布式系统工程师

Distributed Systems Engineer

开发工程限定地区(需当地身份)
公司ThisWay
薪资未公开
工作地点United States
地域资格限定地区(需当地身份)
时区要求日间重叠约 9 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 United States 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

ThisWay Global 正在寻找一名远程职位的分布式系统工程师,工作地点在美国
ThisWay Global, Inc. 是一家以人工智能为核心的科技公司,总部位于德克萨斯州,业务涉及人工智能、数据中心基础设施和人力资源解决方案的交叉领域
公司主要运营三个业务领域:

  • ADCAP — AI 与数据中心加速平台:一个用于数据中心开发和运营的平台,支持加快部署时间以及 NVIDIA NVL72/GB300 GPU 集群
  • Amalgamy.ai — AI 协调软件:一个企业级 AI 协调平台,专注于 AI 环境中的 GPU 和计算资源利用率
  • 招聘与人力资源解决方案:利用人工智能进行人才匹配的解决方案,可大规模连接雇主与候选人

该职位专注于构建支持 AI 和艾克萨级计算环境的基础分布式系统和运维基础设施。工作重点包括系统编程、分布式架构、容错性和高性能计算级别的可靠性
地点:远程 – 美国
部门:工程部
雇佣类型:全职、非豁免职位
职责

  • 设计和构建能够容忍延迟、带宽限制和间歇性连接的分布式系统
  • 实现具有容错能力的通信策略、重试逻辑、背压、缓存和最终一致性模式
  • 按照开发标准和方法编写可维护、有弹性且经过测试的代码
  • 调试和改进系统行为,包括网络和分布式协调问题
  • 参与主要使用 Rust 编写的系统开发
  • 处理系统级问题,包括调度、内存管理、I/O 优化、存储层次管理以及系统可靠性
  • 在资源受限环境中优化性能和内存使用
  • 调试并发问题和分布式协调挑战
  • 设计和维护分布式组件之间的 API 和通信层
  • 识别并减少服务和系统之间的紧密耦合
  • 诊断和解决生产环境中的跨系统故障
  • 设计和实现符合工程标准的安全、可靠的解决方案
  • 与工程师和计算机科学家合作,研究操作系统内部、编译器内部、容错性、文件系统架构和可信系统
  • 参与解决架构和系统性问题
查看英文原文

ThisWay Global is looking for a Distributed Systems Engineer in a remote role within the United States.
ThisWay Global, Inc. is an AI-first technology company headquartered in Texas, operating at the intersection of artificial intelligence, data center infrastructure, and workforce solutions.
The company operates across three primary business areas:

  • ADCAP — AI & Data Center Acceleration Platform: A data center development and operations platform supporting accelerated deployment timelines and NVIDIA NVL72/GB300 GPU clusters.
  • Amalgamy.ai — AI Orchestration Software: An enterprise AI orchestration platform focused on GPU and compute utilization across AI environments.
  • Staffing & Workforce Solutions: AI-powered talent matching solutions connecting employers with candidates at scale.

This role focuses on building foundational distributed systems and operational infrastructure that support AI and exascale computing environments. The work emphasizes systems programming, distributed architecture, fault tolerance, and HPC-grade reliability.
Location: Remote – United States
Department: Engineering
Employment Type: Full-Time, Exempt
Responsibilities

  • Design and build distributed systems that tolerate latency, bandwidth constraints, and intermittent connectivity.
  • Implement fault-tolerant communication strategies, retry logic, backpressure, caching, and eventual consistency patterns.
  • Write maintainable, resilient, and tested code following development standards and methodologies.
  • Debug and improve system behavior, including networking and distributed coordination issues.
  • Contribute to systems written primarily in Rust.
  • Work with system-level concerns including scheduling, memory management, I/O optimization, storage hierarchy management, and system reliability.
  • Optimize performance and memory usage in resource-constrained environments.
  • Debug concurrency issues and distributed coordination challenges.
  • Design and maintain APIs and communication layers between distributed components.
  • Identify and reduce tight coupling across services and systems.
  • Diagnose and resolve cross-system failures in production environments.
  • Design and implement secure, reliable solutions aligned with engineering standards.
  • Collaborate with engineers and computer scientists on operating systems internals, compiler internals, fault tolerance, file system architecture, and trusted systems.
  • Contribute to resolving architectural and systemic issues.
  • Continue developing expertise in distributed systems, HPC infrastructure, and related tooling.

Requirements

  • Experience building distributed systems in environments with low bandwidth, high latency, or unreliable communication links.
  • Production experience developing systems in Rust or Go.
  • Understanding of distributed systems failure modes and mitigation strategies.
  • Knowledge of consistency models, coordination strategies, and state replication.
  • Experience designing APIs and communication layers between distributed components.
  • Experience working within established architectures and delivering production-quality components.
  • Understanding of systems-level concepts including durability, reliability, and operational behavior.
  • Ability to work independently while collaborating with technical leadership.

Preferred Qualifications

  • Experience with HPC environments, exascale computing, or AI/ML infrastructure.
  • Exposure to operating systems internals, compiler design, or language runtimes.
  • Experience with edge computing or constrained network environments.
  • Familiarity with message queues, event-driven systems, or streaming architectures.
  • Exposure to consensus algorithms or distributed coordination primitives.
  • Experience with concurrency, memory management, or performance optimization in production systems.
  • Experience contributing to developer tooling, internal platforms, or infrastructure-layer components.

Benefits

  • Remote work within the United States.
  • Opportunity to work on distributed systems supporting AI and exascale workloads.
  • Collaboration with engineers experienced in operating systems internals, compiler internals, fault tolerance, file system architecture, and trusted systems.
  • Exposure to AI infrastructure, HPC, and large-scale distributed computing environments.

Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

← 返回全部职位