远程工作雷达

人工智能基础设施工程师

AI Infrastructure Engineer

AI开发工程限定地区(需当地身份)
公司Bright Vision Technologies
薪资$100,000 - $160,000/年
工作地点United States
地域资格限定地区(需当地身份)
时区要求日间重叠约 9 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 United States 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

AI基础设施工程师 – 远程办公

Bright Vision Technologies是一家技术咨询和软件开发公司,为美国各地的客户提供云、AI、数据和企业解决方案。
这是一个加入一家知名且受人尊敬的组织的绝佳机会,提供巨大的职业发展潜力。

职位名称:AI基础设施工程师
工作地点:100%远程(美国)
职位类型:全职,直接W2雇佣
薪资范围:每年10万至16万美元
所需经验:10年以上

赞助:美国公民、绿卡持有者、EAD持有者以及H-1B转签候选人欢迎申请。我们无法为该职位的新H-1B签证申请提供赞助。

职位简介:我们正在寻找一名AI基础设施工程师,负责设计、构建和运营支持大规模AI训练和推理工作负载的平台层。该职位专注于GPU集群、分布式训练框架、调度、存储性能以及机器学习工程师和研究人员的开发体验,强调可靠性、效率和成本控制。理想的候选人具备在大规模环境下构建或运营生产级AI基础设施的经验,了解硬件、内核、调度器和ML框架之间的交互,并为平台工作带来强大的软件工程规范。

主要职责
· 设计和运营用于训练和推理的GPU和加速器基础设施,涵盖本地集群、云托管服务和混合配置。

  • 构建调度、队列和资源共享系统,以最大化多个团队之间的加速器利用率。
  • 将PyTorch、JAX、DeepSpeed、FSDP、Megatron-LM和Ray Train等框架集成到统一的平台中。
  • 运营高性能存储系统和数据流水线,确保加速器以接近线速获取训练数据。
  • 设计支持RDMA、InfiniBand、NCCL和高带宽集体通信的网络架构。
  • 为AI工作负载构建可观测性功能,包括利用率、吞吐量、训练稳定性以及故障模式分析。
  • 实现检查点、重启和容错模式,以支持大规模长时间运行的训练任务。
  • 通过调度、按需容量和合理配置,推动计算、存储和网络的成本优化。
  • 开发开发者工具和预设流程,让研究人员能够安全高效地启动实验。
  • 与研究和应用机器学习团队合作
查看英文原文

AI Infrastructure Engineer – Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: AI Infrastructure Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $100,000–$160,000 Annually
Experience Required: 10+ years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary:We are seeking an AI Infrastructure Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads. The role focuses on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with strong emphasis on reliability, efficiency, and cost control. The ideal candidate has built or operated production AI infrastructure at scale, understands the interaction between hardware, kernel, scheduler, and ML framework, and brings strong software engineering discipline to platform work.

Key Responsibilities
· Design and operate GPU and accelerator infrastructure for training and inference, spanning on-prem clusters, cloud-managed services, and hybrid configurations.

  • Build scheduling, queueing, and resource-sharing systems that maximize accelerator utilization across many teams.
  • Integrate frameworks such as PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform offering.
  • Operate high-performance storage systems and data pipelines that keep accelerators fed with training data at near-line-rate.
  • Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth collective communication.
  • Build observability for AI workloads including utilization, throughput, training stability, and failure-mode analytics.
  • Implement checkpointing, restart, and fault-tolerance patterns for long-running training jobs at scale.
  • Drive cost optimization across compute, storage, and networking through scheduling, spot capacity, and right-sizing.
  • Develop developer tooling and paved-road workflows that let researchers launch experiments safely and efficiently.
  • Partner with research and applied ML teams to plan capacity for upcoming training runs.
  • Implement security controls, isolation, and access management for multi-tenant AI infrastructure.
  • Drive automation across cluster provisioning, lifecycle management, and configuration enforcement.
  • Maintain runbooks, capacity dashboards, and operational documentation for the AI platform.
  • Stay current with AI infrastructure research, accelerator hardware, and emerging open-source AI tooling.

Required Qualifications
· Bachelor’s or Master’s degree in Computer Science or a related field.

  • Ten or more years of experience in infrastructure, platform, or HPC engineering.
  • Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
  • Strong proficiency in Python and at least one systems language such as Go or C++.
  • Deep understanding of distributed training, accelerator architectures, and collective communication.
  • Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads.
  • Strong understanding of Linux internals, networking, and high-performance storage.
  • Experience with at least one major cloud provider’s ML infrastructure offerings.
  • Strong software engineering practices including testing, CI/CD, and code review.
  • Excellent communication and cross-functional collaboration skills.

Preferred Qualifications
· Experience operating InfiniBand or RDMA networking at scale.

  • Contributions to open-source ML infrastructure projects.
  • Familiarity with custom orchestrators or research-grade training stacks.
  • Exposure to frontier model training operations.
  • Experience with FinOps for AI workloads.

How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to or contact us at (908) 505-3899. Learn more about Bright Vision Technologies at .
Bright Vision Technologies is an Equal Opportunity Employer.
Equal Employment Opportunity (EEO) Statement
Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.
BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

云网络工程师

Bright Vision TechnologiesUnited States$100,000 - $150,000/年Full Time今天
开发工程限定地区(需当地身份)

SAP MM顾问

Bright Vision TechnologiesUnited States$130,000 - $190,000/年Full Time今天
其他限定地区(需当地身份)

React开发工程师

Bright Vision TechnologiesUnited States$120,000 - $180,000/年Full Time今天
开发工程限定地区(需当地身份)

Windchill后端开发工程师

Bright Vision TechnologiesUnited States$120,000 - $135,000/年Full Time今天
开发工程限定地区(需当地身份)

SAP HANA 平台工程师

Bright Vision TechnologiesUnited States$100,000 - $150,000/年Full Time今天
开发工程限定地区(需当地身份)

← 返回全部职位