人工智能系统性能专员
AI Systems Performance Specialist
Bright Vision Technologies 是一家技术咨询和软件开发公司,为美国各地提供云、AI、数据和企业解决方案。加入一家成熟且备受尊敬的组织,这是一个绝佳的机会,提供巨大的职业发展潜力。
职位名称
AI 系统性能专家
地点:100% 远程(美国本土)
职位类型:全职,直接 W2
薪资范围:每年 13 万至 18 万美元
经验要求:10+ 年
赞助:美国公民、绿卡持有者、EAD 持有者以及 H-1B 转签候选人欢迎申请。我们无法为该职位新申请 H-1B 签证提供支持。
职位简介
Bright Vision Technologies 正在寻找一位拥有 10+ 年经验的 AI 系统性能专家,专注于 AI 基础设施、机器学习系统、高性能计算(HPC)和性能工程。理想的候选人将优化 AI 训练和推理工作负载,以实现最大性能、可扩展性、可靠性和成本效率。该职位需要在 GPU 优化、分布式训练、大语言模型(LLM)推理、Python、C++、CUDA 和生产级 AI 系统方面具备深厚的专业知识,并能够领导企业级 AI 平台上的性能优化项目。
主要职责
- 优化 AI 训练和推理流程,以实现最大吞吐量、低延迟、可扩展性和基础设施效率。
- 分析并提升 GPU 利用率、内存管理、内核执行和多 GPU 性能,针对生产级 AI 工作负载。
- 设计并实现优化技术,包括量化、剪枝、混合精度、批处理、缓存、推测解码和模型并行。
- 使用行业标准的性能分析工具对 AI 应用进行剖析,并识别计算、内存、网络和存储方面的瓶颈。
- 使用 NCCL、DeepSpeed、PyTorch Distributed、Ray、MPI 或类似的分布式计算框架优化分布式训练和推理。
- 与 AI 研究员、ML 工程师、平台工程师和基础设施团队合作,提升模型性能和生产可靠性。
- 构建自动化基准测试框架、性能仪表板、监控解决方案和回归测试流程。
- 评估新兴 AI 硬件、GPU 架构、推理框架和优化技术,以提升企业级 AI 能力。
查看英文原文
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title
AI Systems Performance Specialist
Location: 100% Remote (Continental United States)
Position Type: Full-time, Direct W2
Salary Range: $130,000–$180,000 Annually
Experience Required:10+ Years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary
Bright Vision Technologies is seeking a highly experienced AI Systems Performance Specialist with 10+ years of experience in AI infrastructure, machine learning systems, High-Performance Computing (HPC), and performance engineering. The ideal candidate will optimize AI training and inference workloads for maximum performance, scalability, reliability, and cost efficiency. This role requires deep expertise in GPU optimization, distributed training, Large Language Model (LLM) inference, Python, C++, CUDA, and production AI systems, along with the ability to lead performance optimization initiatives across enterprise-scale AI platforms.
Key Responsibilities
- Optimize AI training and inference pipelines for maximum throughput, low latency, scalability, and infrastructure efficiency.
- Analyze and improve GPU utilization, memory management, kernel execution, and multi-GPU performance across production AI workloads.
- Design and implement optimization techniques including quantization, pruning, mixed precision, batching, caching, speculative decoding, and model parallelism.
- Profile AI applications using industry-standard performance analysis tools and identify bottlenecks across compute, memory, networking, and storage.
- Optimize distributed training and inference using NCCL, DeepSpeed, PyTorch Distributed, Ray, MPI, or similar distributed computing frameworks.
- Collaborate with AI researchers, ML engineers, platform engineers, and infrastructure teams to improve model performance and production reliability.
- Build automated benchmarking frameworks, performance dashboards, monitoring solutions, and regression testing pipelines.
- Evaluate emerging AI hardware, GPU architectures, inference frameworks, and optimization technologies to improve enterprise AI capabilities.
- Drive AI infrastructure cost optimization through efficient resource utilization, cloud optimization, and FinOps best practices.
- Mentor engineering teams and provide technical leadership on AI systems architecture, GPU optimization, and performance engineering.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, or a related technical discipline.
- 10+ years of professional experience in performance engineering, AI infrastructure, machine learning systems, High-Performance Computing (HPC), or distributed computing.
- Expert-level programming skills in Python and C++.
- Extensive experience optimizing GPU-accelerated AI workloads using CUDA, distributed training frameworks, and modern deep learning libraries.
- Strong knowledge of Large Language Models (LLMs), deep learning frameworks, model serving, and production AI inference.
- Hands-on experience with profiling tools such as NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, TensorBoard, or similar performance analysis tools.
- Experience deploying and optimizing AI workloads on AWS, Microsoft Azure, or Google Cloud Platform (GCP).
- Strong understanding of distributed systems, networking, storage optimization, and AI infrastructure architecture.
- Excellent analytical, troubleshooting, communication, and technical leadership skills.
Preferred Qualifications
- Experience optimizing production-scale LLM inference and serving large foundation models.
- Hands-on experience with vLLM, TensorRT-LLM, DeepSpeed, Triton Inference Server, CUTLASS, FasterTransformer, or similar AI optimization frameworks.
- Knowledge of model compression, KV cache optimization, speculative decoding, and advanced inference optimization techniques.
- Experience implementing FinOps strategies for AI infrastructure cost optimization and resource management.
- Contributions to AI systems research, open-source AI infrastructure projects, patents, or technical publications.
- Familiarity with emerging AI accelerator technologies, including AMD ROCm, Intel oneAPI, or custom AI hardware.
Interested in this opportunity? Apply today for immediate consideration!
Email your updated resume:
Call or Text: (908) 505-3545
Learn more:
Bright Vision Technologies is an Equal Opportunity Employer.
Equal Employment Opportunity (EEO) Statement
Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.
BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
Originally posted on Himalayas