GPU 编程专家 - 全程远程 | 最高 120 美元/小时
GPU Programming Expert - Fully Remote | Upto $120/hr
关于职位
Mercor 将顶尖的创意和技术人才与领先的 AI 研究实验室联系起来。公司总部位于旧金山,我们的投资人包括 Benchmark、General Catalyst、Peter Thiel、Adam D'Angelo、Larry Summers 和 Jack Dorsey。
职位:CUDA 工程专家
类型:合同工
薪酬:80–120 美元/小时
地点:远程
岗位职责
- 分析和优化 GPU 内核以提高性能、效率和硬件利用率。
- 使用 L2 缓存命中率、L2 吞吐量和占用率等分析器指标来指导内核优化。
- 审查 GPU 内核实现,识别瓶颈,而无需深入的算法背景。
- 编写、修改和理解 C++17、Python 和 GPU 编程代码。
- 应用 CUDA、HIP 和着色器编程专业知识以提升性能结果。
- 清晰地记录优化决策,并注明特定分析器指标何时有用。
资格要求
必须具备
- 每周至少可工作 20 小时。
- 精通 C++17 的核心特性。
- 熟悉 Python 和 Git。
- 精通至少一种 GPU 编程模型,如 CUDA、HIP、Slang、HLSL 或 GLSL。
- 至少有一年使用 GPU 的专业或研究生级别研究经验。
- 对 GPU 分析器性能指标有深入了解,用于内核优化。
- 能够在不深入了解每个算法的前提下优化 GPU 内核。
优先考虑
- 具有 CUDA、HIP、CUDA C++ 核心库、内联 PTX 汇编或张量核心级优化的经验。
- 具有针对 NVIDIA Blackwell 硬件优化内核的经验。
- 熟悉 NSight Compute。
- 具有 NVIDIA、AMD 或高通等 GPU 硬件架构的相关经验。
- 与 GPU 内核优化相关的开源贡献。
申请流程(需要 20–30 分钟完成)
- 提交简历或相关技术背景以开始申请。
- 符合条件的申请人可能会被要求完成简短的技术评估或提交更多信息。
资源与支持
- 有关面试流程和平台信息的详细信息,请查看:
- 如有任何帮助或支持需求,请联系:
备注:我们的团队每天都会审核申请。请完成您的 AI 面试和申请步骤,以考虑此机会。
最初发布于 Himalayas
查看英文原文
About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.
Position: CUDA Engineering Expert
Type:Contract
Compensation:$80–$120/hour
Location:Remote
Role Responsibilities
- Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization.
- Use profiler metrics like L2 cache hit rate, L2 throughput, and occupancy to guide kernel improvements.
- Review GPU kernel implementations to identify bottlenecks without needing extensive algorithmic background.
- Write, modify, and reason about C++17, Python, and GPU programming code.
- Apply CUDA, HIP, and shader programming expertise to improve performance outcomes.
- Document optimization decisions clearly, noting when specific profiler metrics are useful.
Qualifications
Must-Have
- Available to work at least 20 hrs/wk.
- Fluent in core C++ features through C++17.
- Working knowledge of Python and Git.
- Fluent in at least one GPU programming model like CUDA, HIP, Slang, HLSL, or GLSL.
- At least 1 year of professional or graduate-level research experience with GPUs.
- Strong understanding of GPU profiler performance metrics for kernel optimization.
- Ability to optimize GPU kernels without deep prior context on every algorithm.
Preferred
- Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization.
- Experience optimizing kernels for NVIDIA Blackwell hardware.
- Familiarity with NSight Compute.
- Prior experience with GPU hardware organizations like NVIDIA, AMD, or Qualcomm.
- Open-source contributions related to GPU kernel optimization.
Application Process (Takes 20–30 mins to complete)
- Submit your resume or relevant technical background to get started.
- Qualified applicants may be asked to complete a brief technical assessment or submit additional information.
Resources & Support
- For details about the interview process and platform information, please check:
- For any help or support, reach out to:
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
Originally posted on Himalayas