软件工程师
Software Engineer
**关于 Positron AI**
Positron AI 专注于开发定制硬件系统以加速 AI 推理。这些推理系统相比传统的 GPU 系统在性能和效率上都有显著提升,带来了每美元和每瓦特性能的优势。Positron 的目标是打造世界上最好的 AI 推理系统。
**职位概述**
**高级软件工程师 – 机器学习系统与高性能 LLM 推理**
我们正在寻找一位高级软件工程师,参与开发用于在我们的定制设备上执行开源大语言模型(LLM)的高性能软件。该设备结合了 FPGA 和 x86 CPU 来加速基于 transformer 的模型。软件栈主要使用现代 C++(C++17/20),并大量依赖模板、SIMD 优化和高效的并行计算技术。
**主要职责**
- 设计和实现在定制硬件上运行的高性能 LLM 推理软件。
- 开发和优化基于 C++ 的库,高效利用 SIMD 指令、线程和内存层次结构。
- 与 FPGA 和系统工程师紧密合作,确保 x86 CPU 和 FPGA 之间的数据传输和计算卸载高效进行。
- 通过低级优化(包括向量化、缓存效率和硬件感知调度)优化模型执行。
- 贡献性能分析工具和方法论,以在指令和数据流层面分析执行瓶颈。
- 应用 NUMA 感知的内存管理技术,优化大规模推理工作负载的内存访问模式。
- 实现 ML 系统级别的优化,如 token 流式传输、KV 缓存优化和 transformer 执行的高效批处理。
- 与 ML 研究人员和软件工程师合作,集成模型量化技术、稀疏性优化和混合精度执行。
- 确保所有代码贡献都包含单元测试、性能测试、验收测试和回归测试,作为基于持续集成的开发流程的一部分。
**所需资格**
- 7 年以上 C++ 软件开发经验,重点在性能关键的应用。
- 对 C++ 模板和现代内存管理有深刻理解。
- 具备 SIMD 编程(AVX-512、SSE 或等效)和基于 intrinsics 的向量化经验。
- 具有高性能计算或相关领域的经验。
查看英文原文
**About Positron AI**
Positron AI specializes in developing custom hardware systems to accelerate AI inference. These inference systems offer significant performance and efficiency gains over traditional GPU-based systems, delivering advantages in both performance per dollar and performance per watt. Positron exists to create the world's best AI inference systems.
**Role Overview**
**Senior Software Engineer – Machine Learning Systems & High-Performance LLM Inference**
We are seeking a Senior Software Engineer to contribute to the development of high-performance software that powers execution of open-source large language models (LLMs) on our custom appliance. This appliance leverages a combination of FPGAs and x86 CPUs to accelerate transformer-based models. The software stack is written primarily in modern C++ (C++17/20) and heavily relies on templates, SIMD optimizations, and efficient parallel computing techniques.
**Key Responsibilities**
- Design and implement high-performance inference software for LLMs on custom hardware.
- Develop and optimize C++-based libraries that efficiently utilize SIMD instructions, threading, and memory hierarchy.
- Work closely with FPGA and systems engineers to ensure efficient data movement and computational offloading between x86 CPUs and FPGAs.
- Optimize model execution via low-level optimizations, including vectorization, cache efficiency, and hardware-aware scheduling.
- Contribute to performance profiling tools and methodologies to analyze execution bottlenecks at the instruction and data flow levels.
- Apply NUMA-aware memory management techniques to optimize memory access patterns for large-scale inference workloads.
- Implement ML system-level optimizations such as token streaming, KV cache optimizations, and efficient batching for transformer execution.
- Collaborate with ML researchers and software engineers to integrate model quantization techniques, sparsity optimizations, and mixed-precision execution.
- Ensure all code contributions include unit, performance, acceptance, and regression tests as part of a continuous integration-based development process.
**Required Qualifications**
- 7+ years of professional experience in C++ software development, with a focus on performance-critical applications.
- Strong understanding of C++ templates and modern memory management.
- Hands-on experience with SIMD programming (AVX-512, SSE, or equivalent) and intrinsics-based vectorization.
- Experience in high-performance computing (HPC), numerical computing, or ML inference optimization.
- Experience with ML model execution optimizations, including efficient tensor computations and memory access patterns.
- Knowledge of multi-threading, NUMA architectures, and low-level CPU optimization.
- Proficiency with systems-level software development, profiling tools (perfetto, VTune, Valgrind), and benchmarking.
- Experience working with hardware accelerators (FPGAs, GPUs, or custom ASICs) and designing efficient software-hardware interfaces.
**Preferred Qualifications**
- Familiarity with LLVM/Clang or GCC compiler optimizations.
- Experience in LLM quantization, sparsity optimizations, and mixed-precision computation.
- Knowledge of distributed inference techniques and networking optimizations.
- Understanding of graph partitioning and execution scheduling for large-scale ML models.
**Leveling & Scope**
While this role is currently posted at a specific level, we are a growth-oriented organization and are open to hiring at a more senior level for the right candidate. Please note that this job description serves as a focused but generalized overview of the role; specific responsibilities and impact expectations will be tailored to the experience and seniority of the final hire.
**Why Join Us?**
- Work on a cutting-edge ML inference platform that redefines performance and efficiency for LLMs.
- Tackle challenging low-level performance engineering problems in AI and HPC.
- Collaborate with a team of hardware, software, and ML experts building an industry-first product.
- Opportunity to contribute to and shape the future of open-source AI inference software.
**Compensation and Benefits**
The base salary range for this role is $150,000 – $250,000.
Please note that the figures provided represent the **base salary range only** and do not include other elements of our total compensation package, equity, or comprehensive benefits.
At Positron AI, we value the unique expertise each candidate brings. While the range above reflects our typical expectation for the position, **we reserve the flexibility to exceed this range** for candidates whose specialized skills, significant experience, or unique qualifications fall outside the standard scope of the role. Final offers are determined based on a variety of factors, including internal equity, and individual impact.
### **Benefits & Perks**
We want you to do your best work and feel confident that you and your family are taken care of. That means comprehensive coverage, real time to rest, and support for your future.
**Health and wellness**
- Fully company-paid medical, dental, and vision insurance for you and your dependents
- Company-paid life and disability coverage, with voluntary options to add more
- Supplemental hospital, critical illness, and accident coverage available
**Time off and flexibility**
- Unlimited paid time off — we encourage everyone to truly unplug and recharge
- 13 paid company holidays
- Remote-first culture with a company-provided computer and home office setup
**Compensation and future**
- Competitive salary and equity
- 401(k) with company matching, eligible from day one
### **Visa Support**
This position is open to candidates currently authorized to work in the U.S. We cannot provide new visa sponsorship for this role but are open to facilitating H-1B visa transfers for eligible candidates.
**Equal Opportunity Employer.** If you’re excited about the role but don’t meet every bullet, we’d still love to hear from you.