高级软件工程师,平台基础设施
Senior Software Engineer, Platform Infrastructure
Moonlite 为运行密集型计算研究、大规模模型训练和高要求数据处理工作负载的组织提供高性能 AI 基础设施。我们提供部署在我们设施中的基础设施,或与您本地共置,提供灵活的按需或预留计算,感觉就像您现有数据中心的延伸。我们的 AI 基础设施专家团队将裸金属性能与云原生操作简便性相结合,使研究团队和企业能够以企业级可靠性与合规性部署高要求的 AI 工作负载。
你的职责:
你将是我们构建综合性基础设施平台的核心成员,该平台将我们的物理基础设施——裸金属服务器、GPU 集群、高性能存储和网络结构——与客户依赖的用于大规模计算、推理、模拟和训练的系统连接起来。你将与产品团队、平台团队成员和基础设施专家紧密合作,设计并实现编排层、API 和自动化框架,使数千台服务器、拍字节级存储和高速网络感觉像一个统一、可编程的平台。
职位职责
- 基础设施抽象层:设计并构建系统,将物理基础设施(裸金属服务器、存储集群、网络结构)与面向客户的业务连接起来,实现对计算、网络和存储的大规模程序化管理。
- 研究集群配置:设计并实现用于配置和管理研究计算环境的系统,包括 Kubernetes 和 SLURM 集群,支持分布式 AI 训练和 HPC 工作负载的自动化部署、资源调度和工作流编排。
- 平台编排:实现全面的编排系统,协调计算、存储和网络,为复杂的研究工作负载提供统一的体验。
- 网络自动化与部署:设计并构建网络配置自动化,包括智能虚拟机部署决策以优化网络拓扑、自动 VLAN 和子网配置,以及用于高性能互连的软件定义网络编排。
- 企业 API 与 SDK:开发强大的 API 和 SDK,使研究人员和工程团队能够跨所有平台领域程序化配置和管理基础设施资源。
- 可观测性与监控:设计并实现系统以监控和分析平台性能、资源使用情况和故障检测,确保稳定性和高效性。
查看英文原文
Moonlite delivers high-performance AI infrastructure for organizations running intensive computational research, large-scale model training, and demanding data processing workloads.We provide infrastructure deployed in our facilities or co-located in yours, delivering flexible on-demand or reserved compute that feels like an extension of your existing data center. Our team of AI infrastructure specialists combines bare-metal performance with cloud-native operational simplicity, enabling research teams and enterprises to deploy demanding AI workloads with enterprise-grade reliability and compliance.
Your Role:
You will be foundational to building the comprehensive infrastructure platform that bridges our physical infrastructure – bare-metal servers, GPU clusters, high-performance storage, and networking fabric – with the systems our customers depend on for large-scale computation, inference, simulations, and training. Working closely with product, your platform team members, and infrastructure specialists, you’ll design and implement the orchestration layer, APIs, and automation framework that make thousands of servers, petabytes of storage, and high-speed networks feel like a unified, programmable platform.
Job Responsibilities
- Infrastructure Abstraction Layer: Design and build systems that bridge physical infrastructure (bare-metal servers, storage clusters, network fabric) with customer-facing services, enabling programmatic management of compute, networking, and storage at scale.
- Research Cluster Provisioning: Design and implement systems for provisioning and managing research computing environments including Kubernetes and SLURM clusters, enabling automated deployment, resource scheduling, and workload orchestration for distributed AI training and HPC workloads.
- Platform Orchestration: Implement comprehensive orchestration systems that coordinate across compute, storage and networking to deliver unified experience for complex research workloads.
- Network Automation & Placement: Design and build network provisioning automation including intelligent VM placement decisions for optimal network topology, automated VLAN and subnet configuration, and software-designed networking orchestration for high-performance interconnects.
- Enterprise APIs & SDKs: Develop robust APIs and SDKs that enable researchers and engineering teams to programmatically provision and manage infrastructure resources across all platform domains.
- Observability & Telemetry: Implement comprehensive observability, telemetry, and logging systems that provide visibility into infrastructure health, performance, and utilization across the infrastructure footprint.
- Performance Engineering: Build and optimize platform services that deliver consistent high-throughput low-latency networking for demand research applications and data-intensive workloads.
- Cross-Team Collaboration: Work closely with engineering, infrastructure, and product to define requirements, drive infrastructure-product-rollouts, and improve resource lifecycle management.
- Compliance & Security: Implement platform-wide compliance and security features supporting SOC 2, ISO 27001, and enterprise regulatory requirements including comprehensive audit logging, access controls, and data residency management.
Requirements
- Experience: 5+ years in software engineering with a proven track record of infrastructure platforms, distributed systems, or cloud platforms for production environments.
- Kubernetes & Container Orchestration: Strong familiarity with Kubernetes architecture, container orchestration concepts, and experience deploying workloads in Kubernetes environments. Understanding of pods, deployments, services, and basic Kubernetes operations.
- Infrastructure Systems: Strong understanding of infrastructure fundamentals including compute orchestration, storage systems, networking technologies, and how they integrate together to deliver complete platform experiences.
- Programming Skills: Experience with systems programming languages (Go, C/C++, Rust, Python) for performance-critical components is a strong plus.
- Linux Production Experience: Strong experience with linux in production environments, including systems administration, performance tuning, and troubleshooting.
- Bare-Metal & Virtualization: Deep knowledge of bare-metal infrastructure, provisioning systems, out-of-band management, and virtualization technologies (KVM, Kubernetes, etc).
- API & Platform Design: Proven experience designing and building APIs, SDKs, and automation frameworks that enable programmatic infrastructure management.
- Cloud Platform Knowledge: Strong familiarity with cloud environments (AWS, GCP, Azure) and understanding of how to translate cloud-native patterns to bare-metal infrastructure.
- Infrastructure Automation: Experience with Infrastructure-as-code tools (Terraform, Ansible) and building automated deployment pipelines.
- Problem Solving & Autonomy: Self-starter who can navigate ambiguity, balance pragmatic shipping with good long-term architecture, and independently drive complex technical initiatives.
- Communication Skills: Strong written and verbal communication skills, including ability to write clear technical communication and collaborate across teams.
- Commitment to Growth: Growth mindset with continuous focus on learning and professional development.
Preferred Qualifications
- Background provisioning or managing research computing environments (Kubernetes, SLURM, or HPC clusters)
- Experience building internal platforms, infrastructure-as-a-service, or developer tooling
- Background with GPU computing platforms and AI/ML infrastructure requirements
- Knowledge of high-performance networking technologies (InfiniBand, RDMA, SR-IOV)
- Experience with observability and monitoring platforms (Prometheus, Grafana, ELK stack)
- Familiarity with both cloud-native and bare-metal infrastructure deployment models
- Understanding of enterprise compliance requirements and security best practices
- Extra points for experience with financial services technology infrastructure and understanding of trading system requirements
Key Technologies
· Go, Python, Kubernetes, Docker, Terraform, Ansible, Linux, Networking (BGP, VXLAN), Storage Systems, FastAPI, PostgreSQL, Redis, NVIDIA GPU Technologies, InfiniBand
Why Moonlite
- Build Next-Generation Infrastructure: Your work will create the platform foundation that enables financial institutions to harness AI capabilities previously impossible with traditional infrastructure.
- Hands-On Ownership: As an early engineer, you’ll have end-to-end ownership of projects and the autonomy to influence our product and technology direction.
- Shape Industry Standards: Contribute to defining how enterprise AI infrastructure should work for the most demanding regulated environments.
- Collaborate with Experts: Work alongside seasoned engineers and industry professionals passionate about high-performance computing, innovation, and problem-solving.
- Start-Up Agility with Industry Impact: Enjoy the dynamic, fast-paced environment of a startup while making an immediate impact in an evolving and critical technology space.
We offer a competitive total compensation package combining a competitive base salary, startup equity, and industry-leading benefits. The total compensation range for this role is $165,000 – $225,000, which includes both base salary and equity. Actual compensation will be determined based on experience, skills, and market alignment. We provide generous benefits, including a 6% 401(k) match, fully covered health insurance premiums, and other comprehensive offerings to support your well-being and success as we grow together.
#li-remote
Originally posted on Himalayas