远程工作雷达

AI解决方案架构师

AI Solution Architect

AI开发工程限定地区(需当地身份)
公司uvation
薪资未公开
工作地点India
地域资格限定地区(需当地身份)
时区要求日间重叠约 6 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 India 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

职位概述

我们正在寻找一位经验丰富的AI解决方案架构师,负责设计和领导涵盖计算、高性能网络、存储、Kubernetes、云和AI/ML平台的端到端企业AI工厂和GPU基础设施解决方案。该职位需要在NVIDIA GPU技术、AI工作负载、可扩展基础设施架构、安全、可观测性、性能工程和容量规划方面具备深厚的专业知识。

主要职责

  • 从需求到生产就绪,负责AI工厂和企业AI解决方案的端到端架构。
  • 评估AI/ML工作负载需求,包括训练、微调、推理、批量处理和高性能计算。
  • 设计GPU计算架构,包括NVIDIA HGX/DGX/OEM平台、多GPU系统、NVLink/NVSwitch以及GPU资源分配。
  • 使用100/200/400/800G以太网、EVPN/VXLAN和叶脊架构设计高性能AI网络。
  • 使用对象存储、并行文件系统如Ceph、WEKA等设计AI存储和数据架构。
  • 在Kubernetes、HPC、容器运行时、模型服务框架和企业AI框架上定义AI平台架构。
  • 建立安全、身份、租户隔离、数据保护、可观测性、灾难恢复和运营弹性的架构标准。
  • 开发参考架构、高层/低层设计、容量模型、物料清单、技术评估和实施路线图。
  • 领导技术评估、概念验证、供应商评估和架构评审委员会。
  • 与基础设施、网络、安全、存储、云、数据、应用和运维团队协作。
  • 定义性能、可用性、可扩展性、安全性和成本目标,并根据可衡量的验收标准验证架构。
  • 在部署、迁移、集成、故障排除和生产过渡期间提供技术领导。

• 所需技术技能

AI / ML架构

• NVIDIA AI Enterprise、NGC、CUDA、NCCL、DCGM、GPU Operator和AI平台生态系统。

• PyTorch、TensorFlow、JAX以及对训练和推理工作负载的操作理解。

• GPU调度、多租户、MIG/vGPU、GPU利用率和工作负载放置。

• 大型语言模型、生成式AI、RAG、微调、模型服务和推理架构。

GPU与AI工厂基础设施

• NVIDIA A100/H100/H200/B200或同等GPU

查看英文原文

Job Overview

We are seeking an experienced AI Solution Architectto design and lead end-to-end enterprise AI Factory and GPU infrastructure solutions spanning compute, high-performance networking, storage, Kubernetes, cloud, and AI/ML platforms. The role requires strong expertise in NVIDIA GPU technologies, AI workloads, scalable infrastructure architecture, security, observability, performance engineering, and capacity planning.

Key Responsibilities

  • Own end-to-end architecture for AI Factory and enterprise AI solutions from requirements through production readiness.
  • Assess AI/ML workload requirements for training, fine-tuning, inference, batch processing, and high-performance computing.
  • Design GPU compute architectures including NVIDIA HGX/DGX/OEM platforms, multi-GPU systems, NVLink/NVSwitch, and GPU resource allocation.
  • Design high-performance AI networking using 100/200/400/800G Ethernet, EVPN/VXLAN, and leaf-spine architectures.
  • Design AI storage and data architectures using object storage, parallel file systems like Ceph, WEKA, , or equivalent platforms.
  • Define AI platform architecture across Kubernetes, HPC, container runtimes, model-serving platforms, and enterprise AI frameworks.
  • Establish architecture standards for security, identity, tenant isolation, data protection, observability, disaster recovery, and operational resilience.
  • Develop reference architectures, high-level/low-level designs, capacity models, bills of materials, technology evaluations, and implementation roadmaps.
  • Lead technical evaluations, proof-of-concepts, vendor assessments, and architecture review boards.
  • Collaborate with infrastructure, network, security, storage, cloud, data, application, and operations teams.
  • Define performance, availability, scalability, security, and cost objectives and validate architecture against measurable acceptance criteria.
  • Provide technical leadership during deployment, migration, integration, troubleshooting, and production transition.

• Required Technical Skills

AI / ML Architecture

• NVIDIA AI Enterprise, NGC, CUDA, NCCL, DCGM, GPU Operator and AI platform ecosystem.

• PyTorch, TensorFlow, JAX and operational understanding of training and inference workloads.

• GPU scheduling, multi-tenancy, MIG/vGPU, GPU utilization and workload placement.

• LLM, generative AI, RAG, fine-tuning, model serving and inference architecture.

GPU & AI Factory Infrastructure

• NVIDIA A100/H100/H200/B200 or equivalent GPU platforms; familiarity with next-generation systems.

• NVLink, NVSwitch, PCIe topology and multi-GPU performance architecture.

• DGX/HGX/OEM GPU server architecture and lifecycle management.

• AI Factory capacity planning, rack density, power, cooling, commissioning and lifecycle strategy.

High-Performance Networking

• 100/200/400/800G Ethernet, InfiniBand, RoCEv2 and RDMA, Netris

• NVIDIA ConnectX/SuperNIC, Spectrum/Spectrum-X, Quantum and BlueField DPU technologies.

• BGP, EVPN/VXLAN, VRF, ECMP, VLAN, MTU, PFC, ECN, QoS and congestion management.

• GPU east-west traffic, GPUDirect RDMA and network performance troubleshooting.

AI Storage & Data Architecture

• Parallel file systems, object storage, NFS, NVMe/NVMe-oF and high-throughput data pipelines.

• Ceph, WEKA, VAST, Dell PowerScale, Pure FlashBlade, NetApp or equivalent technologies.

• Data lake/lakehouse concepts, metadata, lineage, data movement and data lifecycle.

• GPUDirect Storage and storage/network performance optimization.

AI Platform & Orchestration

• Kubernetes, GPU Operator, container runtimes and Kubernetes GPU scheduling.

• HPC or other equivalent workload schedulers.

• Model serving/inference platforms and MLOps platform architecture.

• API gateways, service discovery, secrets management and platform integration.

Cloud & Hybrid Architecture

• AWS and/or Azure AI infrastructure and security services.

• Hybrid cloud connectivity, IAM, private networking, cloud storage and workload placement.

• Cloud cost optimization, capacity planning and FinOps considerations for GPU workloads.

Security & Governance

• Zero Trust, network segmentation, IAM/RBAC, PAM and workload identity.

• GPU, DPU, container, Kubernetes, firmware and supply-chain security.

• Encryption at rest/in transit, secrets management, audit logging and compliance controls.

• AI-specific risks including data/model protection, tenant isolation and secure model access.

Observability & Reliability

• Prometheus, Grafana, OpenTelemetry, NVIDIA DCGM and infrastructure telemetry.

• Monitoring across GPU, CPU, memory, network, storage, power and thermal domains.

• High availability, backup/restore, disaster recovery, business continuity and failure-domain design.

• Performance engineering, bottleneck analysis, SLO/SLA design and capacity forecasting.

Architecture Deliverables

• AI Factory reference architecture and solution blueprints

• High-Level Design (HLD) and Low-Level Design (LLD)

• Network, compute, GPU and storage architecture diagrams

• Capacity, performance and scalability models

• Technology evaluation and vendor comparison documents

• Security architecture and threat-model inputs

• Bill of Materials (BOM) and infrastructure sizing

• Migration/deployment strategy and implementation roadmap

• Operational readiness checklist, runbooks and acceptance criteria

Experience & Qualifications

• 10+ years of infrastructure, cloud, enterprise architecture or solution architecture experience, with significant AI/GPU infrastructure exposure.

• Proven experience designing large-scale enterprise platforms and translating business requirements into technical architectures.

• Hands-on understanding of physical infrastructure, GPU systems, networking, storage and Linux platforms.

• Bachelor's degree in Computer Science, Engineering, Information Technology or related field preferred.

Preferred Certifications

• NVIDIA certifications or equivalent GPU/AI infrastructure credentials

• AWS Solutions Architect / Azure Solutions Architect

• TOGAF or equivalent enterprise architecture certification

• CCNP/CCIE or equivalent networking certification

• CISSP or equivalent security certification

• Kubernetes certifications such as CKA/CKAD

• Red Hat / Linux certifications

Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

← 返回全部职位