AI解决方案架构师
AI Solution Architect
职位概述
我们正在寻找一位经验丰富的AI解决方案架构师,负责设计和领导涵盖计算、高性能网络、存储、Kubernetes、云和AI/ML平台的端到端企业AI工厂和GPU基础设施解决方案。该职位需要在NVIDIA GPU技术、AI工作负载、可扩展基础设施架构、安全、可观测性、性能工程和容量规划方面具备深厚的专业知识。
主要职责
- 从需求到生产就绪,负责AI工厂和企业AI解决方案的端到端架构。
- 评估AI/ML工作负载需求,包括训练、微调、推理、批量处理和高性能计算。
- 设计GPU计算架构,包括NVIDIA HGX/DGX/OEM平台、多GPU系统、NVLink/NVSwitch以及GPU资源分配。
- 使用100/200/400/800G以太网、EVPN/VXLAN和叶脊架构设计高性能AI网络。
- 使用对象存储、并行文件系统如Ceph、WEKA等设计AI存储和数据架构。
- 在Kubernetes、HPC、容器运行时、模型服务框架和企业AI框架上定义AI平台架构。
- 建立安全、身份、租户隔离、数据保护、可观测性、灾难恢复和运营弹性的架构标准。
- 开发参考架构、高层/低层设计、容量模型、物料清单、技术评估和实施路线图。
- 领导技术评估、概念验证、供应商评估和架构评审委员会。
- 与基础设施、网络、安全、存储、云、数据、应用和运维团队协作。
- 定义性能、可用性、可扩展性、安全性和成本目标,并根据可衡量的验收标准验证架构。
- 在部署、迁移、集成、故障排除和生产过渡期间提供技术领导。
• 所需技术技能
AI / ML架构
• NVIDIA AI Enterprise、NGC、CUDA、NCCL、DCGM、GPU Operator和AI平台生态系统。
• PyTorch、TensorFlow、JAX以及对训练和推理工作负载的操作理解。
• GPU调度、多租户、MIG/vGPU、GPU利用率和工作负载放置。
• 大型语言模型、生成式AI、RAG、微调、模型服务和推理架构。
GPU与AI工厂基础设施
• NVIDIA A100/H100/H200/B200或同等GPU
查看英文原文
Job Overview
We are seeking an experienced AI Solution Architectto design and lead end-to-end enterprise AI Factory and GPU infrastructure solutions spanning compute, high-performance networking, storage, Kubernetes, cloud, and AI/ML platforms. The role requires strong expertise in NVIDIA GPU technologies, AI workloads, scalable infrastructure architecture, security, observability, performance engineering, and capacity planning.
Key Responsibilities
- Own end-to-end architecture for AI Factory and enterprise AI solutions from requirements through production readiness.
- Assess AI/ML workload requirements for training, fine-tuning, inference, batch processing, and high-performance computing.
- Design GPU compute architectures including NVIDIA HGX/DGX/OEM platforms, multi-GPU systems, NVLink/NVSwitch, and GPU resource allocation.
- Design high-performance AI networking using 100/200/400/800G Ethernet, EVPN/VXLAN, and leaf-spine architectures.
- Design AI storage and data architectures using object storage, parallel file systems like Ceph, WEKA, , or equivalent platforms.
- Define AI platform architecture across Kubernetes, HPC, container runtimes, model-serving platforms, and enterprise AI frameworks.
- Establish architecture standards for security, identity, tenant isolation, data protection, observability, disaster recovery, and operational resilience.
- Develop reference architectures, high-level/low-level designs, capacity models, bills of materials, technology evaluations, and implementation roadmaps.
- Lead technical evaluations, proof-of-concepts, vendor assessments, and architecture review boards.
- Collaborate with infrastructure, network, security, storage, cloud, data, application, and operations teams.
- Define performance, availability, scalability, security, and cost objectives and validate architecture against measurable acceptance criteria.
- Provide technical leadership during deployment, migration, integration, troubleshooting, and production transition.
• Required Technical Skills
AI / ML Architecture
• NVIDIA AI Enterprise, NGC, CUDA, NCCL, DCGM, GPU Operator and AI platform ecosystem.
• PyTorch, TensorFlow, JAX and operational understanding of training and inference workloads.
• GPU scheduling, multi-tenancy, MIG/vGPU, GPU utilization and workload placement.
• LLM, generative AI, RAG, fine-tuning, model serving and inference architecture.
GPU & AI Factory Infrastructure
• NVIDIA A100/H100/H200/B200 or equivalent GPU platforms; familiarity with next-generation systems.
• NVLink, NVSwitch, PCIe topology and multi-GPU performance architecture.
• DGX/HGX/OEM GPU server architecture and lifecycle management.
• AI Factory capacity planning, rack density, power, cooling, commissioning and lifecycle strategy.
High-Performance Networking
• 100/200/400/800G Ethernet, InfiniBand, RoCEv2 and RDMA, Netris
• NVIDIA ConnectX/SuperNIC, Spectrum/Spectrum-X, Quantum and BlueField DPU technologies.
• BGP, EVPN/VXLAN, VRF, ECMP, VLAN, MTU, PFC, ECN, QoS and congestion management.
• GPU east-west traffic, GPUDirect RDMA and network performance troubleshooting.
AI Storage & Data Architecture
• Parallel file systems, object storage, NFS, NVMe/NVMe-oF and high-throughput data pipelines.
• Ceph, WEKA, VAST, Dell PowerScale, Pure FlashBlade, NetApp or equivalent technologies.
• Data lake/lakehouse concepts, metadata, lineage, data movement and data lifecycle.
• GPUDirect Storage and storage/network performance optimization.
AI Platform & Orchestration
• Kubernetes, GPU Operator, container runtimes and Kubernetes GPU scheduling.
• HPC or other equivalent workload schedulers.
• Model serving/inference platforms and MLOps platform architecture.
• API gateways, service discovery, secrets management and platform integration.
Cloud & Hybrid Architecture
• AWS and/or Azure AI infrastructure and security services.
• Hybrid cloud connectivity, IAM, private networking, cloud storage and workload placement.
• Cloud cost optimization, capacity planning and FinOps considerations for GPU workloads.
Security & Governance
• Zero Trust, network segmentation, IAM/RBAC, PAM and workload identity.
• GPU, DPU, container, Kubernetes, firmware and supply-chain security.
• Encryption at rest/in transit, secrets management, audit logging and compliance controls.
• AI-specific risks including data/model protection, tenant isolation and secure model access.
Observability & Reliability
• Prometheus, Grafana, OpenTelemetry, NVIDIA DCGM and infrastructure telemetry.
• Monitoring across GPU, CPU, memory, network, storage, power and thermal domains.
• High availability, backup/restore, disaster recovery, business continuity and failure-domain design.
• Performance engineering, bottleneck analysis, SLO/SLA design and capacity forecasting.
Architecture Deliverables
• AI Factory reference architecture and solution blueprints
• High-Level Design (HLD) and Low-Level Design (LLD)
• Network, compute, GPU and storage architecture diagrams
• Capacity, performance and scalability models
• Technology evaluation and vendor comparison documents
• Security architecture and threat-model inputs
• Bill of Materials (BOM) and infrastructure sizing
• Migration/deployment strategy and implementation roadmap
• Operational readiness checklist, runbooks and acceptance criteria
Experience & Qualifications
• 10+ years of infrastructure, cloud, enterprise architecture or solution architecture experience, with significant AI/GPU infrastructure exposure.
• Proven experience designing large-scale enterprise platforms and translating business requirements into technical architectures.
• Hands-on understanding of physical infrastructure, GPU systems, networking, storage and Linux platforms.
• Bachelor's degree in Computer Science, Engineering, Information Technology or related field preferred.
Preferred Certifications
• NVIDIA certifications or equivalent GPU/AI infrastructure credentials
• AWS Solutions Architect / Azure Solutions Architect
• TOGAF or equivalent enterprise architecture certification
• CCNP/CCIE or equivalent networking certification
• CISSP or equivalent security certification
• Kubernetes certifications such as CKA/CKAD
• Red Hat / Linux certifications
Originally posted on Himalayas