技术架构师 - 机器学习
Technical Architect - ML
虽然技术是我们的业务核心,但全球多元的文化是成功的核心。我们热爱我们的员工,并自豪地为他们提供一个以透明、多样、诚信、学习和成长为基础的文化。
如果在一个鼓励你不仅在职业上而且在个人生活中创新和卓越的环境中工作吸引你,那么你将在Quantiphi拥有愉快的职业生涯!
必备技能与资格:
- 8年以上在机器学习/人工智能工程或MLOps角色的工作经验,具备扎实的架构经验。
- 精通AWS云原生机器学习堆栈,包括:SageMaker(主要)、EKS、Lambda、API Gateway、CI/CD(CodeBuild/CodePipeline或类似工具)
- 至少掌握一个主流的MLOps工具集并了解其他替代方案:MLflow、Kubeflow、SageMaker Pipelines、Airflow、BentoML、KServe、Seldon。
- 深入理解模型生命周期管理(特征工程→训练→注册→部署→监控)。
- 有实现或支持LLMOps流水线的经验,包括:提示版本控制、评估指标、自动化框架。
- 深入理解机器学习生命周期:数据摄入、特征工程、训练、评估、模型打包、CI/CD、漂移检测、监控和治理。
- 熟练使用AWS SageMaker(Pipelines、Feature Store、Model Registry、Model Monitor)。
- 有实施机器学习CI/CD流水线的经验,包括自动化训练、测试、验证、模型晋升和端点部署。
- 有使用基础设施即代码(IaC)工具和CI/CD流水线的经验。
- 有基于Kubernetes的开发经验。
- 有特征工程流水线和Feature Store管理经验。
- 了解血缘追踪:训练数据快照、特征版本、代码版本控制、元数据追踪、可复现性。
- 有使用AWS Bedrock和Agentcore服务的实际经验。
- 有CloudWatch、SageMaker Model Monitor、Prometheus/Grafana的经验。
- 扎实的Python基础和云原生开发模式理解。
- 对安全最佳实践、IAM、密钥管理和制品治理有深入理解。
加分技能:
- 有向量数据库、RAG流水线或多智能体AI系统的经验。
- 有DevOps和基础设施即代码(Terraform、Helm、CDK)的接触经验。
- 对模型漂移检测、A/B测试、金丝雀发布和蓝绿部署有实际理解。
- 熟悉可观测性堆栈(Prometheus
查看英文原文
While technology is the heart of our business, a global and diverse culture is the heart of our success. We love our people and we take pride in catering them to a culture built on transparency, diversity, integrity, learning and growth.
If working in an environment that encourages you to innovate and excel, not just in professional but personal life, interests you- you would enjoy your career with Quantiphi!
Must have skills & Qualifications:
- 8+ years working in ML/AI engineering or MLOps roles with strong architecture exposure.
- Strong expertise in AWS cloud-native ML stack, including: SageMaker(primary), EKS, Lambda, API Gateway, CI/CD (CodeBuild/CodePipeline or equivalent)
- Hands-on experience with at least one major MLOps toolset and awareness of alternatives: MLflow, Kubeflow, SageMaker Pipelines, Airflow, BentoML, KServe, Seldon.
- Deep understanding of model lifecycle management (feature engineering->training → registry → deployment → monitoring).
- Experience implementing or supporting LLMOps pipelines, including: prompt versioning, evaluation metrics, automation frameworks
- Deep understanding of ML lifecycle: data ingestion, feature engineering, training, evaluation, model packaging, CI/CD, drift detection, monitoring, and governance.
- Strong experience with AWS SageMaker (Pipelines, Feature Store, Model Registry, Model Monitor).
- Experience implementing ML CI/CD pipelines including automated training, testing, validation, model promotion, and endpoint deployment.
- Experience working on Infrastructure as Code (IaC) tools and CI/CD pipelines
- Experience with Kubernetes based development
- Experience with feature engineering pipelines and Feature Store management.
- Understanding of lineage tracking: training data snapshot, feature versions, code versioning, metadata tracking, reproducibility.
- Hands-on experience with AWS Bedrock and Agentcore service
- Experience with CloudWatch, SageMaker Model Monitor, Prometheus/Grafana.
- Strong foundation in Python and cloud-native development patterns.
- Solid understanding of security best practices, IAM, secrets management, and artifact governance.
Good to have skills:
- Experience with vector databases, RAG pipelines, or multi-agent AI systems.
- Exposure to DevOps and infrastructure-as-code (Terraform, Helm, CDK).
- Hands-on understanding of model drift detection, A/B testing, canary rollouts, and blue-green deployments.
- Familiarity with Observability stacks (Prometheus, Grafana, CloudWatch, OpenTelemetry).
- SQL and data transformation experience using Snowflake, Databricks, Spark.
- Ability to translate business goals into scalable AI/ML platform designs.
- Strong communication and cross-team collaboration skills.
- Ability to guide engineering teams through technical uncertainty and design choices.
Key Responsibilities:
- Architect and implement the MLOps strategy for the programme, ensuring alignment with the project proposal and delivery roadmap.
- Design and own enterprise-grade ML/LLM pipelines covering model training, validation, deployment, versioning, monitoring, and CI/CD automation.
- Build container-oriented ML platforms (EKS-first) while evaluating alternative orchestration tools with similar capabilities (Kubeflow, SageMaker, MLflow, Airflow, etc.).
- Implement hybrid MLOps + LLMOps workflows, including prompt/version governance, evaluation frameworks, and monitoring for LLM-based systems.
- Serve as a technical authority across multiple internal and customer projects, contributing architectural patterns, best practices, and reusable frameworks.
- Enable observability, monitoring, drift detection, lineage tracking, and auditability across ML/LLM systems.
- Define and implement standards for model deployment, monitoring, governance, and automation to ensure production-grade reliability and scalability.
- Collaborate with cross-functional teams — data engineering, platform, DevOps, and client stakeholders — to deliver production-ready ML solutions.
- Ensure all solutions adhere to security, governance, and compliance expectations, particularly around handling cloud services, Kubernetes workloads, and MLOps tools.
- Conduct architecture reviews, troubleshoot complex ML system issues, and guide teams through implementation across cloud-native ML platforms.
- Mentor engineers and provide guidance on modern MLOps tools, platform capabilities, and best practices.
If you like wild growth and working with happy, enthusiastic over-achievers, you'll enjoy your career with us!
Originally posted on Himalayas