远程工作雷达

高级AI ML运维工程师

Senior AI ML Operations Engineer

AI开发工程职能支持限定地区(需当地身份)
公司Smartsheet
薪资未公开
工作地点India
地域资格限定地区(需当地身份)
时区要求日间重叠约 7 小时,基本正常作息
用工类型permanent
发布时间2026-08-11
数据来源4dayweek.io
前往 4dayweek.io 查看并投递 →
注意地域限制:该职位明确限定在 India 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。

超过20年,Smartsheet 一直赋能团队无缝管理任务并更智能地扩展解决方案。现在,在我们最雄心勃勃的篇章中,我们将人类团队与AI代理结合在一起。通过协调代理擅长的工作,自动化手动任务,并在大规模中发现见解,我们为人们腾出空间,专注于真正重要的事情:判断力、创造力和大思维。这就是工作中的魔法,也是我们每天所追求的。

职位描述/职责:

- 设计、开发和维护稳定可靠的AI/ML Ops平台/流程
- 模型部署:将AI/ML服务打包并部署到生产环境,确保其可复现和可解释
- CI/CD流程开发:设计并实现自动化CI/CD(持续集成/持续部署)流程,使用工具加速模型部署
- 基础设施管理:为训练和推理提供并优化基础设施,使用Docker、Kubernetes或无服务器平台
- 监控与可观测性:使用工具实施部署后的监控,包括模型性能、数据漂移和延迟。有Monte Carlo经验者优先
- 自动化:自动化重新训练和数据流程工作流,确保模型随时间保持准确性。
- 管理基础模型的部署、微调流程以及检索增强生成(RAG)堆栈(向量数据库、知识图谱)。有AWS Bedrock经验者优先
- 资源优化:管理GPU/CPU利用率,以最小化云成本同时保持用户的低延迟推理
- 协作:与数据科学家、数据工程师和软件工程师紧密合作,弥合模型开发与生产之间的差距。
- 版本控制与治理:使用MLflow等工具管理数据、代码和模型的版本控制。
- 安全与合规:实施数据安全措施,确保符合数据治理政策,保护敏感数据
- 技术评估与创新:了解新兴数据技术,探索创新机会以改进组织的数据基础设施
- 故障排除与问题解决:诊断和解决复杂的数据相关问题,确保数据平台的稳定性和可靠性
- 执行其他指派的任务

**所需技能:**

- 企业SaaS软件解决方案,具有高可用性和可扩展性
- 处理大规模结构化和非结构化数据的经验

查看英文原文

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.

Job Description/ Responsibilities:

- Designing, developing and maintaining stable and reliable AI/ML Ops platforms / pipelines
- Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable
- CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools
- Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms
- Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable
- Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time.
- Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable
- Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users
- Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production.
- Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow.
- Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data
- Technology Evaluation and Innovation: Staying abreast of emerging data technologies and exploring opportunities for innovation to improve the organisation’s data infrastructure
- Troubleshooting and Problem Solving: Diagnosing and resolving complex data-related issues, ensuring the stability and reliability of the data platform
- Perform other duties as assigned

**Required Skills:**

- Enterprise SaaS software solutions with high availability and scalability
- Solution handling large scale structured and unstructured data from varied data sources
- Experience in building and maintaining AI/ML Ops platform systems ensuring scalability, reliability, efficiency and security
- **In depth experience in AI/ML Frameworks and tools involving large Petabytes of data with Databricks Lakehouse ecosystem**
- **AI/MLOps workflows on Databricks , MLFlow, Mosaic AI Agent Framework, Unity Catalog, Vector Search, Knowledge Graph**
- **Knowledge of AI/ML frameworks like LangChain, LangGraph for AI/ML Ops pipeline integration**
- Cloud Platforms: Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP). Experience in AWS hosted data platform is preferable
- Programming languages like Python and SQL
- Modern software engineering practices like Kubernetes, CI/CD, IAC tools (Preferably Terraform), Observability, monitoring and alerting
- Solution Cost Optimisations and design to cost
- Legally eligible to work in India on an ongoing basis

**Get to Know Us:**

At Smartsheet, your ideas are heard, your potential is supported, and your contributions have real impact. You’ll have the freedom to explore, push boundaries, and grow beyond your role. We welcome diverse perspectives and nontraditional paths—because we know that impact comes from individuals who care deeply and challenge thoughtfully. When you’re doing work that stretches you, excites you, and connects you to something bigger, that’s magic at work. Let’s build what’s next, together.

**Equal Opportunity Employer:**

Smartsheet is an Equal Opportunity (EEO) employer committed to fostering an inclusive environment with the best employees. It is our policy to provide equal employment opportunities to all qualified applicants in accordance with applicable laws in the US, UK, Australia, Germany, Costa Rica, Japan, Bulgaria, India, and Singapore. All qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran or disabled status, or genetic information.

If there are preparations we can make to help ensure you have a comfortable and positive interview experience, please let us know.

#LI-Remote

本页面信息整理自 4dayweek.io,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

该公司其他在招职位

公司律师

SmartsheetAustraliapermanent今天
职能支持限定地区(需当地身份)

高级产品设计师 I

SmartsheetUnited States$135,000 - $190,000/年permanent5 天前
设计限定地区(需当地身份)与中国几乎无重叠,需长期倒时差

← 返回全部职位