数据科学家 / ML 平台工程师
Data Scientist / ML Platform Engineer
职责
我们正在寻找一名数据科学家/机器学习平台工程师,参与整个机器学习开发生命周期的工作——从模型构建和实验到生产部署和监控。核心职责包括应用数据科学和MLOps,同时对数据工程和轻量级平台运维有次要贡献。该职位在既定的平台模式下工作,与专门的基础设施工程师协作,无需他们参与日常的机器学习和数据任务。所有工作均在受HIPAA监管、符合FedRAMP标准的医疗分析环境中进行。
你将负责:
- 开发、训练和评估机器学习模型(分类、回归、聚类、异常检测),并为基于大语言模型的能力(如RAG管道和提示评估)做出贡献。
- 使用MLFlow支持模型治理和部署实践,包括实验跟踪、模型版本控制、注册表提升流程以及机器学习生命周期中的自动化测试。
- 参与生产机器学习运维:模型性能监控、漂移检测、自动化警报以及事件升级,以确保可靠性和SLA合规性。
- 构建和改进模型服务基础设施、特征管道和生命周期自动化,以支持可复现、可扩展的模型开发和推理。
- 应用可解释性技术(例如SHAP、LIME)并生成技术文档,以支持利益相关者的透明度和合规性要求。
- 在Snowflake和Databricks环境中使用Spark和基于SQL的框架,参与数据摄取、ELT/ETL转换和管道可靠性工作。
- 支持管道编排、银币架构规范和数据管理实践(元数据管理、PII处理、Unity Catalog中的血缘追踪)。
- 与平台团队合作执行偶尔的系统管理任务,包括环境配置、访问管理、计算故障排除和使用平台原生工具处理密钥。
资格要求
基本要求:
- 2年本科/学士学位;0年硕士/硕士学位;6年高中文凭/同等学历
- 具有SQL和Python的经验,包括使用Python-based的机器学习框架(例如scikit-learn、XGBoost、PyTorch或TensorFlow)。
- 具有使用MLFlow或其他类似工具进行实验跟踪、模型治理和生命周期管理的实际经验。
- 对SDLC基础有深刻理解,并有相关经验
查看英文原文
Responsibilities
We are looking for a Data Scientist / ML Platform Engineer to contribute across the full ML development lifecycle — from model building and experimentation to production deployment and monitoring. Core responsibilities are in applied data science and MLOps, with secondary contributions to data engineering and light platform operations. This role works within established platform patterns alongside dedicated infrastructure engineers, without requiring their involvement for routine ML and data tasks. All work is performed in a HIPAA-governed, FedRAMP-compliant healthcare analytics environment.
What you'll do:
- Develop, train, and evaluate ML models (classification, regression, clustering, anomaly detection) and contribute to LLM-based capabilities such as RAG pipelines and prompt evaluation.
- Support model governance and deployment practices using MLFlow, including experiment tracking, model versioning, registry promotion workflows, and automated testing across the ML lifecycle.
- Contribute to production ML operations: model performance monitoring, drift detection, automated alerting, and incident escalation to maintain reliability and SLA compliance.
- Build and improve model serving infrastructure, feature pipelines, and lifecycle automation to support reproducible, scalable model development and inference.
- Apply explainability techniques (e.g., SHAP, LIME) and produce technical documentation to support stakeholder transparency and compliance requirements.
- Contribute to data ingestion, ELT/ETL transformation, and pipeline reliability using Spark and SQL-based frameworks within Snowflake and Databricks environments.
- Support pipeline orchestration, medallion architecture conventions, and data stewardship practices (metadata management, PII handling, lineage tracking in Unity Catalog).
- Perform occasional system administration tasks in collaboration with platform teams, including environment configuration, access management, compute troubleshooting, and secrets handling using platform-native tools.
Qualifications
Basic Qualifications:
- 2 years with BS/BA; 0 years with MS/MA; 6 years with HS Diploma/equivalent
- Demonstrated experience with SQL and Python, including Python-based ML frameworks (e.g., scikit-learn, XGBoost, PyTorch, or TensorFlow).
- Hands-on experience with MLFlow or equivalent tools for experiment tracking, model governance, and lifecycle management.
- Strong understanding of SDLC fundamentals and experience with GitHub or equivalent version control.
- Experience with distributed compute environments (e.g., Spark, Databricks) and cloud-native services.
- Basic proficiency with Bash or shell scripting for automation and environment setup.
- Ability to collaborate across multidisciplinary teams and communicate technical concepts to varied audiences.
- Ability to obtain and maintain a Public Trust clearance
- US citizenship required
Preferred Qualifications:
- Experience with MLOps practices including CI/CD for ML, containerization, feature pipeline automation, and model deployment frameworks.
- Experience with Databricks E2 components (Unity Catalog, Feature Store, Delta Live Tables) and/or model serving and drift monitoring tools (e.g., Databricks Model Serving, Evidenly, etc.).
- Experience with LLM frameworks (e.g., LangChain, LlamaIndex, Hugging Face Transformers) and familiarity with model explainability libraries (e.g., SHAP, LIME).
- Advanced Spark performance optimization experience and/or API development using Databricks REST APIs.
- Experience with healthcare analytics data (preferably Medicare or Medicaid) and familiarity with HIPAA or FedRAMP compliance constraints.
- Experience building data pipelines in a Snowflake or Databricks environment.
- Familiarity with orchestration tools (Airflow, Databricks Workflows).
- Exposure to streaming data patterns using Spark Structured Streaming, Delta Live Tables, or Kafka.
- Familiarity with environment reproducibility tooling (Docker, conda) and scripting (Python, Bash) to support automation and CI/CD tasks
Peraton Overview
Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees do the can’t be done by solving the most daunting challenges facing our customers. Visit peraton.com to learn how we’re keeping people around the world safe and secure.
Target Salary Range
$80,000 - $128,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual’s experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.EEO
EEO: Equal opportunity employer, including disability and protected veterans, or other characteristics protected by law.Originally posted on Himalayas