湖仓机器学习工程师
Lakehouse Machine Learning Engineer
简介
你的未来,由你守护。ISC2是一支向善的力量。作为全球领先的非营利网络安全专业人士组织,我们的核心价值观——诚信、倡导、承诺、包容和卓越——推动我们为实现安全可靠的网络世界而努力。我们广受认可的获奖认证组合为各个职业阶段的网络安全知识、技能和经验提供了独立且全球认可的证明。我们的慈善机构——网络安全与教育中心,使ISC2和我们的成员能够通过教育最脆弱的人群了解网络风险,并赋予他们进入并在这个网络领域蓬勃发展所需的能力。在ISC2官网了解更多,并在Twitter、Facebook和LinkedIn上与我们联系。加入ISC2,你将展示你对包容和公平环境的承诺。你对全球网络安全从业人员和专业领域中独特观点和经验的支持将得到认可。我们邀请你积极参与帮助我们打造一种真正的归属感——一个真实、信任、赋权和连接的环境,让所有人的成功都得以实现。了解更多。
职位概述
ISC2正在构建一个现代的数据湖架构,并利用它在会员业务中运行机器学习。这个职位涵盖从生成数据的管道到利益相关者实际使用的最终结果的整个流程。湖架构机器学习(ML)工程师在受控的数据湖上编写生产级的Python和Spark代码,花时间探索数据后再决定哪些内容值得建模,并在模型上线后进行部署和监控。随着平台的成熟,也有机会参与应用大语言模型(LLM)的工作。该数据湖架构运行在Azure Databricks上;具备类似技术栈的深入经验更佳。
**此职位不面向加利福尼亚州居民**。
职责
- 通过青铜层、银层和金层构建和维护Python/Spark管道,以及消费这些数据的语义数据集和ML模型。
- 在建模之前深入分析数据。确定数据能支持什么,构建并测试由此产生的特征,然后与利益相关者协调哪些问题值得回答,哪些不值得。
- 涉及多种模型类型:生存分析和事件时间模型、预测模型、分类和倾向模型、序列模型、推荐系统
查看英文原文
Overview
Your Future. Secured. ISC2 is a force for good. As the world’s leading nonprofit member organization for cybersecurity professionals, our core values — Integrity, Advocacy, Commitment, Inclusion, and Excellence — drive everything we do in support of our vision of a safe and secure cyber world. Our globally recognized, award-winning portfolio of certifications provide an independent and globally recognized endorsement of cybersecurity knowledge, skills and experience for all career levels. Our charitable arm, the Center for Cyber Safety and Education, enables ISC2 and our members to serve the public by educating the most vulnerable about cyber risks and empowering access to enter and thrive in the cyber profession. Learn more at ISC2 online and connect with us on Twitter, Facebook and LinkedIn. When you join ISC2, you’ll demonstrate your commitment to an inclusive and equitable environment. Your support of the unique perspectives and experiences shared by our global cybersecurity workforce and profession will be recognized. We invite you to take an active role in helping us create a true sense of belonging across our organization — an environment of authenticity, trust, empowerment and connectedness that empowers all of our successes. Learn more.
Position Summary
ISC2 is building a modern lakehouse and using it to run machine learning across the membership business. This role covers the whole path, from the pipelines that produce the data through to results a stakeholder actually uses. The Lakehouse Machine Learning (ML) Engineer writes production Python and Spark against a governed lakehouse, spends real time exploring data before deciding what’s worth modeling, and deploys and monitors models once they’re live. As the platform matures there is room to take on applied LLM work. The lakehouse runs on Azure Databricks; deep experience on a comparable stack transfers.
**This position is not available to residents of California**.
Responsibilities
- Build and maintain Python/Spark pipelines through bronze, silver, and gold layers, plus the semantic datasets and ML models that consume them.
- Dig into the data before modeling it. Work out what it can support, build and test the features that come out of that exploration, then coordinate with stakeholders about which questions are worth answering and which aren’t.
- Work across a range of model types: survival and time-to-event, forecasting, classification and propensity, sequence models, recommenders, causal evaluation.
- Put models into production and keep them there - Experiment tracking, model registry, scheduled inference, and monitoring for drift and decay once they’re live.
- Get results to the people who need them. That means landing output in governed semantic tables feeding dashboards and CDP systems, and being able to walk a business team through what the numbers mean.
- Take on applied LLM work as the platform matures, including structured extraction from free text, and retrieval over governed data.
- Build inside the platform’s security and governance requirements rather than around them. Access controls, data protection, auditability, and human review where model output drives a decision that affects a member.
- Turn what works into reusable patterns, including project templates, shared feature and evaluation code, and implementation standards, so the next model doesn’t start from a blank notebook.
- Prove things out before they get built for real - Small proofs of concept that establish whether the data supports the theory, whether the approach holds up, and whether the result can actually be operated once it’s live.
- Perform miscellaneous duties, as required.
Behavioral Competencies
- Highly organized with strong attention to detail and documentation rigor.
- Collaborative, intellectually curious, and proactive in identifying analytical opportunities.
Qualifications
- Strong Extract/Transform/Load (ETL) skills, with the ability to assemble a dataset rather than request one.
- Fluent Python and SQL skills, comfortable working across enterprise source systems, and fluent with the standard ML stack — scikit-learn at minimum, and at least one deep learning framework such as PyTorch or TensorFlow.
- Knowledge range in ML, including knowing when timing matters enough to warrant survival analysis over a churn classifier, and being able to tell a causal question from a predictive one.
- Familiarity with hyperparameter tuning, cross-validation, and the techniques used to confirm a model holds up on data it has not seen.
- Ability to perform careful model validation, and to provide clarity about uncertainty. Also, to perform model evaluation and bias mitigation, judgment about where a person needs to stay in the loop, and the ability to explain the result to executives in non-technical terms.
- Understanding of data security, privacy, and compliance constraints, as well as how they shape what can be built with member data. Ability to work within the confines of access controls, data protection, and auditability.
- Knowledge of Databricks, including Unity Catalog, Workflows, MLflow, a plus.
- Ability to perform cohort-based or hierarchical forecasting at scale, a plus.
- Working knowledge of Salesforce, a plus.
- Relevant certifications: Databricks Data Engineer or Machine Learning Associate/Professional, Azure AI Engineer or Data Scientist Associate, or equivalent, a plus.
Education and Work Experience
- Bachelor’s or Master's degree in an IT field preferred. Will consider candidates with a high school diploma or equivalent and 7+ years of hands-on experience in data engineering and applied machine learning, preferably in enterprise environments.
- 3+ years of hands-on experience in data engineering and applied machine learning, preferably in enterprise environments.
- Experience deploying and monitoring models in production. ML flow or an equivalent tracking, registry, and scheduled-inference stack.
- Production experience building in a medallion architecture — bronze, silver, and gold, or an equivalent layered model — in a data-catalog-governed environment. Distributed processing with Spark or a comparable engine, an open table format such as Delta or Iceberg, and catalog-managed schemas, lineage, and access control. Demonstrated experience building and operating in that environment, not just querying it.
- Practical Large Language Model (LLM) experience including embeddings and retrieval, structured extraction, evaluation, a plus.
- Experience with subscription or membership-lifecycle data, a plus.
- Experience running a build-versus-buy evaluation. Hands-on assessment of tools and vendors, and a recommendation that can be defended, a plus.
Physical and Mental Demands
- Up to 5% travel may be required.
- Work normal business hours and extended hours when necessary.
- Remain in a stationary position, often standing or sitting, for prolonged periods.
- Regular use of office equipment in a remote environment such as a computer/laptop and monitor computer screens.
- Dexterity of hands and fingers to operate a computer keyboard, mouse, and other computer components.
Total Rewards
The pay range for this position is $93,000 - $118,900/Yr.
Final pay is based on several factors including but not limited to internal equity, market data, and the applicant’s education, work experience, certifications, etc.
Information regarding our comprehensive benefits package is available here.
This position will be posted for a minimum of 5 calendar days. This is a current vacancy, and the employer intends to fill this position within approximately 30 days.
Equal Employment Opportunity Statement
All qualified applicants will receive consideration for employment without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic as protected by applicable law. Job candidates will not be obligated to disclose sealed or expunged records of conviction or arrest as part of the hiring process.
Originally posted on Himalayas