[8BE] 数据科学家 (AI + ML)
[8BE] Data Scientist (AI + ML)
我们是Software Mind,一支充满激情的工程师团队,随时准备为任何顶级公司的项目注入动力!我们的目标是始终领先一步。加入一家持续增长的多元文化公司,这里的工作环境获得Great Place To Work认证!
我们正在寻找一位在概率AI和统计机器学习方面有深厚专业知识的数据科学家,以支持客户基于分布式微服务架构的电商平台。该平台的发展路线包括一系列需要严格统计建模而非标准监督学习的智能功能。该职位负责设计、验证和将概率模型投入生产,并与后端工程团队紧密合作,将这些模型转化为平台现有微服务生态系统中的服务化生产架构。
项目时长:3 - 6个月。
你将负责
· 贝叶斯建模与推断:设计并实现贝叶斯统计模型——先验、似然和后验推断,以支持定价、细分和需求相关的不确定性决策。
· 马尔可夫链和隐马尔可夫模型:构建马尔可夫链和隐马尔可夫模型公式,用于序列和行为模式(例如客户生命周期阶段、状态转换),生成下游服务可以使用的输出。
· MCMC和Metropolis-Hastings采样:应用马尔可夫链蒙特卡洛方法,包括Metropolis-Hastings采样,以估计没有闭式解的模型的后验分布,并验证收敛性和采样质量。
· 混合建模:开发混合模型——特别是高斯混合模型——以支持细分用例,从交易和行为数据中识别潜在的客户或产品分组。
· 期望最大化:实现期望最大化算法,用于混合模型及相关无监督学习任务中的潜在变量估计。
· 生产转化:与后端工程团队合作,将统计模型转化为生产服务架构——定义API、数据契约和平台现有微服务及事件驱动流程中的集成点。
· 模型生命周期管理:制定模型训练、验证、版本控制、监控/漂移检测和重新训练频率的方法,一旦模型投入生产。
· 路线图协作:参与与产品和工程团队的协作,确保模型符合业务目标和技术要求。
查看英文原文
We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!
We are seeking a Data Scientist with deep expertise in probabilistic AI and statistical machine learning to support a client's e-commerce platform, built on a distributed microservices architecture. The platform roadmap includes a set of intelligence capabilities that require rigorous statistical modeling rather than standard supervised ML. This role is responsible for designing, validating, and productionizing probabilistic models, and for working closely with backend engineering to translate those models into service-oriented production architecture within the platform's existing microservices ecosystem.
Project Length: 3 - 6 months.
What you will do
· Bayesian modeling and inference: Design and implement Bayesian statistical models — priors, likelihoods, and posterior inference — to support decisioning under uncertainty across pricing, segmentation, and demand-related use cases.
· Markov chains and Hidden Markov Models: Build Markov chain and Hidden Markov Model formulations for sequential and behavioral patterns (e.g., customer lifecycle stages, state transitions), producing outputs that downstream services can consume.
· MCMC and Metropolis-Hastings sampling: Apply Markov Chain Monte Carlo methods, including Metropolis-Hastings sampling, to estimate posterior distributions for models without closed-form solutions, and validate convergence and sampling quality.
· Mixture modeling: Develop mixture models — Gaussian Mixture Models in particular — to support segmentation use cases, identifying latent customer or product groupings from transactional and behavioral data.
· Expectation-Maximization: Implement Expectation-Maximization for latent-variable estimation underlying mixture models and related unsupervised learning tasks.
· Production translation: Work with backend engineering to translate statistical models into production service architecture — defining APIs, data contracts, and integration points within the platform's existing microservices and event-driven pipelines.
· Model lifecycle management: Define the approach for model training, validation, versioning, monitoring/drift detection, and retraining cadence once models are in production.
· Roadmap collaboration: Partner with delivery and engineering leads to size, sequence, and estimate probabilistic/statistical modeling initiatives on the product roadmap.
· Documentation and handoff: Document modeling assumptions, methodology, and validation results, and provide clear hand-off guidance so models remain maintainable by the engineering team after the engagement.
- +90% English written and oral (at least B2 level) with excellent communication skills
- Strong, demonstrable background in Bayesian statistics/Bayesian inference, Markov chains, Hidden Markov Models, MCMC methods (including Metropolis-Hastings sampling), mixture models (ideally Gaussian Mixture Models), and Expectation-Maximization.
- Proven experience building and deploying statistical/ML models into production systems, not just research notebooks or offline analysis.
- Proficiency in Python (or R) with standard probabilistic/statistical libraries (e.g., PyMC, Stan, scikit-learn, NumPy/SciPy) for model development and validation.
- Ability to translate statistical/mathematical models into service-oriented production architecture — defining APIs and data contracts and working directly with backend engineers to integrate them.
- Solid understanding of version control, testing practices, and CI/CD, sufficient to collaborate effectively with an engineering team on production delivery.
- Strong written and verbal communication skills, with the ability to explain model behavior, assumptions, and uncertainty to non-technical stakeholders.
Preferred Qualifications
· Experience in e-commerce or retail domains, particularly pricing optimization, customer segmentation, or demand forecasting.
· Experience integrating ML models with microservices architectures (REST/GraphQL) and event-driven systems (e.g., message queues/pub-sub), and deploying to cloud infrastructure.
· Familiarity with common backend service ecosystems (e.g., .NET, Java, or Node.js) — even if modeling itself is done in Python — for a smoother handoff to the production engineering team.
· Experience with MLOps tooling such as model registries, monitoring, and feature stores.
· Background in pricing science, recommendation systems, or marketing analytics.
Our Benefits
· Flexible schedule and Work From Anywhere
· Referral Program
· Supportive and chill atmosphere
We are accepting applications from LATAM countries