研究科学家
Research Scientist
### 公司简介
模型以数据为食。但大量训练计算被浪费在已经学习过的、无关的甚至有害的数据上,导致模型质量下降,训练和部署成本增加。
在 DatologyAI,我们构建了一个最先进的数据整理套件,可以自动整理和优化 PB 级的数据,为您的模型创建最佳的训练数据。在整理后的数据上进行训练可以大幅减少训练时间和成本(根据使用场景,训练速度提高 **7-40 倍**),大幅提升模型性能,**相当于您在 10 倍以上的原始数据上训练**,而不会增加训练成本,并且允许参数少于一半的小型模型在推理时使用更少的计算资源的情况下超越大型模型,显著降低部署成本。更多细节,请查看我们最近关于合成数据扩展的研究([BeyondWeb](https://www.datologyai.com/blog/beyondweb))以及使用领域特定数据进行预训练的研究([The Finetuner’s Fallacy](https://www.datologyai.com/blog/finetuners-fallacy))。
我们在两轮中总共筹集了 5750 万美元,包括种子轮和 A 轮融资。我们的投资方包括 Felicis Ventures、Radical Ventures、Amplify Partners、微软、亚马逊,以及像 Geoff Hinton、Yann LeCun、Jeff Dean 等人工智能领域的先驱者,他们深刻理解识别和优化模型最佳训练数据的重要性和难度。我们的团队在这一前沿研究领域进行了开创性工作,拥有数据研究和数据工程方面的深厚专业知识,能够解决这个极具挑战性的问题,并让任何想要用自己的数据训练模型的人都能轻松实现数据整理。
该职位位于加利福尼亚州圣马特奥。我们每周有 4 天在办公室工作。
### 职位简介
我们正在寻找一名研究科学家,探索如何干预训练数据以提升深度学习模型的质量并塑造其行为。您将从文献中获取和实现想法,开展基于真实客户需求的研究,并与工程师和产品团队紧密合作,将研究成果转化为实际影响。该职位需要强大的科学判断力,熟悉深度学习文献,并具备在快速发展的初创环境中自主工作的动力。
### 您将参与的工作
- 研究文献浩如烟海,充满歧义且不断演变。您将收集、验证并实现这些文献中的想法,同时结合实际客户的需求进行研究,并与工程师和产品团队密切合作,将研究成果转化为实际影响。
查看英文原文
### About the Company
Models are what they eat. But a large portion of training compute is wasted training on data that are already learned, irrelevant, or even harmful, leading to worse models that cost more to train and deploy.
At DatologyAI, we’ve built a state of the art data curation suite to automatically curate and optimize petabytes of data to create the best possible training data for your models. Training on curated data can dramatically reduce training time and cost ( **7-40x faster training** depending on the use case), dramatically increase model performance **as if you had trained on >10x more raw data** without increasing the cost of training, and allow smaller models with **fewer than half the parameters** to outperform larger models despite using far less compute at inference time, substantially reducing the cost of deployment. For more details, check out our recent research on synthetic data scaling ( [BeyondWeb](https://www.datologyai.com/blog/beyondweb)) and pretraining with domain-specific data ( [The Finetuner’s Fallacy](https://www.datologyai.com/blog/finetuners-fallacy)).
We raised a total of $57.5M in two rounds, a Seed and Series A. Our investors include Felicis Ventures, Radical Ventures, Amplify Partners, Microsoft, Amazon, and AI visionaries like Geoff Hinton, Yann LeCun, Jeff Dean, and many others who deeply understand the importance and difficulty of identifying and optimizing the best possible training data for models. Our team has pioneered this frontier research area and has the deep expertise on both data research and data engineering necessary to solve this incredibly challenging problem and make data curation easy for anyone who wants to train their own model on their own data.
This role is based in San Mateo, CA. We are in office 4 days a week.
### About the Role
We're looking for a Research Scientist to investigate how intervening on training data can improve the quality and shape the behavior of deep learning models. You'll source and implement ideas from the literature, conduct research grounded in real customer needs, and collaborate closely with engineers and product teams to turn findings into tangible impact. This role requires strong scientific judgment, fluency with the deep learning literature, and the drive to work autonomously in a fast-moving startup environment.
### What You'll Work On
- The research literature is vast, rife with ambiguity, and constantly evolving. You'll source, vet, implement, and improve promising ideas from the literature and your own thinking.
- Our research is guided by concrete customer needs and product outcomes, not conference benchmarks. You'll have clear context on why your work matters and who it serves.
### How You'll Work
- We believe researchers do their best work with autonomy. You'll have the freedom to pursue problems in the way that works best for you, with the resources and context to back it up.
- We expect Research Scientists to collaborate closely with engineers, talk to customers, and shape the product vision.
### About You
- 3+ years of deep learning research experience
- Strong fundamentals in deep learning
- Practical experience and/or publications in one or more of the following areas:
- Data pruning and curation
- Curriculum learning
- Synthetic data generation
- Dataset distillation
- Effects of training data on model behavior
- Embedding models and semantic search
- Training large vision (including video), language, or multimodal models
- Efficient ML
- Enough software engineering and PyTorch experience (or willingness to learn) to run large-scale experiments and build production prototypes
- A demonstrated track record in deep learning research, whether through papers, tools, or other artifacts
**Nice to have:**
- Experience with distributed data processing tools like Spark or Snowflake
- Experience building and shipping ML products
Candidates do not need a PhD or extensive publications. Some of the best researchers we've worked with have no formal training in machine learning, and obtained all of their experience working in industry and building products. We believe adaptability, combined with exceptional communication and collaboration skills, are the most important ingredients for successful research in a startup environment.
### **Compensation**
At DatologyAI, we are dedicated to rewarding talent with competitive salary and meaningful equity. The salary for this position ranges from $180,000 to $300,000.
- Starting pay is based on job-related skills, experience, qualifications, and interview performance.
**Benefits:**
- 100% covered health benefits (medical, vision, and dental).
- 401(k) plan with a generous 4% company match.
- Unlimited PTO policy
- Paid Parental Leave of 12 weeks, plus 6 months of WFH flexibility.
- Annual $2,000 wellness stipend.
- Annual $1,000 learning and development stipend.
- Daily lunches and snacks are provided in our office!
- Relocation assistance for employees moving to the Bay Area.