首席数据科学家(自然语言处理 + 应用人工智能)
Principal Data Scientist (NLP + Applied AI)
职位描述:
我们相信大胆的想法、多元的视角,以及将知识转化为影响的动力。在这里,你的求知欲推动进步,你的声音塑造创新,你的雄心帮助重新定义科学和学习中可能的边界。我们是一个专注于影响、挑战并推动未来发展的文化,为我们的客户、同事以及整个社会创造无限可能。
职位简介:
我们正在构建系统,将世界上最大的科学文献库转化为研究智能。这意味着在数百万篇期刊文章上运行的生产级NLP流水线,提取实体、分类、主张元组和摘要,以优化下游代理应用的使用。我们正在寻找一位高级数据科学家,负责从评估集到最终部署的端到端领域特定内容建模工作。
你将加入一个小型、资深的团队,数据科学家在生产环境中负责自己的模型。你将编写代码、负责评估、部署变更,并对结果负责。这是一个需要亲自动手的角色,适合希望在快速变化的市场中看到自己的模型真正服务于用户的人员。
你将负责的工作:
- 设计并构建NLP增强流水线,从科学全文中大规模提取实体、分类、主张和摘要。
- 对比NLP方法与基于大语言模型(LLM)的方法在提取和增强方面的表现,为每项任务选择合适的工具。这意味着将传统NLP(命名实体识别、序列标注、分类)、基于嵌入的检索、LLM提示和微调的小型模型放在同一平台上进行比较,并用评估、成本和运营权衡来支持每项选择。这是工作的核心部分,而不是偶尔进行的活动。
- 负责评估。与领域专家和供应商合作建立黄金数据集,选择指标,并在速度、质量和成本之间做出有效的权衡。
- 编写高质量的Python代码。管理高负载LLM工作流的并发性和成本。编写结构清晰的代码,便于工程师和其他数据科学家进行扩展。
- 与数据工程师团队协作,使用Airflow和Dagster等数据流水线和数据构建工具协调工作。设计可重试、可评估、可靠的流水线阶段,在大规模运行失败时仍能保持稳定。
- 为代理AI应用工作做出贡献:能够使用工具并在增强后的语料库上进行推理的系统,你的NLP和评估背景将在这里发挥关键作用。
查看英文原文
Job Description:
We believe in bold ideas, diverse perspectives, and the drive to transform knowledge into impact. Here, your curiosity fuels progress, your voice shapes innovation, and your ambition helps redefine what’s possible within science and learning. We are a culture that obsesses over impact, challenges, and drives what’s next to power infinite possibilities for our customers, colleagues and society at large.
About the Role:
About the role
We'rebuilding the systems that turn one of the world's largest scientific corpora intoresearchintelligence. That meansproductionNLP pipelines running over millions of journal articles, extracting entities, classifications, claim tuples, and summariesoptimizedfor use by downstream agentic applications.We'relooking for a senior data scientist to owndomain-specificcontentmodeling work end to end, from the eval set through the pipeline stage that ships it.
You'lljoin a small, senior team where data scientists own their models in production.You'llwrite the code, own the evaluations, ship the changes, and stay accountable for the outcomes. This is a hands-on role for someone who wants to see their models through to real usersin a rapidly evolving market.
Whatyou'lldo
- Design and build NLP enrichment pipelines that extract entities, classifications, claims, and summaries from scientificfull-textat scale.
- Compare NLP approaches to extraction and enrichment against LLM-basedapproaches, andpick the right tool for each task. That means putting traditional NLP (NER, sequence labeling, classification), embedding-based retrieval, LLM prompting, and fine-tuned smaller models on the same table, and defending each choice with evaluation, cost, and operational tradeoffs. This is a core part of the job, not an occasional exercise.
- Own evaluation. Build the golden setsin consultation with SMEs and vendors, choose the metrics, and make productive tradeoffs between speed, quality, and cost.
- Write production-quality Python. Manage concurrency and cost for high-volume LLM workloads. Structure code that engineers canshipand other data scientists can extend.
- Collaborate with a team of data engineers to orchestrate work in datapipelineand data build tools like Airflow andDagster. Design idempotent,retryable, evaluable pipeline stages that stay reliable when a run fails at scale.
- Contribute to agentic AI application work: tool-using systems that reason over the enriched corpus, where your NLP and evaluation background will shape how the agent grounds and defends its answers.
- Work directly with editors, product managers, and engineers. Bring the modeling perspective into productdecisions, andtranslate stakeholderpushbackinto concrete modeling work.
Whatyou'llbring
- Deep Python.You'vewritten it in production, at scale, for years. You know when to reach forasyncioversus threads versus a queue, and you can explain the tradeoff clearly.
- Strong NLP background across modern (LLMs, transformers, embeddings, retrieval) and classical (NER, classification, sequence labeling) approaches.You'vebuilt evaluationsand learned from the results.
- A habit of comparing approaches and choosing the right one for the task. You can defend "prompt a large LLM" and "train a small classifier on 2,000 labels" with equal seriousness, back the choice with an eval and a costestimate, andknow what to do when performance drifts.
- A track recordof shipping – not just prototypes and papers, but systems that deliver value to real users.
Nice to have
- Experience working with scientific or scholarly text.
- Familiarity with AWS (S3, Batch, Lambda, SageMaker) and Parquet or Iceberg data lake patterns.
- Experience running LLMs underreal costand latency budgets in production.
- Some exposure to agentic AI applications: tool use, multi-step reasoning, guardrails, and evaluation of trajectories rather than single-turn outputs.
Why us
We publish some of the world's most-read research, andwe'renow in a rare position: applying modern AI to a corpus of trusted scientific knowledge that spans two centuries. Researchers will use the systems you build here to move faster and get closer to the answers theycame for.That'sthe work:from knowledge to impact.
We power infinite possibilities.
For more than 200 years, we've transformed knowledge into discoveries that shape the world. Today, our global team of innovators, creators, and experts is driving what's next in science, education, and publishing—creating impact that reaches everywhere.
We're not just observers of progress. We're the ones accelerating scientific breakthroughs, advancing learning, and sparking innovation that redefines entire fields and improves lives.
Here, your talent matters. Your ideas have room to grow. And your work creates breakthroughs that can change everything.
Wiley is an equal opportunity/affirmative action employer. We evaluate all qualified applicants and treat all qualified applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability, protected veteran status, genetic information, or based on any individual's status in any group or class protected by applicable federal, state or local laws. Wiley is also committed to providing reasonable accommodation to applicants and employees with disabilities. Applicants who require accommodation to participate in the job application process may contact for assistance.
We are proud that our workplace promotes continual learning and internal mobility. We offer meeting-free Friday afternoons allowing more time for heads down work and professional development, and through a robust body of employee programing we facilitate a wide range of opportunities to foster community, learn, and grow.
We are committed to fair, transparent pay, and we strive to provide competitive compensation in addition to a comprehensive benefits package. The range below represents Wiley's good faith and reasonable estimate of the base pay for this role at the time of posting roles in the United Kingdom, Canada, USA, Austria, Czechia, Denmark, France, Greece, Italy, Netherlands, Romania, or Spain. It is anticipated that most qualified candidates will fall within the range, however the ultimate salary offered for this role may be higher or lower and will be set based on a variety of non-discriminatory factors, including but not limited to, geographic location, skills, and competencies.
When applying, please attach your resume/CV to be considered.
Salary Range:
59,100.00 GBP to 84,633.33 GBPOriginally posted on Himalayas