高级应用科学家
Senior Applied Scientist
关于团队
Agentic Foundations 团队是 Zillow 更大规模 Agentic AI 组织中的一个专注于应用科学的小组,成员约 15-20 人,致力于构建支撑下一代购房体验的基础 AI 技术。该高级应用科学家将专注于模型后训练和奖励建模——这项科学使我们的代理能够作为自我演化的 Agentic 系统的一部分,持续提升性能。该职位将从头到尾推动复杂且模糊的工作,与产品、工程、科学、平台和数据团队紧密合作,将最新的大语言模型(LLM)、强化学习和后训练技术转化为实际生产影响。现在加入 Zillow 是一个激动人心的时刻,因为公司正在整个购房旅程中构建 Agentic 技能——从搜索和个性化指导,到报价策略和融资——每天处理数亿次请求。
职位描述
职责和要求:
- 负责 LLM 后训练流程——SFT、DPO 和强化微调(RFT/GRPO)——在 GPU 基础设施上端到端运行训练。
- 构建并训练快速、多类别的奖励模型(PRMs),从对比偏好对中输出标量信号。
- 设计在策略和在线评估:使用 LLM-as-a-Judge 框架对实时代理轨迹进行评分。
- 将离线评估标准转化为可泛化的奖励函数,适用于任意生产轨迹。
- 独立推动跨团队合作的复杂工作。
- 提供技术领导力和指导,同时帮助将最前沿的 AI 研究转化为实际生产影响。
该职位被归类为远程岗位。“远程”员工没有固定的公司办公场所,而是从自己选择的物理位置工作,必须向公司说明该地点。美国员工可以居住在美国 50 个州中的任何一州,有少数例外。在加利福尼亚州、康涅狄格州、马里兰州、马萨诸塞州、新泽西州、纽约州、华盛顿州和华盛顿特区,该职位的标准基本薪资范围为每年 160,900.00 至 257,100.00 美元。该薪资范围仅适用于这些地区,可能不适用于其他地区。在科罗拉多州、夏威夷州、伊利诺伊州、缅因州、明尼苏达州、内华达州、俄亥俄州、罗德岛州、佛蒙特州和弗吉尼亚州,该职位的标准基本薪资范围为每年 152,900.00 至 244,300.00 美元。该薪资范围仅适用于这些地区,可能不适用于其他地区。
查看英文原文
About the team
The Agentic Foundations team is a focused group of 15–20 applied scientists within Zillow's broader Agentic AI organization, building the foundational AI technologies that power the next generation of home shopping experiences. This Senior Applied Scientist will focus on model post-training and reward modeling — the science that lets our agents measurably improve over time as part of a self-evolving agentic system. The role will drive complex, ambiguous work end-to-end, partnering closely with product, engineering, science, platform, and data teams to translate the latest advances in LLMs, reinforcement learning, and post-training into production impact. This is an exciting time to join as Zillow is building agentic skills across the full home-shopping journey — from search and personalized guidance to offer strategy and financing — at a scale of hundreds of millions of requests per day.About the role
Role and responsibilities:
- Own LLM post-training pipelines — SFT, DPO, and reinforcement fine-tuning (RFT/GRPO) — running training end-to-end on GPU infrastructure.
- Build and train fast, multi-category reward models (PRMs) that emit scalar signals from contrastive preference pairs.
- Design on-policy and online assessment: score live agent trajectories using LLM-as-a-Judge frameworks.
- Translate offline evaluation rubrics into generalizable reward functions that work on arbitrary production traces.
- Independently drive complex work end-to-end by collaborating across partner teams.
- Provide technical leadership and mentorship while helping translate state-of-the-art AI research into production impact.
This role has been categorized as a Remote position. “Remote” employees do not have a permanent corporate office workplace and, instead, work from a physical location of their choice, which must be identified to the Company. U.S. employees may live in any of the 50 United States, with limited exceptions.In California, Connecticut, Maryland, Massachusetts, New Jersey, New York, Washington state, and Washington DC the standard base pay range for this role is $160,900.00 - $257,100.00 annually. This base pay range is specific to these locations and may not be applicable to other locations.In Colorado, Hawaii, Illinois, Maine, Minnesota, Nevada, Ohio, Rhode Island, Vermont, and Virginia the standard base pay range for this role is $152,900.00 - $244,300.00 annually. The base pay range is specific to these locations and may not be applicable to other locations.In addition to a competitive base salary this position is also eligible for equity awards based on factors such as experience, performance and location. Actual amounts will vary depending on experience, performance and location. Employees in this role will not be paid below the salary threshold for exempt employees in the state where they reside.Who you are
- Master's degree or higher in Computer Science or a related field.
- Hands-on experience with LLM post-training and RL fine-tuning: SFT, DPO, RFT/GRPO, running end-to-end training runs on GPU infrastructure.
- Exposure to reward model development — research-level experience is sufficient; production deployment is not required.
- Strong knowledge of generative AI, including foundation models, transformers, reinforcement learning, and preference learning.
- Ability to independently scope and solve ambiguous problems end-to-end while providing technical leadership to scientists and MLEs.
- Strong programming skills, especially Python, plus experience with ML frameworks such as PyTorch or TensorFlow.
Nice-to-haves:
- PhD in Computer Science, Machine Learning, or a related field.
- Experience evaluating agentic AI: multi-turn trace analysis, LLM-as-a-Judge, and the interaction between offline evals and online monitoring.
- Published work in post-training, RLHF/RLAIF, preference learning, or reward modeling.
- Experience with GPU training platforms such as Databricks or Fireworks.
Get to know us
At Zillow, we’re reimagining how people move—through the real estate market and through their careers. As the most-visited real estate platform in the U.S., we help people navigate buying, selling, financing and renting with greater ease and confidence. Whether you're working in tech, sales, operations, or design, you’ll be part of a company reshaping an industry and helping more people make home a reality.
How we work is a key part of that transformation. Through Cloud HQ, our strategic approach to distributed work, most employees have the flexibility to work from wherever they’re most productive—enabling us to move fast, stay connected and deliver on our people promise.
Zillow is honored to be recognized among the best workplaces in the U.S. Zillow was named one of FORTUNE 100 Best Companies to Work For® in 2026, and included on TIME’s America’s Best Companies 2026 list, reflecting our commitment to creating an innovative, inclusive, and engaging culture where employees are empowered to grow.
No matter where you sit in the organization, your work will help drive innovation, support our customers, and move the industry—and your career—forward, together.
Zillow Group is an equal opportunity employer committed to fostering an inclusive, innovative environment with the best employees. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. If you have a disability or special need that requires accommodation, please contact your recruiter directly.
Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable state and local law.
Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.
Originally posted on Himalayas