远程工作雷达

提示工程师(LLM系统、评估与安全)

Prompt Engineer (LLM Systems, Evals & Safety)

AI开发工程未标注地域
公司webook.com
薪资未公开
工作地点Jordan
地域资格未标注地域
时区要求日间重叠约 4 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
前往 Himalayas 查看并投递 →

你是否希望热爱自己的工作?是否希望有所作为,产生影响,改变人们的生活?是否希望与一个相信颠覆常规、枯燥和普通的团队一起工作?
如果是这样,那么这就是你正在寻找的工作。webook.com 是沙特阿拉伯在技术、功能、敏捷性和收入方面排名第一的活动票务和体验预订平台,服务着王国最大的大型活动,销售额超过 20 亿。

职位概述
设计高质量的提示、系统指令和工具,使我们的 LLM 功能准确、安全且成本可控。你将负责评估、提示版本管理和持续改进。

主要职责:

  • 编写、重构和组合提示(系统/工具/策略)以完成各种任务。
  • 创建离线/在线评估框架(评分标准、黄金数据集、指标)。
  • 构建带有版本控制、A/B 测试和遥测功能的提示库。
  • 通过验证、约束解码和工具使用来减少幻觉。
  • 实现安全性:破解/提示注入测试、内容策略检查、个人身份信息处理。
  • 与工程师合作,将提示集成到生产功能中。

要求:

  • 在多种任务类型和模型上展示过提示设计能力。
  • 有构建评估数据集和自动化评分的经验(例如准确性、忠实性、实用性、成本/延迟)。
  • 熟悉检索增强生成概念和工具/函数调用。
  • 强大的脚本编写能力(Python/TypeScript),用于数据准备、评估和分析。
  • 清晰的写作能力;能够将业务目标转化为可衡量的提示规范。

加分项:

  • 有 LangChain/LLM 协调、向量存储和重排序器的经验。
  • 了解安全工具和红队技术。
  • 实验平台(功能标志、A/B 测试)、分析工具。
查看英文原文

Do you want to love what you do at work? Do you want to make a difference, an impact, and transform peoples lives? Do you want to work with a team that believes in disrupting the normal, boring, and average?
If yes, then this is the job you are looking for , webook.com is Saudi’s #1 event ticketing and experience booking platform in terms of technology, features, agility, revenue serving some of the largest mega events in the Kingdom surpassing over 2 billion in sales.
Role Overview
Design high-quality prompts, system instructions, and tooling that make our LLM features accurate, safe, and cost-effective. You’ll own evaluation, prompt versioning, and continuous improvement.
Key Responsibilities:

  • Author, refactor, and chain prompts (system/tool/policy) for varied tasks.
  • Create offline/online evaluation harnesses (rubrics, golden sets, metrics).
  • Build prompt libraries with versioning, A/B testing, and telemetry.
  • Reduce hallucinations via verification, constrained decoding, and tool use.
  • Implement safety: jailbreak/prompt-injection tests, content policy checks, PII handling.
  • Partner with engineers to integrate prompts into production features.

Requirements

  • Demonstrated prompt design across multiple task types and models.
  • Experience building eval datasets and automated scoring (e.g., accuracy, faithfulness, utility, cost/latency).
  • Familiarity with retrieval-augmented generation concepts and tool/function calling.
  • Strong scripting (Python/TypeScript) for data prep, evals, and analysis.
  • Clear writing; ability to translate business goals into measurable prompt specs.

Nice-to-Haves

  • Experience with LangChain/LLM orchestration, vector stores, and rerankers.
  • Knowledge of safety tooling and red-teaming techniques.
  • Experiment platforms (feature flags, A/B tests), analytics.

Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

← 返回全部职位