远程工作雷达

高级人工智能工程师(荷兰)

Senior AI Engineer (Netherlands)

AI开发工程限定地区(需当地身份)日间重叠约 2 小时,需偶尔早起或晚睡
公司Starlims Corporation
薪资未公开
工作地点Netherlands
地域资格限定地区(需当地身份)
时区要求日间重叠约 2 小时,需偶尔早起或晚睡
用工类型Full Time
发布时间昨天
数据来源Himalayas
前往 Himalayas 查看并投递 →
注意地域限制:该职位明确限定在 Netherlands 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。
作息提示:日间重叠约 2 小时,需偶尔早起或晚睡。

职位描述:
我们正在将人工智能融入STARLIMS,这是一个用于质量制造、生命科学、公共卫生、法医学和环境科学的平台。
此职位专注于代理系统:能够对任务进行推理、调用工具、完成多个步骤,并将结果交给人员审核和批准的软件。
我们的用户在严格的准确性、可追溯性和验证要求下工作。工程挑战在于使非确定性系统足够可靠、可观测和可控,以被信任、测试和发布。
你将同时参与代理执行的平台和运行时,以及基于其上的生产代理。
你将负责的工作:
代理平台与运行时(核心重点)

  • 设计并构建代理执行的运行时:规划和执行循环、工具调用、状态管理、持久化执行和故障恢复
  • 构建代理安全访问平台数据和外部系统的层
  • 在需要时设计代理和工作流之间的协调、委派和交接机制
  • 使代理行为在版本控制、测试、测量和发布中的回归安全
  • 构建可重用的基本组件,使新代理可以通过配置而非重新开发来创建

构建代理(核心重点)

  • 从专家对话中将领域工作流程转化为可用代理:目标、动作、执行流程、故障处理和成功标准
  • 将代理决策和输出建立在权威企业数据基础上,而不是仅依赖模型知识
  • 通过设计实现人机协同:包括审批节点、覆盖捕获、不确定性处理和代理决策的明确证据。代理推荐并起草;人员决定
  • 闭合循环:将用户的更正和覆盖作为可衡量提升代理性能的信号

评估与可靠性

  • 构建多步骤行为的评估框架,而非单次响应的准确性:任务完成度、工具调用正确性、依据性、轨迹质量,以及模型、提示和工具变化后的回归情况
  • 定义代理质量、可靠性、延迟、成本和人工干预率的生产指标
  • 实现防护机制、回退方案、超时、成本上限,以及代理运行的端到端可观测性和追踪
  • 设计针对提示注入、不安全工具使用、权限过度、数据泄露等代理特定安全风险的防护措施
  • 管理提示演变和模型漂移
查看英文原文

About the Role:
We’re building AI into STARLIMS, a platform used across quality manufacturing, life sciences, public health, forensics, and environmental sciences.
This role is focused on agentic systems: software that reasons over a task, calls tools, works through multiple steps, and hands the result to a person to review and approve.
Our users work under strict accuracy, traceability, and validation requirements. The engineering challenge is making non-deterministic systems reliable, observable, and controllable enough to be trusted, tested, and shipped.
You’ll work on both the platform and runtime our agents execute on and the production agents built on top of it.
What You’ll Work On:
Agent Platform & Runtime (Core Focus)

  • Design and build the runtime our agents execute on: planning and execution loops, tool calling, state management, durable execution, and failure recovery
  • Build the layer through which agents reach platform data and external systems safely
  • Design coordination, delegation, and handoff across agents and workflows where needed
  • Make agent behavior versionable, testable, measurable, and regression-safe across releases
  • Build reusable primitives so new agents are configured rather than rebuilt from scratch

Building Agents (Core Focus)

  • Take a domain workflow from expert conversation to a working agent: goals, actions, execution flow, failure handling, and success criteria
  • Ground agent decisions and outputs in authoritative enterprise data rather than relying on model knowledge alone
  • Implement human-in-the-loop by design, including approval gates, override capture, uncertainty handling, and clear evidence for agent decisions. Agents recommend and draft; people decide
  • Close the loop: turn user corrections and overrides into signals that measurably improve the agent

Evaluation & Reliability

  • Build evaluation harnesses for multi-step behavior, not single-response accuracy: task completion, tool-call correctness, groundedness, trajectory quality, and regression across model, prompt, and tool changes
  • Define production metrics for agent quality, reliability, latency, cost, and human intervention rates
  • Implement guardrails, fallbacks, timeouts, cost ceilings, and end-to-end observability and tracing across agent runs
  • Design safeguards against prompt injection, unsafe tool use, excessive permissions, data leakage, and other agent-specific security risks
  • Manage prompt evolution, model drift, and non-determinism while maintaining consistent, measurable system behavior across releases

Integration & Data

  • Integrate agents with platform APIs and third-party enterprise systems already running in our customers’ environments
  • Build retrieval and context pipelines that turn fragmented enterprise data into reliable, permission-aware agent context
  • Design controlled execution paths for automated actions, with a complete, traceable audit trail

Platform & Infrastructure

  • Build and operate backend services on AWS (Lambda, API Gateway, DynamoDB, Step Functions, etc.)
  • Own significant parts of the system architecture and contribute to key technical decisions
  • Contribute to infrastructure-as-code and deployment pipelines

Tech Stack

  • Languages: TypeScript, Python
  • Backend: Node.js, Python, AWS Lambda, Step Functions
  • AI: OpenAI, Anthropic, MCP and related agent/tool protocols, embeddings and vector search
  • Frontend: React, Next.js, Tailwind CSS
  • Infrastructure: AWS, Terraform
  • Testing: Jest, Playwright, pytest

What We’re Looking For
Must Have

  • 6+ years of software engineering experience, including production systems
  • Experience building production LLM systems, including tool-using or multi-step agentic workflows beyond simple prompting and chat interfaces
  • Strong understanding of LLM behavior, limitations, and failure modes, especially how errors compound across a multi-step run
  • Experience with LLM APIs, tool and function calling, and designing planning and execution loops
  • Experience evaluating and debugging non-deterministic systems
  • Solid backend and cloud experience (AWS or equivalent)
  • Proficiency in TypeScript and/or Python

You Should Be Comfortable With

  • Debugging across distributed and non-deterministic systems
  • Making explicit tradeoffs between accuracy, latency, reliability, and cost
  • Working in ambiguous problem spaces where the right architecture isn't obvious yet
  • Owning production systems end-to-end
  • Choosing conventional software over AI when AI isn't the right solution

Nice to Have

  • C#, Microsoft .NET Framework
  • Tool and interop protocols such as MCP
  • Evaluation pipelines and metrics built specifically for agentic systems
  • Experience in regulated or domain-heavy systems (validation, audit trails, controlled change)
  • Retrieval and grounding techniques for supplying agent context
  • Workflow and durable-execution platforms (Temporal, Step Functions, n8n, etc.)
  • Containerization and orchestration (ECS, EKS, Kubernetes)
  • Infrastructure as Code (Terraform or similar)

STARLIMS is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, creed, religion, color, national or ethnic origin, citizenship, sex, sexual orientation, gender identity and expression, genetic information, veteran status, age or disability status. Originally posted on Himalayas

本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。 本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。

← 返回全部职位