高级应用AI工程师, 产品与代理性能
Staff Applied AI Engineer, Product & Agent Performance
Arcadia 是最值得信赖的医疗保健平台,推动成果实现。我们把复杂的医疗数据转化为可信的智能,帮助医疗机构、保险公司和生命科学组织清晰地行动,做出自信的决策,并实现可衡量的临床、运营和财务成果。
基于覆盖数千万患者生命的全面数据基础,Arcadia 结合先进的分析和负责任的 AI,挖掘有意义的见解,协调行动,并在大规模上提升绩效。我们对 AI 和自动化的做法是受控且透明的——旨在增强人类专业知识,而不是取代判断或掩盖责任。
数百家机构依赖 Arcadia 来改善成本、质量和成果。由 Nordic Capital 支持,我们继续投资于我们的平台、AI 能力和人才,以实现我们的使命:帮助医疗保健为每个人、每个社区和每一代人带来更好的成果。
为什么这个职位对 Arcadia 至关重要
Arcadia 的数据和分析平台被数百家医疗系统、ACO(基层医疗组织)、保险公司和生命科学组织使用,影响着数千万患者的健康。这个职位负责确保我们的代理功能在相同规模下表现良好:准确、对其自身信心透明,并且对依赖它们的临床医生、护理团队和患者安全。
作为一位资深个人贡献者,你将负责塑造代理行为的产品层决策,包括提示、检索和上下文、记忆和状态、评估和升级,同时与产品和工程团队合作,构建支持这些功能的系统。你的工作将帮助 Arcadia 做出基于证据的发布决策,并扩展可操控、可信赖且适合真实医疗工作流程的负责任的 AI。
什么是成功的表现
3 个月内
- 你已建立优先代理工作流的生产环境基准,包含文档化的失败模式、按严重性加权的评估标准以及明确的测量计划
- 你已映射当前的检索、上下文、记忆和升级模式,并识别出提高可靠性、校准和成本的最高价值机会
- 你通过将生产证据转化为清晰、可操作的建议,赢得了产品和工程团队的信任
6 个月内
- 代表生产环境的评估套件和回归检查为优先代理模型变更决策提供依据
查看英文原文
Arcadia is the most trusted healthcare platform powering outcomes. We transform complex healthcare data into trusted intelligence, helping providers, payers, and life sciences organizations act with clarity, make confident decisions, and achieve measurable clinical, operational, and financial outcomes.
Built on a comprehensive data foundation spanning tens of millions of patient lives, Arcadia combines advanced analytics and responsible AI to surface meaningful insights, coordinate action, and improve performance at scale. Our approach to AI and automation is governed and transparent — designed to strengthen human expertise, not replace judgment or obscure responsibility.
Hundreds of organizations rely on Arcadia to improve cost, quality, and outcomes. Backed by Nordic Capital, we continue to invest in our platform, AI capabilities, and people as we pursue our purpose: helping healthcare deliver better outcomes for every person, every community, and every generation.
Why This Role Is Important to Arcadia
Arcadia’s data and analytics platform is used by hundreds of health systems, ACOs, payers, and life sciences organizations, touching tens of millions of patient lives. This role owns how our agentic capabilities perform at that same scale: accurate, transparent about their own confidence, and safe for the clinicians, care teams, and patients who depend on them.
As a staff-level individual contributor, you will own the product-layer decisions that shape agent behavior, including prompting, retrieval and context, memory and state, evaluation, and escalation, while partnering with Product and Engineering on the systems that support them. Your work will help Arcadia make evidence-based launch decisions and scale responsible AI that is steerable, trustworthy, and ready for real healthcare workflows.
What Success Looks Like
In 3 months
- You have established a production-grounded baseline for priority agentic workflows, with documented failure modes, severity-weighted evaluation rubrics, and a clear measurement plan
- You have mapped the current retrieval, context, memory, and escalation patterns and identified the highest-value opportunities to improve reliability, calibration, and cost
- You have earned trust across Product and Engineering by turning production evidence into clear, actionable recommendations
In 6 months
- Production-representative evaluation suites and regression checks inform model-change decisions for priority agentic workflows
- You have delivered measurable improvements in accuracy, reliability, steerability, latency, or cost for one or more priority workflows
- Human-review and escalation behavior has been validated under adversarial and edge-case conditions, with decision criteria and ownership boundaries clearly documented
In 12 months
- Arcadia has a repeatable product-layer AI performance practice that moves from production failure to diagnosis, experiment, evaluation, and release decision
- High-severity regressions are caught earlier, and agent behavior is more transparent, calibrated, and trustworthy at scale
- Model cards, intended-use guidance, limitations, and performance documentation are current and useful to product and customer-facing teams
About Arcadia
Arcadia.io helps innovative providers and payers across the country transform healthcare to reduce cost while improving patient health. We do this by aggregating large amounts of disparate data, applying algorithms to identify opportunities to provide better patient care, and making those opportunities actionable by physicians at the point of care in near-real time. We are passionate about helping our customers drive meaningful outcomes. We are growing fast and have emerged as a market leader in the highly competitive population health management software market and have been recognized by industry analysts KLAS, IDC, Forrester, and Chilmark for our leadership. For a better sense of our brand and products, please explore our website.
Protect Yourself
If you have concerns about the authenticity of a job offer or recruitment-related communication claiming to be from Arcadia, we encourage you to verify by contacting us directly at (781) 202-3600 and select option 3. For more information, visit our website.
This position is responsible for following all Security policies and procedures in order to protect all PHI under Arcadia's custodianship as well as Arcadia Intellectual Properties. For any security-specific roles, the responsibilities would be further defined by the hiring manager.
What You'll Be Doing
- Design and iterate on agent behavior across real, live workflows, including long-horizon, multi-turn agentic tasks
- Design retrieval and context architecture so the right source data reaches a model in the right structure and agents remain grounded in real data rather than filling gaps with assumptions
- Design memory and state handling across multi-turn and multi-agent flows, determining what is carried forward, summarized, or dropped and why
- Create context and prompt templates that combine few-shot examples, structured formatting, and reasoning scaffolding for consistent agent behavior
- Improve performance through prompting, tool-use strategy, and context construction, validated through direct experimentation rather than guesswork
- Build and run evaluations against real production conditions to measure performance, regressions, failure modes, and edge cases
- Author evaluation rubrics, quality heuristics, and thresholds that weight failures by severity and cost, not just frequency, and monitor those measures against production behavior
- Design and validate escalation paths that route agents to human review based on confidence and uncertainty while preserving safety and consistency under adversarial and edge-case conditions
- Design for cost-aware performance alongside latency, reliability, and accuracy through efficient context construction and tool-call economy
- Evaluate and sign off on model changes by baselining current behavior, running comparative evaluations, and making the go/no-go call before a change reaches a customer
- Maintain product-level AI documentation, including model cards, intended use, limitations, and known failure modes, so customer-facing teams work from actual agent behavior
- Partner closely with Product and product managers to ensure agents are not just capable, but steerable, trustworthy, and ready to scale
What You'll Bring
- We value equivalent practical experience that demonstrates the depth required for this staff-level role
- 8+ years of production software engineering experience, including 3+ years of hands-on ownership of ML, LLM, or agentic systems in production, with direct experience in healthcare, finance, or another regulated industry
- Demonstrated ability to diagnose why an agent failed, correctly attribute the fix to instruction, retrieval, context, or memory design, and weigh failures by severity and cost rather than frequency alone
- Hands-on experience with RAG architecture, production-grounded evaluation frameworks, and fallback or human-in-the-loop logic for automated systems
- Working familiarity with AWS AI/ML services, including Bedrock and SageMaker, sufficient to build and evaluate effectively in Arcadia’s environment
- Evidence-led judgment and the credibility to push back on launch decisions, paired with a builder’s instinct to run the experiment and move from a production failure to a fix
Would Love for You to Have
- Experience applying AI to healthcare data or workflows where safety, transparency, and calibrated uncertainty directly affect care teams or patients
- Experience with long-horizon, multi-turn or multi-agent workflows and product-level AI documentation such as model cards
What You'll Get
- The opportunity to define how agent performance, safety, and readiness are measured for production healthcare workflows
- Meaningful ownership across prompts, context, memory, evaluations, and escalation patterns at product scale
- A cross-functional role translating production evidence into AI improvements used across Arcadia’s platform
- A mission-driven company working to improve how patients receive care
- A flexible, remote-friendly culture with personality and heart
- Employee-driven programs and initiatives for personal and professional development
- Membership in the talented, energized, diverse, and purpose-driven Arcadian community