AI 应用全栈工程师
AI Application Full Stack Engineer
职位描述
我们正在构建一个内部AI平台,公司各产品团队都基于此进行开发。它是数据、AI和工程团队与他们使用的模型之间的中间层,这些模型可能来自外部供应商,也可能在我们的GPU上运行。通过这些团队,它服务于成千上万的企业和数百万用户。
你将负责从头到尾完成工作:API、其背后的的服务、数据模型、人们操作它的界面、测试,以及在真实负载下的生产环境行为。
你将负责的工作
- 设计并构建其他工程团队每天依赖的安全REST API和实时流式端点
- 构建在并发、部分失败和流量激增情况下仍能保持正确的后端服务,并且在业务高峰时段达到99.9%的可用性目标
- 构建使复杂系统易于理解的网页界面,使操作人员无需阅读源代码即可了解状态并作出响应
- 在PostgreSQL中仔细建模数据,并编写可以在成本和正确性方面经得起推敲的查询
- 构建并运行统一的API网关,覆盖外部模型供应商和你部署和运营的自托管模型,具备低延迟路由、负载均衡和故障转移功能,使使用我们API的团队永远看不到后端差异
- 构建多租户边界:认证(OAuth2)、基于角色的访问控制、配额和限速策略,确保在失败时关闭而非泄露
- 准确测量使用情况和成本,以支持报告和计费
- 为所交付的内容添加监控,并在事件中利用这些监控找到真正原因,而不是可能的原因
- 参与交付流程:代码审查、CI/CD、渐进式发布,以及后续的调试
- 与产品经理和内部使用该平台的团队紧密合作,将他们的需求转化为实际采用的解决方案
- 与技术项目经理合作管理开发周期:概念、设计、测试、发布和支持
- 与DevOps紧密合作,在我们的基础设施中运营和维护该平台
- 保持团队知识的记录:技术需求、API契约、部署说明和事后分析
- 指导其他工程师,通过评审提升设计和代码质量的标准
要求
工程
- 4年以上在团队环境中进行软件工程的经验,构建和运行生产级Web应用
- 精通JavaScript,
查看英文原文
About the role
We are building the internal AI platform that product teams across the company build on. It is the layer between our data, AI and engineering teams and the models they use, whether those come from external providers or run on our own GPUs, and through those teams it reaches thousands of businesses and millions of users.
You will own work end to end: the API, the services behind it, the data model, the interface people operate it through, the tests, and how it behaves in production under real load.
What you will do
- Design and build secure REST APIs and real-time streaming endpoints that other engineering teams depend on daily
- Build backend services that stay correct under concurrency, partial failure, and traffic spikes, and that hold a 99.9% availability target through peak business hours
- Build web interfaces that make a complex system legible, so an operator can understand state and act on it without reading source code
- Model data carefully in PostgreSQL and write queries you can defend on cost as well as correctness
- Build and run a unified API gateway over external model providers and the self-hosted models you deploy and operate on our own GPUs, with low-latency routing, load balancing, and failover, so the teams who consume our APIs never see the differences between backends
- Build multi-tenant boundaries that hold: authentication (OAuth2), role-based access control, quotas, and rate limits that fail closed rather than leak
- Measure usage and cost accurately enough to report and bill from
- Instrument what you ship, and use that instrumentation during incidents to find the real cause rather than a plausible one
- Take part in delivery: code review, CI/CD, progressive rollout, and the debugging that follows a bad release
- Work closely with Product Managers and the internal teams who consume the platform to turn their needs into solutions they actually adopt
- Work with the Technical Program Manager to run the development lifecycle: concept, design, test, release, and support
- Work closely with DevOps to operate and maintain the platform in our infrastructure
- Keep team knowledge written down: technical requirements, API contracts, deployment notes, and post-mortems
- Mentor other engineers and raise the bar on design and code quality through review
Requirements
Engineering
- 4+ years of software engineering experience in a team setting, building and running production web applications
- Strong JavaScript and TypeScript, with production Node.js experience
- Production experience with at least one modern frontend framework. Vue is preferred; React or Next.js also works, and we will expect you to become effective in Vue regardless of which you arrive with.
- Working knowledge of Go language
- PostgreSQL in production: schema design, indexing, transactions, and diagnosing a slow query rather than guessing at it
- REST API design, plus practical experience with streaming responses and long-lived connections
- Testing as part of the change rather than a later cleanup (TDD or close to it), and comfort with code review as a two-way conversation
- Good understanding of microservices design patterns and where they cost more than they return
Production and infrastructure
- Demonstrated ownership of scalability and reliability in high-traffic systems, including API gateways, load balancing, and operating against availability and error-rate targets
- Security fundamentals in day-to-day work: OAuth2 and role-based access control, credential handling, tenant isolation, input validation, and least privilege
- Docker, and container orchestration with Kubernetes and Helm
- CI/CD pipelines and Git-based workflows, including release and rollback
- A public cloud in production. Experience with Alibaba Cloud, AWS, GCP, or Azure.
AI application experience
- You have shipped LLM-backed features to real users, not only prototypes
- Familiarity with multiple model providers and their APIs (OpenAI, Anthropic, and others), including aggregators, and an understanding of where their contracts differ in practice rather than in documentation
- Practical grasp of streaming responses, tool and function calling, embeddings and retrieval (RAG with a vector database), multimodal input, and provider batch and file APIs
- Experience deploying and operating self-hosted models in production (LLMs, embedding, speech-to-text, text-to-speech, or multimodal) with an inference server such as vLLM, SGLang, TGI, or Triton, including the trade-offs between GPU capacity, latency, and throughput
- Experience building agent workflows that automate multi-step processes, and knowing where they need guardrails
- Some way of telling whether model output is actually good, whether that is evaluation sets, human review, or production signals
- Awareness of cost and latency as product constraints, not afterthoughts
- Working proficiency with AI-assisted development tools such as Claude Code or Codex
Collaboration and ways of working
- Experience collaborating directly with product teams and AI engineers on technical development of features and services
- Familiarity with Scrum and Kanban
- Strong written and verbal communication, with a habit of sharing context with teammates and stakeholders
- Ability to build and deploy solutions independently, from problem framing to production
Nice to have
- Experience deploying models across multiple GPUs or multiple nodes, or fine-tuning models for a specific use case
- Experience with microfrontend architectures (e.g. Module Federation, single-spa)
- Experience with Ruby frameworks (e.g. Rails, Sinatra)
- Experience building internal developer platforms or APIs consumed by other engineering teams
Originally posted on Himalayas