高级软件工程师(后端与基础设施)
Senior Software Engineer (Backend & Infrastructure)
职位描述
你将负责为Forager的平台客户以大规模方式提供数据的后端、数据和基础设施——从数据摄入到客户集成的API和数据流。这是一个以后端为主的职位:你大部分时间将用于构建API功能,运营我们的搜索和ETL数据基础设施,并负责其下的DevOps和平台层。你也会对面向客户的React/TypeScript网页应用做出贡献,但这只是工作的一小部分。
这是在小型团队中的高级、高自主权职位。你不仅会贡献功能——你还将负责平台本身的设计、可靠性和演进。你将直接参与产品方向的决策,并拥有(以及责任)发布能显著提升我们覆盖范围、准确率、延迟和系统可用性指标的工作的自由。
你将负责的内容
按你花费时间的优先级排序:
- 后端与API开发 —— 实时增强API(人员/组织查找、联系人数据、反向搜索)、搜索和筛选端点、瀑布增强逻辑,以及客户计量的计费界面。匹配率、延迟和数据新鲜度直接决定客户价值。
- 数据工程、ETL与搜索基础设施 —— 全流程负责Elasticsearch/ECK堆栈(索引和映射设计、数据摄入管道、相关性调优、集群健康、扩展、迁移),以及每天处理数十亿数据点的ETL应用程序,将数据存入可搜索存储、Snowflake数据产品和仓库导出,同时保持数据填充率和准确性。
- DevOps与平台(核心负责人) —— 你负责日常AWS基础设施、基础设施即代码、CI/CD和可观测性。这不是辅助性的运维工作:你负责基础设施现代化、部署流水线、安全扫描、日志/指标整合,以及你所构建系统的可靠性。
- 可观测性与数据质量工具 —— 构建Grafana仪表盘、负载测试和数据准确性评估,使覆盖范围、数据新鲜度、延迟和正确性在数据质量范围内可测量并持续监控。
- 面向客户的网页应用(次要部分) —— 当工作需要时,为React/TypeScript应用、文档、上手指南和自助服务界面做出贡献。
核心职责
后端与API开发
- 设计、构建和维护用于增强、搜索、数据流和平台客户流程的RESTful API。
- 开发可扩展的后端服务——工作者、任务队列、水槽
查看英文原文
The Role
You'll own the backend, data, and infrastructure that deliver Forager's data to platform customers at scale — from ingestion all the way through to the APIs and feeds customers integrate against. This is a backend-first role: most of your time is spent building API features, operating our search and ETL data infrastructure, and owning the DevOps and platform layer underneath it. You'll also contribute to our customer-facing React/TypeScript web app, but that's the minority of the work.
This is a senior, high-ownership role on a small team. You won't just contribute features — you'll be responsible for the design, reliability, and evolution of the platform itself. You'll have direct input into product direction and the freedom (and responsibility) to ship work that materially moves our coverage, accuracy, latency, and uptime metrics.
What You'll Own
Ranked by where your time goes:
- Backend & API development — real-time enrichment APIs (person/org lookup, contact data, reverse search), search and filter endpoints, waterfall enrichment logic, and the credit/billing surfaces customers meter against. Match rate, latency, and freshness directly drive customer value.
- Data engineering, ETL & search infrastructure — own the Elasticsearch/ECK stack end-to-end (index and mapping design, ingestion pipelines, relevance tuning, cluster health, scaling, migrations), plus the ETL applications that move billions of data points daily into searchable stores, Snowflake data products, and warehouse exports while holding fill rates and accuracy.
- DevOps & platform (core owner) — you own day-to-day AWS infrastructure, infrastructure-as-code, CI/CD, and observability. This is not shared-on-the-side ops: you're responsible for infrastructure modernization, deployment pipelines, security scanning, log/metric consolidation, and the reliability of the systems you build.
- Observability & data-quality tooling — build Grafana dashboards, load tests, and data-accuracy evals that make coverage, freshness, latency, and correctness measurable and continuously monitored across the data-quality surface.
- Customer-facing web app (minority) — contribute to the React/TypeScript app, docs, onboarding, and self-serve surfaces when the work calls for it.
Core Responsibilities
Backend & API Development
- Design, build, and operate RESTful APIs for enrichment, search, feeds, and platform-customer workflows.
- Develop scalable backend services — workers, task queues, waterfall lookups, data pipelines — that keep refresh cycles predictable and fill rates high.
- Build the metering, credit-logging, and billing logic that customers rely on for accurate consumption.
- Participate actively in product planning; help shape which features have the highest customer impact.
Data Engineering, Search & ETL
- Own the Elasticsearch / ECK search stack end-to-end: index design, ingestion, relevance, cluster health, scaling, and migrations.
- Design and operate ETL applications moving data into searchable stores, bulk feeds, and warehouses (Snowflake, S3).
- Optimize PostgreSQL — query performance, indexing, cache utilization.
- Drive measurable improvements in latency, uptime, error rate, fill rate, and scalability.
DevOps & Infrastructure (Core Owner)
- Own AWS infrastructure (ECS, S3, etc.) and infrastructure-as-code.
- Own CI/CD pipelines.
- Own observability — Grafana, CloudWatch, Sentry — and on-call response for the surfaces you build.
- Drive infrastructure modernization and share crawler infrastructure maintenance with the team.
Collaboration & Quality
- Code review with high standards for readability, security, and performance.
- Write unit, integration, and E2E tests — test reliability is a quality contributor, not overhead.
- Document features, architecture, and API contracts; great developer docs are how our customers succeed.
Requirements
What We're Looking For
Required Experience
- 5+ years building and operating production backend systems and APIs, with clear examples of strategic technical problem-solving (not just tenure).
- Strong proficiency in Python / Django REST Framework.
- Hands-on experience operating Elasticsearch / OpenSearch at scale — index/mapping design, query and relevance tuning, cluster management, migrations.
- Demonstrated track record building and operating ETL pipelines that move significant data volumes reliably.
- Production experience with PostgreSQL, Redis, and async task systems (Celery / RabbitMQ or equivalent).
- Core-owner-level comfort with AWS (ECS, S3, CloudWatch), infrastructure-as-code, and CI/CD (GitHub Actions or equivalent) — you're comfortable owning the platform, not just deploying to it.
- Experience operating services in production — observability, on-call, incident response.
- Working proficiency in React / TypeScript — comfortable shipping product UI when the work calls for it. (This is the minority of the role, not its center.)
- Strong written communication; comfortable owning documentation as a deliverable.
AI & Agentic Workflows (Required)
This is non-negotiable. AI coding tools are core to how we build, and you're expected to operate effectively in an AI-augmented engineering environment. You must demonstrate strong, hands-on fluency with:
- AI coding tools (Claude Code, Cursor, Copilot, or equivalent) used daily for implementation, refactoring, and code review — while validating all generated output before it reaches production.
- Agentic workflows — designing, orchestrating, and debugging multi-step agent pipelines (e.g., research → plan → implement → verify loops, MCP server integration, tool-use design).
- Judgment about where AI helps vs. hurts — knowing when to delegate to an agent, when to write the code yourself, and how to keep an agent on rails for production work.
We evaluate this in interviews with live exercises. Candidates without demonstrable agentic workflow experience will not be considered.
Nice to Have
- Experience with Snowflake or other data warehouses.
- Background in B2B data products — enrichment, contact data, company data, search/discovery.
- Experience with data-privacy/compliance engineering — GDPR/PII handling, US state privacy regimes.
- Building MCP servers.
- Experience with web crawling, data sourcing, or large-scale ingestion systems.
- Open-source contributions or public technical writing.
Benefits
- Competitive pay for your location — calibrated to your experience and where you're based.
- Flexible working hours — we care about output and overlap, not clock-watching.
- Untracked PTO — take the time off you need; there's no accrual to count.
- Relaxed working environment — small team, high autonomy, low bureaucracy. Anyone can take an idea to production without jumping through hoops.
Originally posted on Himalayas