软件工程总监 (Node.js & 网络爬虫专家)
Director of Software Engineering (Node.js & Web Scraping Expert)
我们正在寻找一位软件工程总监,具备深厚的 Node.js 开发和大规模网络爬虫经验。该职位将领导工程团队,设计和优化高性能、分布式网络爬虫系统。理想的候选人应具备处理反爬虫措施、数据管道优化和可扩展云架构的丰富经验。
关键职责 - 软件工程与网络爬虫领导:
- 使用 Node.js 架构、开发并维护可扩展的分布式网络爬虫系统。
- 设计并实现数据抽取管道,以处理大量结构化和非结构化数据。
- 开发解决方案以绕过反爬虫机制,包括 CAPTCHA 处理、会话管理、指纹识别和 IP 旋转。
- 优化爬虫流程的性能、可靠性和效率,同时管理代理服务(住宅、数据中心、旋转)。监督数据存储和处理策略,确保高可用性和一致性。
- 与产品、DevOps 和数据科学团队合作,将提取的数据集成到分析和业务应用中。
- 实施微服务、API 集成和实时数据流的最佳实践。
关键职责 - 可扩展性、安全性和 DevOps:
- 领导向云原生、容器化和无服务器架构的迁移,用于网络爬虫。
- 确保符合法律和道德标准(robots.txt、GDPR、CCPA 等)。优化云资源(AWS、GCP 或 Azure)以支持高吞吐量的爬虫。
- 管理实时监控和警报系统,以检测爬虫失败、IP 封禁或性能瓶颈。
- 与 DevOps 团队紧密合作,优化 CI/CD 流水线、自动化部署和系统可扩展性。
关键职责 - 工程团队管理与战略:
- 领导、指导并发展一个高绩效的工程团队。
- 制定并执行技术路线图,与业务目标保持一致。
- 培养持续学习、协作和创新的文化。
- 实施敏捷开发方法(Scrum、Kanban)以优化项目执行。
- 确保所有工程工作中代码质量、安全性和最佳实践。
任职资格与经验 - 技术专长:
- 10 年以上软件工程经验,其中至少 5 年在网络爬虫和大规模数据提取领域。
- 在 Node.js、Puppeteer、Playwright 等工具上有扎实的实战经验。
查看英文原文
We are seeking a Director of Software Engineering with deep expertise in Node.js development and large-scale web scraping. This role will lead the engineering team, designing and optimizing high-performance, distributed web scraping systems. The ideal candidate has extensive experience in handling anti-bot measures, data pipeline optimization, and scalable cloud-based architectures.
Key Responsibilities- Software Engineering & Web Scraping Leadership:
- Architect, develop, and maintain scalable and distributed web scraping systems using Node.js.
- Design and implement data extraction pipelines to process large volumes of structured and unstructured data.
- Develop solutions to bypass anti-bot mechanisms, including CAPTCHA handling, session management, fingerprinting, and IP rotation.
- Optimize scraping processes for performance, reliability, and efficiency while managing proxy services(residential, datacenter, rotating).Oversee data storage and processing strategies, ensuring high availability and consistency.
- Collaborate with Product, DevOps, and Data Science teams to integrate extracted data into analytics and business applications.
- Implement best practices for microservices, API integrations, and real-time data streaming.
Key Responsibilities- Scalability, Security & DevOps:
- Lead the transition to cloud-native, containerized, and serverless architectures for web scraping.
- Ensure compliance with legal and ethical standards (robots.txt, GDPR, CCPA, etc.).Optimize cloud resources (AWS, GCP, or Azure) to support high-throughput scraping.
- Manage real-time monitoring and alerting systems to detect scraping failures, IP bans, or performance bottlenecks.
- Work closely with DevOps teams to optimize CI/CD pipelines, automated deployments, and system scalability.
Key Repsonsibilities- Engineering Team Management & Strategy:
- Lead, mentor, and grow a high-performance engineering team.
- Define and execute the technology roadmap, aligning with business objectives.
- Foster a culture of continuous learning, collaboration, and innovation.
- Implement agile development methodologies (Scrum, Kanban) to optimize project execution.
- Ensure code quality, security, and best practices across all engineering efforts.
Qualifications & Experience- Technical Expertise:
- 10+ years of experience in software engineering, with at least 5+ years in web scraping and large-scale data extraction.
- Strong hands-on expertise in Node.js, Puppeteer, Playwright, Cheerio, Selenium, and headless browser automation.
- Extensive experience in handling CAPTCHAs, IP rotation, session management, and anti-bot evasion techniques.
- Deep knowledge of proxy management (residential, datacenter, rotating, and VPNs).Experience with NoSQL/SQL databases (MongoDB, PostgreSQL, Redis, Elasticsearch, etc.).
- Familiarity with data processing frameworks (Kafka, RabbitMQ, Spark, Airflow, etc.).Strong experience with CI/CD, containerization (Docker, Kubernetes), and cloud deployment (AWS/GCP/Azure).
Qualifications & Experience- Leadership & Soft Skills:
- Proven track record of scaling engineering teams and leading complex projects.
- Strong problem-solving and debugging skills, especially for scraping challenges and performance bottlenecks.
- Excellent communication and stakeholder management skills.
- Passion for mentorship, team development, and continuous learning.
Preferred Qualifications:
- Experience with machine learning for data extraction and NLP.
- Knowledge of browser fingerprinting and bot detection mechanisms.
- Familiarity with enterprise-scale web crawling frameworks (Scrapy, Colly, Apify, etc.).
- Prior leadership experience in data-driven businesses or web scraping startups.
Originally posted on Himalayas