高级 Python 数据抓取工程师(自由职业)
Senior Python Data Scraping Engineer (Freelance)
Mindrift 正在寻找高技能的 Vibecode 专家加入 Tendem 项目(),并为现实世界的应用场景驱动专业的数据抓取流程。Mindrift 正在寻找高技能的高级 Python 数据抓取工程师加入 Tendem 项目,驱动针对现实世界应用的专业数据抓取流程。在这个职位中,你将运用你在网页抓取、数据提取和数据处理方面的专业知识,交付准确、可靠且高质量的结果。这是一份兼职远程工作机会,适合在网页抓取、数据提取和处理方面有实际经验的技术专业人士。
我们做什么
Mindrift 平台将专家与创新技术项目连接起来。我们的使命是通过结合全球专业人员的实践经验与先进的 AI 开发工作,帮助开发高质量的 AI 技术。
职位介绍
这是一个 Tendem 项目的自由职业职位。作为高级 Python 数据抓取工程师,你将处理需要技术精确度的网页提取和处理任务,使用 Apify、OpenRouter 等工具,以及你自己的技术专长和方法。
主要职责
- 负责跨复杂网站的端到端数据提取流程,确保全面覆盖、准确性以及结构化数据集的可靠交付。
- 利用现有工具和自定义流程加速数据收集、验证和任务执行,同时满足既定要求。
- 确保从动态和交互式网页来源中可靠提取数据,根据需要调整方法以处理 JavaScript 渲染内容和变化的网站行为。
- 通过验证检查、跨源一致性控制、遵循格式规范以及交付前的系统验证来确保数据质量标准。
- 使用高效的批量处理或并行化扩展抓取操作,监控故障,并在小型网站结构变化时保持稳定性。
教育背景
- 至少 5+ 年的数据工程、网页抓取、自动化或软件开发相关经验(必需)。
- 工程、应用数学、计算机科学或相关技术领域的学士或硕士学位是加分项。
学术和/或专业经验
候选人应具备扎实的技术基础和脚本编写、自动化方面的实践经验
查看英文原文
Mindrift is looking for highly skilled Vibecode specialists to join the Tendem project () and drive specialized data scraping workflows for real-world use cases. Mindrift is looking for highly skilled Senior Python Data Scraping Engineers to join the Tendem project and drive specialized data scraping workflows for real-world applications. In this role, you'll apply your expertise in web scraping, data extraction, and data processing to deliver accurate, reliable, and high-quality results. This part-time remote opportunity is ideal for technical professionals with hands-on experience in web scraping, data extraction and processing.
What We Do
The Mindrift platform connects specialists with innovative technology projects. Our mission is to help develop high-quality AI technologies by combining real-world expertise from professionals across the globe with advanced AI development efforts.
About the Role
This is a freelance role for a Tendem project. As a Senior Python Data Scraping Engineer, you'll handle data scraping tasks requiring technical precision for web extraction and processing, utilizing tools such as Apify, OpenRouter, and other technologies, alongside your own technical expertise and approaches.
Key Responsibilities
- Own end-to-end data extraction workflows across complex websites, ensuring complete coverage, accuracy, and reliable delivery of structured datasets.
- Leverage available tools and custom workflows to accelerate data collection, validation, and task execution while meeting defined requirement.
- Ensure reliable extraction from dynamic and interactive web sources, adapting approaches as needed to handle JavaScript-rendered content and changing site behavior.
- Enforce data quality standards through validation checks, cross-source consistency controls, adherence to formatting specifications, and systematic verification prior to delivery.
- Scale scraping operations for large datasets using efficient batching or parallelization, monitor failures, and maintain stability against minor site structure changes.
Educational qualifications
- At least 5+ years of relevant experience in data engineering, web scraping, automation, or software development (required).
- Bachelor’s or Master’s Degree in Engineering, Applied Mathematics, Computer Science, or related technical fields is a plus.
Academic and/or Professional Experience
Candidates should have a strong technical foundation and practical experience with scripting, automation, and data extraction workflows. We are looking for specialists who can solve non-trivial problems, work confidently with modern development tools and technologies, and systematically collect, structure, and validate data from diverse sources. A methodical, detail-oriented approach and the ability to work independently are essential.
Technical Skills (Essential)
- Strong experience in Python web scraping (BeautifulSoup, Selenium or similar), including dynamic content (JS, AJAX, infinite scroll) and APIs via proxies
- Proven ability to extract data from complex structures (hierarchies, archived pages, inconsistent HTML)
- Solid background in data cleaning, normalization, and validation, delivering structured datasets (CSV, JSON, Google Sheets)
Additional requirements
- Demonstrated experience handling anti-bot mechanisms and dynamic site structures at scale
- Experience with cloud infrastructure (AWS or equivalent) and containerization (Docker) as part of real workflows
- Hands-on experience with LLM frameworks (LangChain, OpenRouter, or similar) applied to automation tasks
- Strong attention to detail and commitment to data accuracy
- Self-directed work ethic with ability to troubleshoot independently
- A link to GitHub is a plus
- English proficiency: Upper-intermediate (B2) or above (required)
Project time expectations
For this project, tasks are estimated to require around 10–20 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active.
Compensation
On this project, contributors can earn up to $25 per hour equivalent, depending on their level and pace of contribution. Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.
Originally posted on Himalayas