网络爬虫工程师 - 全球团队(印度)
Web Scraping Engineer - Global Team (India)
我们分析每天数十亿个替代数据点,为拼车、电子商务市场、支付等提供准确、详细的洞察。我们的按需洞察团队使用专有技术来识别、授权、清理和分析许多世界顶级投资基金和公司所依赖的数据。
连续三年,我们被列为Inc最佳工作场所之一。我们是一家快速增长的技术公司,由The Carlyle Group和Norwest Venture Partners支持。我们的办公室位于纽约、奥斯汀、迈阿密、丹佛、山景城、西雅图、香港、上海、北京、广州和新加坡。我们培养以员工为中心的文化,注重精通、所有权和透明度。
为什么现在申请:
- 高影响力:你的工作将直接影响多个业务部门的关键报告和战略决策。
- 充满挑战:解决弹性网络爬虫的设计问题,应对动态网站结构,并优化大规模数据提取。
- 成长机会:作为我们不断扩展的网络爬虫工程团队的早期成员,你将在我们的策略、流程和团队文化方面有重要发言权。
职位简介:
我们正在寻找一名网络爬虫工程师加入我们不断壮大的工程团队。在这个实践性很强的职位中,你将负责设计、构建和维护强大的网络爬虫,这些爬虫为我们的组织中的关键报告和客户体验提供支持。你将处理复杂且高影响力的爬取挑战,并与跨职能团队紧密合作,确保我们的数据摄入流程具有弹性、高效且可扩展,同时向我们的产品和利益相关者提供高质量的数据。
作为我们的网络爬虫工程师,你将:
重构和维护网络爬虫
- 重新设计现有的爬虫脚本,提高可靠性、可维护性和效率。
- 实施最佳编码实践(干净代码、模块化架构、代码审查等),以确保质量和可持续性。
实施高级爬取技术
- 使用复杂的指纹方法(cookies、headers、user-agent轮换、代理)以避免被检测和阻止。
- 处理动态内容,导航复杂的DOM结构,并有效管理会话/cookie生命周期。
查看英文原文
About YipitData:
YipitData is the leading market research and analytics firm for the disruptive economy and recently raised up to $475M from The Carlyle Group at a valuation over $1B.
We analyze billions of alternative data points every day to provide accurate, detailed insights on ridesharing, e-commerce marketplaces, payments, and more. Our on-demand insights team uses proprietary technology to identify, license, clean, and analyze the data many of the world’s largest investment funds and corporations depend on.
For three years and counting, we have been recognized as one of Inc’s Best Workplaces. We are a fast-growing technology company backed by The Carlyle Group and Norwest Venture Partners. Our offices are located in NYC, Austin, Miami, Denver, Mountain View, Seattle, Hong Kong, Shanghai, Beijing, Guangzhou, and Singapore. We cultivate a people-centric culture focused on mastery, ownership, and transparency.
Why You Should Apply NOW:
- High Impact: Your work will directly influence key reports and strategic decisions across multiple business units.
- Exciting Challenges: Tackle the design of resilient web scrapers, navigate dynamic website structures, and optimize large-scale data extraction.
- Growth Opportunities: As an early member of our expanding Web Scraping Engineering team, you will have significant input on our strategies, processes, and team culture.
About The Role:
We are seeking a Web Scraping Engineer to join our growing engineering team. In this hands-on role, you’ll take ownership of designing, building, and maintaining robust web scrapers that power critical reports and customer experiences across our organization. You will work on complex, high-impact scraping challenges and collaborate closely with cross-functional teams to ensure our data ingestion processes are resilient, efficient, and scalable, while delivering high-quality data to our products and stakeholders.
As Our Web Scraping Engineer You Will:
Refactor and Maintain Web Scrapers
- Overhaul existing scraping scripts to improve reliability, maintainability, and efficiency.
- Implement best coding practices (clean code, modular architecture, code reviews, etc.) to ensure quality and sustainability.
Implement Advanced Scraping Techniques
- Utilize sophisticated fingerprinting methods (cookies, headers, user-agent rotation, proxies) to avoid detection and blocking.
- Handle dynamic content, navigate complex DOM structures, and manage session/cookie lifecycles effectively.
Collaborate with Cross-Functional Teams
- Work closely with analysts and other stakeholders to gather requirements, align on targets, and ensure data quality.
- Provide support, documentation, and best practices to internal stakeholders to ensure effective use of our web scraped data in critical reporting workflows.
Monitor and Troubleshoot
- Develop robust monitoring solutions, alerting frameworks to quickly identify and address failures.
- Continuously evaluate scraper performance, proactively diagnosing bottlenecks and scaling issues.
Drive Continuous Improvement
- Propose new tooling, methodologies, and technologies to enhance our scraping capabilities and processes.
- Stay up to date with industry trends, evolving bot-detection tactics, and novel approaches to web data extraction.
This is a fully-remote opportunity based in India. Standard work hours are from 11am to 8pm IST, but there is flexibility here.
You Are Likely To Succeed If:
- Effective communication in English with both technical and non-technical stakeholders.
- 3+ years of experience with web scraping frameworks (e.g., Selenium, Playwright, or Puppeteer).
- Strong understanding of HTTP, RESTful APIs, HTML parsing, browser rendering, and TLS/SSL mechanics.
- Expertise in advanced fingerprinting and evasion strategies (e.g., browser fingerprint spoofing, request signature manipulation).
- Deep experience managing cookies, headers, session states, and proxy rotations, including the deployment of both residential and data center proxies.
- Experience with logging, metrics, and alerting to ensure high availability.
- Troubleshooting skills to optimize scraper performance for efficiency, reliability, and scalability.
What We Offer:
Our compensation package includes comprehensive benefits, perks, and a competitive salary:
- We care about your personal life and we mean it. We offer vacation time, parental leave, team events, learning reimbursement, and more!
- Your growth at YipitData is determined by the impact that you are making, not by tenure, unnecessary facetime, or office politics. Everyone at YipitData is empowered to learn, self-improve, and master their skills in an environment focused on ownership, respect, and trust.
We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, marital status, disability, gender, gender identity or expression, or veteran status. We are proud to be an equal-opportunity employer.
Job Applicant Privacy Notice
<img height="1" width="1" style="display:none;" alt="" src="https://px.ads.linkedin.com/collect/?pid=4341228&conversionId=10486642&fmt=gif" />