高级软件工程师,数据基础设施
Senior Software Engineer, Data Infrastructure
ZoomInfo 是职业加速的地方。我们行动迅速,敢于创新,并赋能你完成一生中最出色的工作。你将与一群充满热情、互相挑战并庆祝胜利的同事一起工作。拥有能放大你影响力的工具和一个支持你抱负的文化,你不仅仅是在做出贡献,而是在快速实现目标。
**职位机会**
我们正在寻找一名高级软件工程师加入我们的网络数据团队,帮助构建 ZoomInfo 下一代的网络爬取和数据提取基础设施。
这是一个具有高影响力的工程岗位,你将与经验丰富的工程师一起开发处理数十亿页面的大规模爬取和提取系统。你的日常工作将涉及解决实际的分布式系统问题,编写生产级代码,并推出直接影响平台数据质量的功能。
你将与一位负责人力资源和行政事务的专职经理合作,一位将业务需求与技术工作联系起来的产品经理,以及一位消除障碍并支持职业发展的高级经理合作。你的重点是工程执行、技术卓越和协作。
**你将负责**
作为高级软件工程师,你将为处理海量网络数据的企业级爬取和提取平台做出贡献。
你将:
- 设计和实现可扩展、容错的网络爬取和提取流水线组件
- 使用 Java 和 Python 编写干净、生产级别的代码
- 构建和运营大规模数据提取和转换的 ETL/ELT 流水线
- 在 GCP 和 AWS 上使用云基础设施,主要在 GKE 上
- 提升你所参与系统的可观测性、可靠性和运营卓越性
- 与产品和数据科学团队合作,交付有影响力解决方案
- 参与代码审查、文档编写和团队内的知识分享
- 跟踪不断发展的网络技术、反爬机制和人工智能驱动的提取方法
**必备资格**
我们更看重扎实的软件工程基础,而非深入的爬虫相关经验。合适的候选人首先是一名优秀的工程师;该领域知识可以在团队中学习。
软件工程基础
- 5 年以上构建生产系统的专业软件工程经验
- 扎实的计算机科学基础:算法、数据结构、并发、分布式系统
查看英文原文
ZoomInfo is where careers accelerate. We move fast, think boldly, and empower you to do the best work of your life. You’ll be surrounded by teammates who care deeply, challenge each other, and celebrate wins. With tools that amplify your impact and a culture that backs your ambition, you won’t just contribute. You’ll make things happen–fast.
**The Opportunity**
We're looking for a Senior Software Engineer to join our Web Data team and help build the next generation of ZoomInfo's web crawling and data extraction infrastructure.
This is a hands-on engineering role with high impact. You'll work alongside experienced engineers building large-scale crawling and extraction systems that process billions of pages. Your day-to-day will involve solving real distributed systems problems, writing production code, and shipping features that directly impact data quality across the platform.
You'll partner with a dedicated people manager who handles HR and administrative responsibilities, a product manager who connects business needs with technical work, and a senior manager who removes roadblocks and supports career growth. Your focus is on engineering execution, technical excellence, and collaboration.
**What You'll Do**
As a Senior Software Engineer, you'll contribute to enterprise-scale crawling and extraction platforms that process massive volumes of web data.
You will:
- Design and implement components of scalable, fault-tolerant web crawling and extraction pipelines
- Write clean, production-grade code in Java and Python
- Build and operate ETL/ELT pipelines for large-scale data extraction and transformation
- Work with cloud infrastructure on GCP and AWS, primarily on GKE
- Improve observability, reliability, and operational excellence across the systems you contribute to
- Partner with product and data science teams to deliver impactful solutions
- Contribute to code reviews, documentation, and knowledge sharing across the team
- Stay current with evolving web technologies, anti-crawling mechanisms, and AI-powered extraction approaches
**Must-Have Qualifications**
We're prioritizing strong software engineering fundamentals over deep crawling-specific experience. The right candidate is a great engineer first; the domain can be learned on the team.
Software Engineering Fundamentals
- 5+ years of professional software engineering experience building production systems
- Strong CS fundamentals: algorithms, data structures, concurrency, distributed systems
- Proficiency in Java and/or Python
- Track record of owning features end-to-end from design through deployment and operation
- Comfortable making sound architectural decisions at the component level
**Data Engineering**
- Hands-on experience with cloud data warehouses such as BigQuery or Snowflake
- Experience designing and operating large-scale ETL/ELT pipelines
- Experience with orchestration tools such as Apache Airflow
- Experience with streaming or event-driven systems such as Apache Kafka
**Cloud and Infrastructure**
- Production experience on GCP (preferred) or AWS; multi-cloud exposure is a plus
- Hands-on experience with Kubernetes (GKE/EKS) for distributed workloads
- Familiarity with infrastructure-as-code tooling such as Terraform
**Background and Mindset**
- Strong communicator who can explain technical decisions clearly
- Comfortable operating in ambiguity and iterating quickly
- Bias toward action and pragmatic problem solving
- Self-starter who thrives in fast-paced, evolving environments
**Nice to Have**
- Experience with web crawling at scale (Scrapy or similar frameworks)
- Familiarity with proxy infrastructure, rotation strategies, or anti-bot evasion techniques
- Experience extracting structured and unstructured web data from diverse site architectures
- Knowledge of SERP (Search Engine Results Page) extraction
- Comfort with AI/LLM-based extraction approaches, applying language models to HTML at scale
- Experience working in a B2B data company or data-as-a-product environment
**Core Technical Stack**
- Java and Python
- Apache Kafka
- GCP (BigQuery, GKE, Vertex AI)
- Snowflake and Starburst/Trino
- Terraform
- Scrapy and web scraping frameworks
- Proxy management systems
- Distributed systems and Kubernetes
- Apache Airflow
- Large-scale ETL pipelines
**Ideal Candidate Profile**
- 5+ years of software engineering experience
- Experience operating systems that process meaningful volumes of data
- Strong CS fundamentals (algorithms, data structures, distributed systems)
- Excellent communicator who can explain complex ideas to diverse audiences
- Passion for solving hard problems and building elegant, scalable systems
- Self-starter who thrives in fast-paced, evolving environments
- Experience working in a B2B data company or data-as-a-product environment is a strong plus
**Why This Role Matters**
This role sits at the heart of ZoomInfo's data platform. You'll contribute to the infrastructure that powers our business intelligence products and help shape how web data acquisition is done at ZoomInfo, while working at massive scale with cutting-edge technologies.
You'll be joining at a pivotal moment: the team is growing, the ambition is high, and there's real opportunity to make a lasting impact and grow into deeper technical leadership over time.
LI-REMOTE
LI-AR2
Actual compensation offered will be based on factors such as the candidate’s work location, qualifications, skills, experience and/or training. Your recruiter can share more information about the specific salary range for your desired work location during the hiring process. We want our employees and their families to thrive.
In addition to comprehensive benefits we offer holistic mind, body and lifestyle programs designed for overall well-being. Learn more about ZoomInfo benefits [here](https://www.zoominfo.com/careers#benefits).
Below is the US base salary for this position. Additional compensation such as Bonus, Commission, Equity and other benefits may also apply.
$140,000—$220,000 USD
**About us:**
ZoomInfo (NASDAQ: GTM) is the Go-To-Market Intelligence Platform that empowers businesses to grow faster with AI-ready insights, trusted data, and advanced automation. Its solutions provide more than 35,000 companies worldwide with a complete view of their customers, making every seller their best seller.
ZoomInfo is committed to protecting your privacy when you apply for jobs with us. Please review our Job Applicant [Privacy Notice](https://drive.google.com/file/d/1yiG-k0YX_sW10PiJk_xliDc3W-veJCrK/view?usp=drive_link) for more details on how we handle your personal information.
ZoomInfo may use a software-based assessment as part of the recruitment process. More information about this tool, including the results of the most recent bias audit, is available [here](https://www.zoominfo.com/legal/nyc-local-law-144-notice).
ZoomInfo is proud to be an equal opportunity employer, hiring based on qualifications, merit, and business needs, and does not discriminate based on protected status. We welcome all applicants and are committed to providing equal employment opportunities regardless of sex, race, age, color, national origin, sexual orientation, gender identity, marital status, disability status, religion, protected military or veteran status, medical condition, or any other characteristic protected by applicable law. We also consider qualified candidates with criminal histories in accordance with legal requirements.
For Massachusetts Applicants: It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability. ZoomInfo does not administer lie detector tests to applicants in any location.