资深软件工程师- 数据采集
Staff Software Engineer- Data Ingestion
作为资深软件工程师,你将主导我们后端架构的演进,主要专注于重构和优化现有的数据流水线。你将推动下一代分布式数据存储和处理系统的发展,这些系统旨在无限扩展并超越传统查询性能。在现代化的基础上,你将设计简洁、直观的接口,为各种数据消费者(从核心网络应用到高级业务分析和人工智能)抽象复杂性。你的专业知识将在将我们的基础设施转变为强大、高性能的基础方面发挥关键作用。
主要职责:
- 识别并开发可扩展且高性能的解决方案。
- 跨领域协作,塑造产品策略和执行。
- 建立代码架构和质量的基础。
- 指导和培养工程师。
- 设定并维护工程流程标准,以支持高质量的工程实践。
最低要求:
- 计算机科学、工程或相关领域的学士/学士以上学位。
- 8年以上作为工程师构建高度可扩展系统的生产级经验。
- 4年以上在团队环境中担任可信技术决策者的经验,兼顾短期和长期业务价值。
- 4年以上使用SQL或其他数据库查询语言处理大型多表数据集的经验。
- 具有大规模分布式系统架构、开发和部署经验。
- 熟悉云技术,例如AWS、Azure、GCP。
- 具有构建持续集成和持续开发(CI/CD)流水线的经验。
- 熟悉服务器端网络技术(例如:Java、Python、Scala、C#、C++、Go)。
优先考虑的技能和经验:
- 8年以上构建高度可扩展和可靠基础设施的经验。
- 在设计、优化和编排强大的数据流水线(ETL/ELT)和摄取系统方面有专长,适用于大规模、实时和批处理场景。
- 管理数据仓库(例如Snowflake、Redshift)并利用分析工具(例如Spark、SQL、Python、Databricks)的经验。
- 具有容器化(Docker、Kubernetes)、CI/CD流水线和分布式架构(事件驱动、内存计算)的实际操作经验。
- 对现代数据库系统有深入熟练的掌握,包括复制、分片、分区、索引和缓存策略,用于高性能查询优化。
- 深入了解数据架构和数据工程原则,能够设计和实现高效的数据处理和存储方案。
查看英文原文
As a Staff Software Engineer, you will lead the evolution of our backend architecture, with a primary focus on refactoring and optimizing existing data pipelines. You will drive the development of next-generation distributed data storage and processing systems designed to scale indefinitely and surpass traditional query performance. Beyond modernization, you will design clean, expressive interfaces that abstract complexity for a wide range of data consumers—from core web applications to advanced business analytics and AI. Your expertise will be instrumental in transforming our infrastructure into a robust, high-performance foundation.
Primary Duties:
- Identify and develop scalable and performant solutions.
- Work across discipline to shape product strategy and execution.
- Develop the foundations of code architecture and quality.
- Mentor and coach engineers.
- Set and uphold the standard for engineering processes to support high-quality engineering.
Minimum Qualifications:
- BS/BTech (or higher) in Computer Science, Engineering or a related field required.
- 8+ years of production-level experience as an engineer building highly scalable systems.
- 4+ years of experience acting as a trusted technical decision-maker in a team setting, solving for short-term and long-term business value.
- 4+ years of experience working with SQL or other database querying languages on large multi-table data sets.
- Experience architecting, developing, and deploying large-scale distributed systems at scale.
- Experience with cloud technologies, e.g., AWS, Azure, GCP.
- Experience building continuous integration and continuous development (CI/CD) pipelines.
- Strong familiarity with server-side web technologies (eg: Java, Python, Scala, C#, C++, Go).
Preferred KSAs:
- 8+ years experience building highly scalable and reliable infrastructure.
- Expertise in designing, optimizing, and orchestrating robust data pipelines (ETL/ELT) and ingestion systems for large-scale, real-time, and batch processing.
- Experience managing data warehouses (e.g., Snowflake, Redshift) and leveraging analytics tools (e.g., Spark, SQL, Python, Databricks).
- Hands-on experience with containerization (Docker, Kubernetes), CI/CD pipelines, and distributed architectures (event-driven, in-memory computing).
- Deep proficiency with modern database systems, including replication, sharding, partitioning, indexing, and caching strategies for high-performance query optimization.
- Strong understanding of data security, governance, and compliance principles.
- Experience with infrastructure monitoring, performance optimization, and active participation in architecture reviews.
Physical Requirements:
Sitting for prolonged periods of time. Extensive use of computers and keyboard. Occasional walking and lifting may be required.