高级数据工程师
Senior Data Engineer
为什么选择Socure?
Socure正在为数字经济构建身份信任基础设施——实时验证100%的合法身份,从源头阻止欺诈。这个使命宏大,问题复杂,每天影响着企业、政府和数百万个人。
我们招聘希望承担这种责任的人。那些行动迅速、批判性思考、像主人一样行事,并且专注于精准解决客户问题的人。如果你想要可预测性或狭窄的职责范围,这里不适合你。如果你希望与一个对自己要求极高的团队一起,共同打造身份的未来——继续阅读。
职位简介
我们正在寻找一名高级数据工程师加入我们的数据自动化团队。你将在设计和构建可扩展的数据平台和管道方面发挥关键作用,这些平台和管道支持Socure的身份验证产品和分析。这个职位适合那些对用数据解决实际业务问题充满热情的人,同时具备深厚的数据工程实践经验和强烈的责任感。
你将负责的工作
• 设计和构建批处理和流式数据管道,以支持跨多个产品领域的自动化数据摄入、机器学习特征工程和分析。
• 负责复杂且模糊的数据项目从头到尾的交付,包括架构、实现、测试、部署、监控和文档编写。
• 开发和演进数据平台,使用现代云原生技术支持大规模数据处理。
• 自动化数据操作(验证、质量检查、警报、回填和恢复流程),以减少人工工作量并提高一致性。
• 优化数据工作负载的成本、性能和可靠性。
• 与跨职能团队(数据科学、产品、工程)紧密合作,理解需求,并将其转化为技术解决方案。
• 评估并采用新技术(新的处理引擎、存储格式、编排工具、GenAI辅助的数据摄入),以保持平台的现代化和高效。
你带来的能力
• 5年以上实际的数据工程经验,构建和维护生产级的数据平台和管道。
• 强大的通用编程语言(如Python或Scala)数据处理技能,以及SQL数据分析技能。
• 在分布式数据处理框架(如Apache Spark)方面有深入经验,包括性能调优和优化。
查看英文原文
Why Socure?
Socure is building the identity trust infrastructure for the digital economy — verifying 100% of good identities in real time and stopping fraud before it starts. The mission is big, the problems are complex, and the impact is felt by businesses, governments, and millions of people every day.
We hire people who want that level of responsibility. People who move fast, think critically, act like owners, and care deeply about solving customer problems with precision. If you want predictability or narrow scope, this won’t be your place. If you want to help build the future of identity with a team that holds a high bar for itself — keep reading.
About the Role
We are looking for a Senior Data Engineer to join our Data Automation team. You will play a critical role in designing and building scalable data platforms and pipelines that power Socure’s identity verification products and analytics. This role is ideal for someone who has a strong passion for solving real business problems with data, and combines deep hands-on data engineering expertise with strong ownership.
What You'll Do
• Design and build batch and streaming data pipelines to support automated data ingestion, ML feature engineering and analytics across multiple product domains.
• Own end-to-end delivery of complex, ambiguous data initiatives, including architecture, implementation, testing, deployment, monitoring, and documentation.
• Develop and evolve the data platform to support large-scale data processing using modern cloud-native technologies.
• Automate data operations (validation, quality checks, alerting, backfills, and recovery workflows) to reduce manual effort and improve consistency.
• Optimize cost, performance, and reliability of data workloads.
• Partner closely with cross-functional teams (Data Science, Product, Engineering) to understand requirements, translate them into technical solutions.
• Evaluate and adopt new technologies (new processing engines, storage formats, orchestration tools, GenAI-assisted ingestion) to keep the platform modern and efficient.
What You Bring
• 5+ years of hands-on data engineering experience, building and maintaining production-grade data platforms and pipelines.
• Strong programming skills in general-purpose language (such as Python or Scala) for data processing, and SQL for data analytics.
• Deep experience with distributed data processing frameworks, such as Apache Spark, including performance tuning and optimization.
• Proven experience building data solutions using services on AWS (EMR, Lambda, s3, etc).
• Strong understanding of data modeling and data warehousing concepts, including partitioning, schema design for large-scale datasets.
• Experience operating and supporting production pipelines, including monitoring, alerting, incident response, and improving reliability over time.
• Solid foundation in software engineering practices, including version control, CI/CD, testing strategies, and code review.
• Strong communication and collaboration skills, with the ability to work effectively with both technical and non-technical stakeholders.
Preferred Qualifications
• Experience with streaming or near-real-time data processing (Kafka, Kinesis, etc).
• Hands-on experience with data orchestration tools (Airflow, Step Functions, etc).
• Familiarity with modern data platform patterns such as Data Lakehouse, Data Mesh, and large-scale data sharing across teams.
• Experience with prompt engineering using modern GenAI, Large Language Models (LLM).
• Experience mentoring other engineers and contributing to engineering-wide standards, best practices.
As a note; Socure cannot provide sponsorship now or in the future for this role.
Socure is an equal opportunity employer that values diversity in all its forms within our company. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
If you need an accommodation during any stage of the application or hiring process—including interview or onboarding support—please reach out to your Socure recruiting partner directly.
Follow Us!
YouTube | LinkedIn | X (Twitter) | Facebook