[GFA] Azure 高级数据工程师
[GFA] Azure Senior Data Engineer
Software Mind 为全球的公司提供具有影响力的技术解决方案。科技巨头与独角兽企业、变革性项目、新兴技术以及无限的机会——这些是我们的日常写照。我们组建跨职能的工程团队,他们拥有主人翁意识并追求卓越,因此我们一直在寻找那些为每个项目带来激情和创造力的优秀人才。我们的文化倡导开放、尊重、坚韧与勇气,并将工作与乐趣结合在一起。
项目 – 你将承担的职责
我们的客户提供了创新的解决方案和洞察力,使我们的客户能够管理风险并雇佣最优秀的人才。他们的先进全球技术平台支持可完全扩展、可配置的筛选程序,满足全球33,000多家客户的独特需求。总部位于美国佐治亚州亚特兰大,他们在19个国家拥有国际化的员工队伍,约有5,500名员工。我们的合作伙伴每年在200多个国家和地区进行超过9300万次筛查。
我们正在寻找一位高级数据工程师,要求具备 Databricks PySpark 开发经验,全面的数据建模经验(包括事实表/维度表、SCD 类型 2 和增量加载),以及基于事件的架构技能,加入我们的数据工程团队,推动基于 Azure 的数据分析平台的发展。
职位 – 你将做出的贡献
· 使用 Databricks Lakehouse 架构和 PySpark 开发可重用、元数据驱动的数据管道
· 设计和实现全面的数据模型
· 构建具有高级功能的 ETL/ELT 解决方案:合并操作、SCD 类型 2 实现等
· 实现具有幂等性模式和重叠连接优化的增量数据加载
· 自动化并优化数据平台流程,重点关注性能和可靠性
· 使用事件驱动模式构建与数据源和消费者的集成
· 与基础设施工程团队合作设置云资源
· 发起并实施数据平台架构的改进
要求 – 你需要具备的经验
· Databricks 专业知识:熟练掌握 Databricks Lakehouse 架构和 PySpark 开发
· 数据建模精通:在数据模型定义方面有丰富经验,包括事实表、维度表、度量、粒度分析、SCD 类型 2、代理键、迟到记录处理、重叠连接以及具有幂等性的增量加载
查看英文原文
Software Mind develops solutions that make an impact for companies around the globe. Tech giants & unicorns, transformative projects, emerging technologies and limitless opportunities – these are a few words that describe an average day for us. Building cross-functional engineering teams that take ownership and crave more means we’re always on the lookout for talented people who bring passion and creativity to every project. Our culture embraces openness, acts with respect, shows grit & guts and combines employment with enjoyment.
Project – the aim you’ll have
Our customer provides innovative solutions and insights that enable our clients to manage risk and hire the best talent. Their advanced global technology platform supports fully scalable, configurable screening programs that meet the unique needs of over 33,000 clients worldwide. Headquartered in Atlanta, GA, they have an internationally distributed workforce spanning 19 countries with about 5,500 employees. Our partner perform over 93 million screens annually in over 200 countries and territories.
We are seeking a Senior Data Engineer with proven expertise in Databricks PySpark development, comprehensive data modeling experience (including fact/dimension tables, SCD Type 2 and incremental loads) and event-based architecture skills to join our Data Engineering Team and drive the evolution of our Azure-based Data Analytics Platform.
Position – how you’ll contribute
· Develop reusable, metadata-driven data pipelines using Databricks Lakehouse architecture and PySpark
· Design and implement comprehensive data models
· Build robust ETL/ELT solutions with advanced features: Merge operations, SCD Type 2 implementations, etc.
· Implement incremental data loads with idempotency patterns and overlap joins optimization
· Automate and optimize data platform processes with focus on performance and reliability
· Build integrations with data sources and consumers using event-driven patterns
· Cooperate with infrastructure engineering team to set up cloud resources
· Initiate and implement improvements to data platform architecture
Expectations – the experience you need
· Databricks expertise: proficient in Databricks Lakehouse architecture and PySpark development
· Data modeling mastery: extensive experience in data model definition including fact tables, dimension tables, measures, grain analysis, SCD Type 2, surrogate keys, late-arriving records handling, overlap joins, and incremental loads with idempotency
· Programming: advanced Python and PySpark skills for ETL/ELT development
· Databricks optimization: deep knowledge of optimize, zOrder, Liquid clustering, ACID transactions and performance tuning
· Event-based architecture: proven experience in designing and implementing event-driven data solutions
· Azure data platform: experience working with Azure-based datasets and data pipelines
· SQL proficiency: strong SQL skills for complex data transformations
· Large-scale data processing: experienced in handling large and complex datasets efficiently
· CI/CD: experience in developing automated deployment pipelines
· Networking fundamentals: understanding of basic networking concepts
· Agile methodology: familiar with Scrum and agile development practices
Additional skills – the edge you have
· Understanding of stream processing challenges and Spark Structured Streaming
· Experience with Infrastructure as Code (Terraform, Bicep)
· Experience with containerized applications (Azure Container Apps, Kubernetes)
· Knowledge of Azure cloud native solutions (Azure Data Factory, Azure Function App, Azure Container Instances)
Our offer – professional development, personal growth:
· Flexible employment and remote work
· International projects with leading global clients
· International business trips
· Non-corporate atmosphere
· Language classes
· Internal & external training
· Private healthcare and insurance
· Multisport card
· Well-being initiatives
Position at: Software Mind Poland
This role requires candidates to be based in Poland.