数据工程师
Data Engineer
XTB是一家来自金融行业的全球公司,专注于金融工具的在线交易。我们是波兰最大的FinTech公司,在中欧和东欧处于领先地位,业务范围涵盖多个国家,包括亚洲和南美洲。在XTB,我们注重员工的发展,为他们提供在各个领域获取知识和技能的机会,以及提供多种培训和发展计划。如果您正在寻找挑战,并希望在国际商业环境中获得宝贵的经验,XTB将是您的理想选择。
我们是一家经过认证的Great Place to Work公司。
我们正在寻找一名数据工程师加入核心数据平台团队,共同创建和开发供公司各团队使用的集中式数据平台。在此职位中,您将负责构建可扩展的数据集成机制,从各种源系统中提取数据,开发共享平台组件和标准,并确保数据的高质量、可靠性和一致性。
职责
- 日常实施、维护和支持MS SQL Server仓库和SSIS流程及作业。
- 使用SQL、Python和Apache Spark / PySpark设计、构建和维护可扩展的数据处理管道。
- 处理原始数据并设计将其集成到数据平台的方法。
- 开发数据工程解决方案的CI/CD管道。
- 创建和维护基础设施即代码解决方案。
- 将数据平台与其他系统和应用程序集成。
- 参与架构决策并定义工程标准。
- 进行代码审查,编写文档,并与团队成员分享知识。
要求
- 至少3年的数据工程师或相关软件工程岗位经验。
- 至少1年使用MS SQL Server、SSIS和SQL Server Agent技术的经验——包括新功能的部署。
- 分析现有遗留代码以进行改进并迁移到新技术。
- 在合作过程中有意愿参与基于Databricks的项目。
- 精通Python并能够编写干净、可测试的生产代码。
- 有Apache Spark的实际操作经验。
- 高级SQL技能(大型数据集的查询优化)。
- 对数据仓库设计原则和数据建模有深入理解。
- 实际知识
查看英文原文
XTB is a global company from the financial industry, focusing on online trading of financial instruments. We are the largest FinTech in Poland and a leader in Central and Eastern Europe, and the range of our operations covers several countries, including Asia and South America. At XTB, we focus on the development of our employees, giving them opportunities to gain knowledge and skills in various fields, as well as offering a number of training and development programs. If you are looking for challenges and want to gain valuable experience in an international business environment, XTB is the right place for you.
We are a certified Great Place to Work company.
We are looking for a Data Engineer to join the Core Data Platform team to co-create and develop a centralized data platform used by teams across the company. In this role, you will be responsible for building scalable data integration mechanisms from various source systems, developing shared platform components and standards, and ensuring high data quality, reliability, and consistency.
Responsibilities
- Implementing, maintaining and support of MS SQL Server warehouse and SSIS processes and jobs on a daily basis.
- Designing, building, and maintaining scalable data processing pipelines using SQL, Python, and Apache Spark / PySpark.
- Working with raw data and designing methods for its integration into the data platform.
- Developing CI/CD pipelines for data engineering solutions.
- Creating and maintaining Infrastructure as Code solutions.
- Integrating the data platform with other systems and applications.
- Participating in architectural decision-making and defining engineering standards.
- Conducting code reviews, documenting solutions, and sharing knowledge with team members.
Requirements
- At least 3 years of experience as a Data Engineer or in a related software engineering role.
- At least 1 year of experience with MS SQL Server, SSIS and SQL Server Agent technologies - including the deployment of new features.
- Analyzing existing legacy code for improvement and transfer to the newer technology.
- Flexibility to be enrolled into the Databricks oriented project at some point of the cooperation.
- Proficiency in Python and the ability to write clean, testable production code.
- Hands-on experience with Apache Spark.
- Advanced SQL skills (query optimization on large datasets).
- Strong understanding of data warehouse design principles and data modeling.
- Practical knowledge of ETL/ELT processes.
- Familiarity with data quality, monitoring, and pipeline reliability.
- Experience working with relational databases, including an understanding of how they work, data modeling, and optimization.
- Understanding of REST and gRPC standards for system integration.
- Experience with workflow orchestration tools (e.g., Airflow, Dagster, or similar).
- Knowledge of CI/CD pipeline building and maintenance principles.
- Strong analytical problem-solving skills and attention to detail.
- Effective communication skills and the ability to collaborate in a team.
- Openness to learning and exploring new technologies and methodologies.
Nice to have
- Experience with dbt.
- Experience with Kafka, Pub/Sub, or other event streaming systems.
- Experience building near-real-time / real-time pipelines.
- Experience working with Infrastructure-as-Code tools.
- Understanding of CDC (Change Data Capture) and data integration from transactional systems.
- Experience with Databricks or other lakehouse platforms.
What we offer
- Real influence on the development of the company and the product.
- Work in an experienced team that is happy to share its knowledge.
- A clear vision of development thanks to regular feedback and clear career paths.
- Regular team-building meetings.
Benefits
- A training budget for courses and conferences that interest you.
- An extra day off on your birthday.
- An extra day off for parents.
- Equipment tailored to your needs.
- Private medical care and group insurance.
- Access to an e-learning platform for learning English and a benefits platform.
- Access to a wellbeing platform and the opportunity to take advantage of workshops and private therapy sessions.
- Remote work, from the office in Warsaw or from a coworking space in your city.
Originally posted on Himalayas