#130529 - 软件/数据工程师 - Spark、AWS EMR 与 AI
#130529 - Software/Data Engineer - Spark, AWS EMR & AI
我们正在寻找一位有经验的软件/数据工程师,负责设计和交付可扩展的数据处理系统和AI驱动的工作流程。该合同职位位于软件工程和数据工程的交叉领域,重点在于Spark、基于云的分布式处理、生产可靠性以及分析和机器学习用例的数据准备
主要职责
· 使用Spark和Amazon EMR或类似的基于云的数据处理平台设计和开发可扩展的数据处理解决方案。
· 构建和维护批处理和分布式数据管道。
· 开发用于数据转换、特征准备和AI或机器学习工作流集成的软件组件。
· 与工程、AI和产品团队合作,实现数据驱动和模型驱动的用例。
· 优化数据管道性能、成本效率、可扩展性和生产可靠性。
· 在开发和生产环境中排查数据和应用问题。
· 参与架构讨论、技术文档和工程标准制定。
· 确保解决方案符合数据质量、治理和安全要求。
必备技能
· 4年以上软件工程或数据工程经验。
· 精通Spark和分布式数据处理。
· 具备Amazon EMR或类似基于云的数据处理平台经验。
· 熟练掌握Java、Python或相关编程语言。
· 了解AI或机器学习工作流、模型集成或智能系统数据准备。
· 熟悉可扩展数据架构和性能优化。
· 具备良好的调试和协作能力。
· 能够在不断变化、数据密集型环境中交付成果。
· 能够兼顾软件工程和数据工程职责。
· 强烈的执行力和实际的架构判断力。
加分项
· 具备Kafka、Airflow、数据湖或数据仓库生态系统经验。
· 熟悉MLOps、特征存储或AI平台集成。
· 具备AWS原生服务和可观测性工具经验。
· 有企业经验者优先。
所需工具与平台
· Apache Spark
· Amazon EMR或类似的基于云的分布式数据处理平台
· Java、Python或相关编程语言
地点、时间与参与方式
· 远程合同职位
· 候选人必须位于拉美地区,不包括墨西哥
查看英文原文
We are seeking an experienced Software/Data Engineer to design and deliver scalable data processing systems and AI-enabled workflows. This contract role sits at the intersection of software engineering and data engineering, with a strong focus on Spark, cloud-based distributed processing, production reliability, and data preparation for analytics and machine learning use cases
Key Responsibilities
· Design and develop scalable data processing solutions using Spark and Amazon EMR or comparable cloud-based data processing platforms.
· Build and maintain batch and distributed data pipelines.
· Develop software components for data transformation, feature preparation, and AI or machine learning workflow integration.
· Collaborate with engineering, AI, and product teams to operationalize data-driven and model-enabled use cases.
· Optimize data pipeline performance, cost efficiency, scalability, and production reliability.
· Troubleshoot data and application issues across development and production environments.
· Contribute to architecture discussions, technical documentation, and engineering standards.
· Ensure solutions align with data quality, governance, and security expectations.
Must-Have Skills
· 4+ years of software engineering or data engineering experience.
· Strong experience with Spark and distributed data processing.
· Experience with Amazon EMR or similar cloud-based data processing platforms.
· Proficiency in Java, Python, or a related programming language.
· Exposure to AI or machine learning workflows, model integration, or data preparation for intelligent systems.
· Strong understanding of scalable data architecture and performance optimization.
· Strong debugging and collaboration skills.
· Comfortable delivering in evolving, data-intensive environments.
· Ability to bridge software engineering and data engineering responsibilities.
· Strong execution focus with practical architecture judgment.
Nice-to-Have Skills
· Experience with Kafka, Airflow, data lakes, or data warehouse ecosystems.
· Familiarity with MLOps, feature stores, or AI platform integration.
· Experience with AWS-native services and observability tooling.
· Enterprise experience strongly preferred.
Required Tools & Platforms
· Apache Spark.
· Amazon EMR or a comparable cloud-based distributed data processing platform.
· Java, Python, or a related programming language.
Location, Time & Engagement
· Remote contract role.
· Candidates must be located in LATAM, excluding Mexico.
· U.S. Central Time coverage is required.
· Full-time allocation of approximately 40 hours per week.
· Current contract end date is March 31, 2027.