Apache Spark 开发工程师
Apache Spark Developer
Apache Spark 开发人员 – 远程办公
Bright Vision Technologies 是一家技术咨询和软件开发公司,为美国各地提供云、人工智能、数据和企业解决方案。
这是一个加入一家知名且受人尊敬的组织的绝佳机会,提供巨大的职业发展潜能。
职位名称:Apache Spark 开发人员
工作地点:100% 远程(美国)
职位类型:全职,直接 W2
薪资范围:每年 125,000 至 185,000 美元
所需经验:6 年以上
赞助:美国公民、绿卡持有者、EAD 持有者以及 H-1B 转移候选人欢迎申请。我们无法为该职位的新 H-1B 签证申请提供赞助。
职位概述
我们正在寻找一位经验丰富的 Apache Spark 开发人员,负责设计、开发和优化支持企业分析、机器学习、实时报告和基于云的数据平台的大规模分布式数据处理应用程序。此职位专注于构建高性能的 Spark 应用程序,能够在结构化和半结构化数据源上处理数十亿条记录,同时提供可扩展、可靠且成本高效的数据显示管道。
您将与数据架构师、数据工程师、云平台团队、机器学习工程师和商业智能开发者紧密合作,利用 Apache Spark、云原生技术和分布式计算框架构建现代数据处理解决方案。理想的候选人具备对 Spark 架构、分布式系统、性能优化和基于云的大数据生态系统的深入理解。
主要职责
· 使用 Apache Spark 设计、开发和维护高性能的分布式数据处理应用程序。
- 构建可扩展的批处理和实时 ETL/ELT 管道,处理大量企业数据。
- 使用 PySpark、Scala 或 Spark SQL 开发 Spark 应用程序,用于数据转换、聚合和分析。
- 优化 Spark 作业的内存使用、分区策略、shuffle 性能和执行效率。
- 从企业数据库、API、Kafka、云存储和数据湖中处理结构化、半结构化和流式数据。
- 开发可重用的 Spark 库、数据处理框架和元数据驱动的摄入管道。
- 与云工程团队合作,将 Spark 工作负载部署到 Databricks、EMR、Azure Synapse 或 Kubernetes。
- 实现数据质量验证、对账和监控
查看英文原文
Apache Spark Developer – Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Apache Spark Developer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $125,000–$185,000 Annually
Experience Required: 6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary
We are seeking an experienced Apache Spark Developer to design, develop, and optimize large-scale distributed data processing applications supporting enterprise analytics, machine learning, real-time reporting, and cloud-based data platforms. This role focuses on building high-performance Spark applications capable of processing billions of records across structured and semi-structured data sources while delivering scalable, reliable, and cost-efficient data pipelines.
You will work closely with data architects, data engineers, cloud platform teams, machine learning engineers, and business intelligence developers to build modern data processing solutions leveraging Apache Spark, cloud-native technologies, and distributed computing frameworks. The ideal candidate possesses deep expertise in Spark architecture, distributed systems, performance optimization, and cloud-based big data ecosystems.
Key Responsibilities
· Design, develop, and maintain high-performance distributed data processing applications using Apache Spark.
- Build scalable batch and real-time ETL/ELT pipelines processing large volumes of enterprise data.
- Develop Spark applications using PySpark, Scala, or Spark SQL for data transformation, aggregation, and analytics.
- Optimize Spark jobs for memory utilization, partitioning strategies, shuffle performance, and execution efficiency.
- Process structured, semi-structured, and streaming data from enterprise databases, APIs, Kafka, cloud storage, and data lakes.
- Develop reusable Spark libraries, data processing frameworks, and metadata-driven ingestion pipelines.
- Collaborate with cloud engineering teams to deploy Spark workloads on Databricks, EMR, Azure Synapse, or Kubernetes.
- Implement data quality validation, reconciliation, monitoring, and automated error handling across distributed pipelines.
- Integrate Spark applications with enterprise data warehouses, lakehouses, and reporting platforms.
- Participate in architecture reviews, code reviews, technical design discussions, and Agile development activities.
- Troubleshoot production issues involving distributed processing, cluster performance, resource utilization, and data quality.
- Support cloud migration initiatives by modernizing legacy ETL workloads into Spark-based architectures.
Required Skills
· Six or more years of professional software or data engineering experience.
- Four or more years of hands-on Apache Spark development experience in enterprise production environments.
- Strong proficiency in PySpark, Scala, or Spark SQL for distributed data processing.
- Deep understanding of Apache Spark architecture including RDDs, DataFrames, Datasets, Catalyst Optimizer, DAG execution, and Tungsten engine.
- Strong experience with distributed computing concepts including partitioning, shuffling, caching, broadcast joins, and fault tolerance.
- Advanced SQL skills with databases such as SQL Server, Oracle, PostgreSQL, Snowflake, or Teradata.
- Experience working with Hadoop ecosystem technologies including Hive, HDFS, YARN, and Parquet.
- Experience processing streaming data using Spark Structured Streaming, Apache Kafka, or Event Hubs.
- Hands-on experience with cloud platforms including Azure Databricks, AWS EMR, AWS Glue, Azure Synapse Analytics, or Google Dataproc.
- Experience integrating Spark applications with Delta Lake, Apache Iceberg, or Apache Hudi.
- Strong understanding of data warehousing concepts, dimensional modeling, and data lake architecture.
- Experience using Git, CI/CD pipelines, Azure DevOps, GitHub Actions, or Jenkins.
- Strong debugging, troubleshooting, and Spark performance tuning skills.
- Experience working in Agile Scrum development environments.
Preferred Qualifications
· Experience building enterprise Lakehouse architectures using Databricks or Delta Lake.
- Familiarity with Apache Airflow, Azure Data Factory, AWS Step Functions, or Control-M for workflow orchestration.
- Experience with machine learning workflows using Spark MLlib, MLflow, or feature engineering pipelines.
- Knowledge of Kubernetes, Docker, and containerized Spark deployments.
- Experience implementing Data Quality frameworks using Great Expectations or Deequ.
- Familiarity with Apache NiFi, Apache Flink, Trino, or Presto.
- Experience working with cloud object storage including Amazon S3, Azure Data Lake Storage (ADLS Gen2), or Google Cloud Storage.
- Knowledge of Infrastructure as Code using Terraform or ARM templates.
- Experience with enterprise monitoring tools including Prometheus, Grafana, Datadog, or OpenTelemetry.
- Cloud certifications in Azure, AWS, Databricks, or Apache Spark-related technologies are highly desirable.
Project Environment
You will be joining a modern data engineering team responsible for building cloud-native big data platforms supporting enterprise analytics, AI, and business intelligence initiatives. Current projects include:
· Enterprise data lakehouse implementation using Databricks and Delta Lake
- Real-time streaming analytics processing billions of daily events
- Large-scale customer analytics and behavioral data platforms
- Financial risk modeling and fraud detection pipelines
- Healthcare clinical and operational analytics solutions
- Cloud migration of legacy Hadoop and ETL workloads
- Machine learning feature engineering and model training pipelines
- Enterprise reporting platforms supporting executive dashboards and self-service analytics
- Distributed data processing infrastructure deployed on Azure and AWS
This is a hands-on engineering role where you will contribute to distributed system architecture, Spark application development, cloud migration, performance optimization, production support, and continuous improvement of enterprise-scale data processing platforms.
How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to or contact us at (908) 676-4399. Learn more about Bright Vision Technologies at .
Bright Vision Technologies is an Equal Opportunity Employer.Equal Employment Opportunity (EEO) Statement
Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.
BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
Originally posted on Himalayas