数据工程师
Data Engineer
职位名称:数据工程师
职位类型:全职
工作地点:远程
职位描述
我们正在寻找一名数据工程师,以构建和扩展支持人工智能驱动产品和研究计划的数据基础设施。在该职位中,您将开发分布式数据流水线,管理跨云环境的大规模数据集,并设计可靠的数据系统,以支持大规模的数据处理、实验和模型开发。
主要职责
- 设计、构建和维护可扩展的数据流水线,用于从多个来源摄取、处理和转换大规模数据集。
- 使用 Spark 和云原生技术开发和优化分布式数据处理工作流。
- 在 SQL 和 NoSQL 系统中构建和维护数据存储解决方案,确保可扩展性、性能和可靠性。
- 在 AWS 上设计和实现数据架构,以支持高吞吐量的数据摄取、处理和分发。
- 编写高效的 Python 和 SQL 代码,用于提取、转换、验证和分析大规模数据集。
- 确保数据管道和存储层中的数据质量、完整性、监控和操作可靠性。
- 与人工智能研究人员、数据科学家和工程团队合作,支持数据密集型应用和实验。
- 实现自动化、编排和监控工作流,以支持可扩展且高效的数据操作。
所需技能和资格
- 精通 Python、SQL 以及分布式数据处理框架,如 Apache Spark。
- 具有 AWS 数据服务和云原生数据架构的实际经验。
- 具有使用 SQL 和 NoSQL 数据库的经验。
- 具有在分布式环境中管理和处理大规模数据集的经验。
- 对数据分区、性能优化和可扩展数据架构有深入理解。
加分项
- 接触过 AI/ML 工作流或研究环境。
- 具有数据可视化工具(如 Matplotlib、Seaborn 或 Plotly)的经验。
- 熟悉与 LLM 相关的数据工作流(用于训练、评估或提示实验的数据集)。
薪酬与福利说明
此全职职位的全国薪资范围为 10 万美元至 15 万美元美元。所有员工均有资格获得股权补偿,员工还可能根据职位和公司政策获得基于绩效的奖金。micro1 提供全面的福利包,包括最高 100%
查看英文原文
Job Title: Data Engineer
Job Type: Full-time
Location: Remote
The Role
We are looking for a Data Engineer to build and scale the data infrastructure that powers AI-driven products and research initiatives. In this role, you will develop distributed data pipelines, manage large-scale datasets across cloud environments, and design reliable data systems that support data processing, experimentation, and model development at scale.
Key Responsibilities
- Design, build, and maintain scalable data pipelines to ingest, process, and transform large-scale datasets from multiple sources.
- Develop and optimize distributed data processing workflows using Spark and cloud-native technologies.
- Build and maintain data storage solutions across SQL and NoSQL systems, ensuring scalability, performance, and reliability.
- Design and implement data architectures on AWS to support high-volume data ingestion, processing, and distribution.
- Write efficient Python and SQL code to extract, transform, validate, and analyze large datasets.
- Ensure data quality, integrity, monitoring, and operational reliability across data pipelines and storage layers.
- Collaborate with AI researchers, data scientists, and engineering teams to support data-intensive applications and experimentation.
- Implement automation, orchestration, and monitoring workflows to support scalable and efficient data operations.
Required Skills and Qualifications
- Strong proficiency in Python, SQL, and distributed data processing frameworks such as Apache Spark.
- Hands-on experience with AWS data services and cloud-native data architectures.
- Experience working with both SQL and NoSQL databases.
- Experience managing and processing large-scale datasets in distributed environments.
- Strong understanding of data partitioning, performance optimization, and scalable data architectures
Nice to Have
- Exposure to AI/ML workflows or research environments.
- Experience with data visualization tools such as Matplotlib, Seaborn, or Plotly.
- Familiarity with LLM-related data workflows (datasets for training, evaluation, or prompt experimentation).
Compensation & Benefits Notice
The national pay range for this full-time position is base salary of $100,000 –$150,000 USD. All employees are eligible for equity compensation, and employees may also receive performance-based bonuses, dependent on role and subject to company policies. micro1 provides a comprehensive benefits package, including up to 100% reimbursement for health-insurance premiums, paid time off, a 401(K) plan with a company match, and additional benefits designed to support a high-performing, remote-first workforce.
micro1 is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, sexual orientation, or gender identity), national origin, age, disability, genetic information, veteran status, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance and/or a reasonable accommodation during the application process, reach out to .
Our hiring process utilizes artificial intelligence tools to assist in candidate screening and assessment. Our AI tools are designed to complement, not replace, human decision-making.
Disclaimer
The information contained in this job posting, including but not limited to role responsibilities, qualifications, compensation, and benefits, is provided for informational purposes only and does not constitute a binding offer of employment. micro1 reserves the right to amend, modify, or withdraw any portion of this posting at its sole discretion and without prior notice. All employment decisions are made in accordance with applicable laws and regulations.
Originally posted on Himalayas