中级数据工程师
Mid-Level Data Engineer
我们正在寻找一名数据工程师,负责构建、改进和维护支持我们产品和组织的数据管道和基础设施。您将使用SQL、Python、Spark/PySpark、Databricks、AWS和Infrastructure as Code来开发可靠、可扩展的数据解决方案,并确保生产数据环境的性能和质量。
您将与工程师、数据科学家以及Crisis Text Line的其他团队合作,将数据需求转化为深思熟虑、可维护的解决方案。我们希望找到具备扎实数据工程基础并有实际经验构建和维护生产数据管道的人才。您不需要熟悉我们环境中每种技术——我们重视可迁移的经验、好奇心以及协作学习和解决问题的能力。
我们的技术栈
您将在一个现代的云基础环境中工作,包括:
- 语言与处理:SQL、Python、Spark/PySpark
- 数据平台:Databricks
- 云服务:AWS
- 基础设施即代码:Terraform
- 可观测性:Datadog
- 数据存储:MySQL、PostgreSQL 和 Redis
- 交付与可靠性:自动化CI/CD、监控以及对生产数据管道的共同责任
职责:
- 作为我们更广泛的数据和工程环境的一部分,设计、构建、维护和改进可靠的生产数据管道。
- 使用SQL、Python、Spark/PySpark和Databricks开发和优化数据处理流程。
- 与工程师、数据科学家和业务团队合作,理解数据需求并将其转化为实用、可维护的解决方案。
- 构建和维护数据源与我们的数据基础设施之间的集成。
- 使用基础设施即代码实践来支持和改进我们的AWS数据环境。
- 监控管道健康状况、数据质量和性能;排查问题并在出现问题时贡献解决方案。
- 自动化重复性流程,以提高数据管道的可靠性、效率和可维护性。
- 参与共享生产支持,并培养独立诊断和解决管道问题的信心。
- 参与代码审查、技术文档和工程标准的制定,帮助团队构建可靠且可维护的系统。
- 带来想法,提出问题,并随着时间推移识别改进数据工具、流程和工程实践的机会。
所需资格:
- 大约三年
查看英文原文
We’re looking for a Data Engineer to build, improve, and support the data pipelines and infrastructure that power our products and organization. You’ll work with SQL, Python, Spark/PySpark, Databricks, AWS, and Infrastructure as Code to develop reliable, scalable data solutions and help ensure the performance and quality of our production data environment.
You’ll collaborate with engineers, data scientists, and teams across Crisis Text Line to translate data needs into thoughtful, maintainable solutions. We’re looking for someone with strong data engineering fundamentals and hands-on experience building and supporting production data pipelines. You don’t need experience with every technology in our environment—we value transferable experience, curiosity, and the ability to learn and solve problems collaboratively.
Our Stack
You’ll work across a modern, cloud-based environment, including:
- Languages & Processing: SQL, Python, Spark/PySpark
- Data Platform: Databricks
- Cloud: AWS
- Infrastructure as Code: Terraform
- Observability: Datadog
- Data Stores: MySQL, PostgreSQL, and Redis
- Delivery & Reliability: Automated CI/CD, monitoring, and shared ownership of production data pipelines
Responsibilities:
- Design, build, maintain, and improve reliable production data pipelines as part of our broader data and engineering environment.
- Develop and optimize data processing workflows using SQL, Python, Spark/PySpark, and Databricks.
- Partner with engineers, data scientists, and business teams to understand data needs and translate them into practical, maintainable solutions.
- Build and maintain integrations between data sources and our data infrastructure.
- Use Infrastructure as Code practices to support and improve our AWS data environment.
- Monitor pipeline health, data quality, and performance; troubleshoot issues and contribute to solutions when something isn’t working as expected.
- Automate repeatable processes that improve the reliability, efficiency, and maintainability of our data pipelines.
- Participate in shared production support and develop confidence diagnosing and resolving pipeline issues independently.
- Contribute to code reviews, technical documentation, and engineering standards that help the team build reliable and maintainable systems.
- Bring ideas, ask questions, and identify opportunities to improve our data tools, processes, and engineering practices over time.
Required Qualifications:
- Approximately three to five years of professional data engineering, software engineering, or related experience, including meaningful responsibility for production data pipelines.
- Professional experience using SQL and Python to build, transform, and troubleshoot data.
- Experience building, maintaining, and supporting production data pipelines or ETL/ELT workflows.
- Experience with distributed data processing technologies such as Spark or PySpark.
- Familiarity with cloud-based environments such as AWS and modern data engineering practices.
- Experience testing, debugging, monitoring, and improving the reliability and performance of production data systems.
- The ability to independently complete well-defined engineering work, investigate problems, and clearly communicate technical decisions, risks, and tradeoffs.
- A collaborative approach to working with engineers, data scientists, and cross-functional partners, with a commitment to building secure and reliable data systems.
Preferred Qualifications
The following experience would be useful, but we encourage you to apply even if you do not meet every item:
- Databricks or a comparable cloud-based data platform.
- AWS services and cloud-based data architecture.
- Terraform or another infrastructure-as-code tool.
- Lakehouse modeling such as medallion architecture and data warehousing.
- Datadog or comparable monitoring and observability tools.
- Data quality, automated testing, CI/CD, or deployment automation.
- Designing data systems for reliability, performance, scalability, and security.
- Supporting data used for analytics, reporting, or machine learning workloads.
- Claude or other AI-assistance code tools.
- Working in a mission-driven, regulated, safety-sensitive, or high-trust environment.
Reliable High-Speed Internet Required: Must have a stable high-speed internet connection to support seamless remote collaboration, virtual meetings, online job tasks, etc.
For United States-based candidates:
The target salary range for this position, across the United States, is $99,704 - $116,990. Starting salary will vary based on location, qualifications, and prior experience. Candidates will learn the range specific to their location during the interview process. We pay competitively in the tech-forward nonprofit space and offer a robust benefits package.
This is a remote position within the United States. At the time of hire and throughout employment, the employee’s primary residence and regular home work location must be in one of the following approved hiring states: California, Colorado, Connecticut, Florida, Georgia, Illinois, Indiana, Maryland, Massachusetts, Michigan, New Jersey, New Mexico, New York, North Carolina, Pennsylvania, Tennessee, Texas, Utah, Virginia, or Washington state.
No visa sponsorship available for this position.
Benefits & Well-Being
Crisis Text Line recognizes that we are all unique human beings with unique life circumstances, and our benefits package aims to be as flexible as possible to support your needs as you work to promote mental well-being for people, wherever they are. Our benefits package is thoughtfully designed using an equity lens, with input from our team and from industry best practices.
Highlights include:
- Comprehensive medical, dental, and vision options that prioritize accessibility and financial peace of mind
- Employer-funded HSA contributions
- Generous PTO, sick time, and 19 paid holidays with a winter break
- 12 weeks of fully paid parental leave after 26 consecutive weeks of service
- Monthly internet and mental health stipends
- Annual Wellness Stipend
- Home office and professional development stipends
- 403(b) retirement plan with employer contribution
- Sabbatical after 3 years of service
Benefits are for U.S.-based employees; international benefits may vary.
This is a remote-only position
Originally posted on Himalayas