生成式AI数据工程师
Gen AI Data Engineer
Tiger Analytics 正在寻找具有生成式 AI 经验的资深机器学习工程师,加入我们快速发展的高级分析咨询公司。我们的员工在机器学习、数据科学和人工智能领域拥有深厚的专业知识。我们是多家财富 500 强公司的可信赖分析合作伙伴,帮助他们从数据中创造商业价值。我们的业务价值和领导力已获得包括 Forrester 和 Gartner 在内的多家市场研究机构的认可。
随着我们继续打造全球最佳的分析咨询团队,我们正在寻找顶尖人才。您将负责以下工作:
所需技术技能:
编程语言:熟练掌握 Python、SQL 和 PySpark。
数据仓库:具备 Snowflake、NOSQL 和 Neo4j 的经验。
数据管道:熟练使用 Apache Airflow。
云平台:熟悉 AWS(S3、RDS、Lambda、AWS batch、SageMaker processing Job、CloudFormation 等)或 GCP(Vertex AI RAG、Data pipeline、Bigquery、GKE)。
操作系统:具备 Linux 经验。
批处理/实时管道:具备构建和部署各种管道的经验。
版本控制:具备 GitHub 经验。
开发工具:熟练使用 VS Code。
工程实践:具备测试、部署自动化、DevOps/SysOps 的技能。
沟通能力:具备出色的演示和沟通能力。
协作能力:具备与本地/海外团队合作的经验。
要求
期望技能:
- 大数据技术:具备 Hadoop 和 Spark 的经验。
- 数据可视化:熟练使用 Streamlit 和仪表盘。
- API:具备构建和维护内部 API 的经验。
- 机器学习:了解 ML 基本概念。
- 生成式 AI:熟悉生成式 AI 工具和方法。
附加专长:
- 知识图谱:具备创建和检索的经验。
- 向量数据库:熟练管理向量数据库。
- 数据持久化:能够开发和维护多种数据持久化和检索方法(RDMBS、向量数据库、存储桶、图数据库、知识图谱等)。
- 云技术:具备 AWS 经验,特别是 SageMaker、Lambda、OpenSearch。
- 自动化工具:具备 Airflow DAGs、AutoSys 和 CronJobs 的经验。
- 非结构化数据管理:具备管理非结构化数据(音频、视频、图像、文本等)的经验。
- CI/CD:具备使用 Jenkins 和 GitHub Actions 进行持续集成和部署的专业知识。
- 基础设施即代码:具备相关经验。
查看英文原文
Tiger Analytics is looking for experienced Machine Learning Engineers with Gen AI experience to join our fast-growing advanced analytics consulting firm. Our employees bring deep expertise in Machine Learning, Data Science, and AI. We are the trusted analytics partner for multiple Fortune 500 companies, enabling them to generate business value from data. Our business value and leadership has been recognized by various market research firms, including Forrester and Gartner.
We are looking for top-notch talent as we continue to build the best global analytics consulting team in the world. You will be responsible for:
Technical Skills Required:
Programming Languages: Proficiency in Python, SQL, and PySpark.
Data Warehousing: Experience with Snowflake, NOSQL and Neo4j.
Data Pipelines: Proficiency with Apache Airflow.
Cloud Platforms: Familiarity with AWS (S3, RDS, Lambda, AWS batch, SageMaker processing Job, CloudFormation, etc.) or GCP (Vertex AI RAG, Data pipeline, Bigquery, GKE)
Operating Systems: Experience with Linux.
Batch/Realtime Pipelines: Experience in building and deploying various pipelines.
Version Control: Experience with GitHub.
Development Tools: Proficiency with VS Code.
Engineering Practices: Skills in testing, deployment automation, DevOps/SysOps.
Communication: Strong presentation and communication skills.
Collaboration: Experience working with onshore/offshore teams.
Requirements
Desired Skills:
- Big Data Technologies: Experience with Hadoop and Spark.
Data Visualization: Proficiency with Streamlit and dashboards.
- APIs: Experience in building and maintaining internal APIs.
- Machine Learning: Basic understanding of ML concepts.
- Generative AI: Familiarity with generative AI tools and techniques.
Additional Expertise:
- Knowledge Graphs: Experience with creation and retrieval.
- Vector Databases: Proficiency in managing vector databases.
- Data Persistence: Ability to develop and maintain multiple forms of data persistence and retrieval methods (RDMBS, Vector Databases, buckets, graph databases, knowledge graphs, etc.).
- Cloud Technologies: Experience with AWS, especially SageMaker, Lambda, OpenSearch.
- Automation Tools: Experience with Airflow DAGs, AutoSys, and CronJobs.
- Unstructured Data Management: Experience in managing data in unstructured forms (audio, video, image, text, etc.).
- CI/CD: Expertise in continuous integration and deployment using Jenkins and GitHub Actions.
- Infrastructure as Code: Advanced skills in Terraform and CloudFormation.
- Containerization: Knowledge of Docker and Kubernetes.
- Monitoring and Optimization: Proven ability to monitor system performance, reliability, and security, and optimize them as needed.
- Security Best Practices: In-depth understanding of security best practices in cloud environments.
- Scalability: Experience in designing and managing scalable infrastructure.
- Disaster Recovery: Knowledge of disaster recovery and business continuity planning.
- Problem-Solving: Excellent analytical and problem-solving abilities.
- Adaptability: Ability to stay up-to-date with the latest industry trends and adapt to new technologies and methodologies.
- Team Collaboration: Proven ability to work well in a team environment and contribute to a positive, collaborative culture.
GenAI Engineer Specific Skills:
- Industry Experience: 8+ years of experience in data engineering, platform engineering, or related fields, with deep expertise in designing and building distributed data systems and large-scale data warehouses.
- Data Platforms: Proven track record of architecting data platforms capable of processing petabytes of data and supporting real-time and batch ingestion processes.
- Data Pipelines: Strong experience in building robust data pipelines for document ingestion, indexing, and retrieval to support scalable RAG solutions. Proficiency in information retrieval systems and vector search technologies (e.g., FAISS, Pinecone, Elasticsearch, Milvus).
- Graph Algorithms: Experience with graphs/graph algorithms, LLMs, optimization algorithms, relational databases, and diverse data formats.
- Data Infrastructure: Proficient in infrastructure and architecture for optimal extraction, transformation, and loading of data from various data sources.
- Data Curation: Hands-on experience in curating and collecting data from a variety of traditional and non-traditional sources.
- Ontologies: Experience in building ontologies in the knowledge retrieval space, schema-level constructs (including higher-level classes, punning, property inheritance), and Open Cypher.
- Integration: Experience in integrating external databases, APIs, and knowledge graphs into RAG systems to improve contextualization and response generation.
- Experimentation: Conduct experiments to evaluate the effectiveness of RAG workflows, analyze results, and iterate to achieve optimal performance.
Benefits
This position offers an excellent opportunity for significant career development in a fast-growing and challenging entrepreneurial environment with a high degree of individual responsibility.
Originally posted on Himalayas