大数据负责人
Big Data Lead
开发工程限定地区(需当地身份)
公司DemandMatrix
薪资未公开
工作地点India
地域资格限定地区(需当地身份)
时区要求日间重叠约 6 小时,基本正常作息
用工类型Full Time
发布时间今天
数据来源Himalayas
注意地域限制:该职位明确限定在 India 招聘。如果你是位于中国大陆的求职者,通常需要当地工作身份才能投递,或需与雇主确认是否接受独立合同(Contractor)形式合作。
你将做什么?
- 为了帮助我们更进一步,我们正在寻找一位能够实际操作的专家,利用大数据技术解决最复杂的数据问题。你将花费近一半的时间进行实际编码。
- 这包括大规模文本数据处理、事件驱动的数据流水线、内存计算,以及考虑CPU核心到网络IO再到磁盘IO的优化。
- 你将使用AWS和GCP上的云原生服务。
你是什么样的人?
- 扎实的计算机工程基础、Unix系统、数据结构和算法知识将使你能够应对这一挑战。
- 设计并构建了多个大数据模块和数据流水线,以处理大量数据。
- 对技术充满热情,并从零开始参与过项目。
必须具备:
- 7年以上软件开发经验,专注于大数据和大规模数据流水线。
- 至少3年使用Python构建服务和流水线的经验。
- 精通多种数据处理系统,包括流处理、事件处理和批处理(Spark、Hadoop/MapReduce)。
- 了解至少一种NoSQL存储,如MongoDB、Elasticsearch、HBase。
- 了解大规模高吞吐和高可用环境中分布式数据存储的数据模型、分片和数据定位策略,以及它们对非结构化文本数据处理的影响。
- 具有在AWS或GCP上运行可扩展且高可用系统的经验。
加分项:
- 具有Docker/Kubernetes经验
- 有CI/CD经验
- 了解爬虫/采集技术
- 全程远程办公
- 生日假期
- 远程工作
在DemandMatrix,我们的愿景是通过领域知识、机器学习和AI来颠覆价值1000亿美元的销售和营销情报行业。微软、谷歌、Adobe、亚马逊、IBM等财富100强公司都信任我们,以识别他们的下一个客户。
最初发布于喜马拉雅山
查看英文原文
What will you do?
- To help us go to the next level we are looking to onboard a hands-on SME in leveraging big data tech to solve the most complex data issues. You will spend almost half of time with hands-on coding.
- It involves large scale text data processing, event driven data pipelines, in-memory computations, optimization considering CPU core to network IO to disk IO.
- You will be using cloud native services in AWS and GCP.
Who Are You?
- Solid grounding in computer engineering, Unix, data structures and algorithms would enable you to meet this challenge.
- Designed and built multiple big data modules and data pipelines to process large volume.
- Genuinely excited about technology and worked on projects from scratch.
Must have:
- 7+ years of hands-on experience in Software Development with a focus on big data and large data pipelines.
- Minimum 3 years of experience to build services and pipelines using Python.
- Expertise with a variety of data processing systems, including streaming, event, and batch (Spark, Hadoop/MapReduce)
- Understanding of at least one NoSQL stores like MongoDB, Elasticsearch, HBase
- Understanding of how data models, sharding and data location strategies for distributed data stores in large scale high-throughput and high-availability environments and their effect in non-structured text data processing
- Experience with running scalable & high available systems with AWS or GCP.
Good to have:
- Experience with Docker / Kubernetes
- Exposure with CI/CD
- Knowledge of Crawling/Scraping
- Entire Work From Home
- Birthday Leave
- Remote Work
At DemandMatrix, our vision is to disrupt the $100 billion sales and marketing intelligence industry by using domain knowledge, machine learning and AI. Fortune 100 companies like Microsoft, Google, Adobe, Amazon, IBM trust us to identify their next customer.
Originally posted on Himalayas
本页面信息整理自 Himalayas,版权归原发布方所有。职位可能随时关闭,投递请以原始页面为准。
本站只做信息聚合展示,不参与招聘流程,也不向求职者收取任何费用。