云基础设施运维工程师
Cloud Infrastructure Operations Engineer
关于Centific
Centific是一家前沿的AI数据工厂,通过我们自研的技术平台,精心整理多样化的高质量数据,为Magnificent Seven及企业客户赋能,实现安全、可扩展的AI部署。我们的团队包括150多名博士和数据科学家,以及4000多名AI从业者和工程师。我们利用集成解决方案生态系统——包括行业领先的合作伙伴和230多个市场中180万名垂直领域专家——创建上下文相关的多语言预训练数据集、定制化的行业专用大语言模型(LLM)以及由向量数据库支持的RAG管道。我们为生成式AI(GenAI)提供的零距离创新™解决方案可将GenAI成本降低多达80%,并将解决方案推向市场的速度提高50%。
我们的使命是通过将GenAI的最佳实践带给独角兽创新者和企业客户,弥合AI创作者与行业领袖之间的差距。我们旨在帮助这些组织通过大规模部署GenAI释放显著的商业价值,确保他们在技术进步的前沿保持领先地位,并在各自市场中保持竞争优势。
职位描述
·
1. 与云服务团队、内部团队和云供应商合作,支持基础设施交付和部署活动。
2. 在云服务器和裸机环境中执行操作系统重新安装和重启。
3. 排查涉及Linux系统、云服务器、网络和硬件异常的基础设施问题。
4. 维护内部系统中的操作数据,包括故障工单系统、维修记录和基础设施跟踪工具。
5. 与云供应商合作解决基础设施事件和大规模运营问题。
6. 参与值班轮换并支持事件管理活动。
7. 监控云基础设施健康状况、操作警报和资产利用率。
8. 使用Shell或Python开发或改进操作工具、脚本和自动化流程。
9. 支持运营流程优化和标准化计划。
10. 执行服务器和云网络排查,包括TCP/IP、VLAN、DNS和IPv6相关问题。
11. 支持云群组运维和基础设施性能管理。
12. 协助文档编写、知识共享和运营报告。
要求
1. 计算机科学、电气工程或相关专业的学士学位。
查看英文原文
About Centific
Centific is a frontier AI data foundry that curates diverse, high-quality data, using our purpose-built technology platforms to empower the Magnificent Seven and our enterprise clients with safe, scalable AI deployment. Our team includes more than 150 PhDs and data scientists, along with more than 4,000 AI practitioners and engineers. We harness the power of an integrated solution ecosystem—comprising industry-leading partnerships and 1.8 million vertical domain experts in more than 230 markets—to create contextual, multilingual, pre-trained datasets; fine-tuned, industry-specific LLMs; and RAG pipelines supported by vector databases. Our zero-distance innovation™ solutions for GenAI can reduce GenAI costs by up to 80% and bring solutions to market 50% faster.
Our mission is to bridge the gap between AI creators and industry leaders by bringing best practices in GenAI to unicorn innovators and enterprise customers. We aim to help these organizations unlock significant business value by deploying GenAI at scale, helping to ensure they stay at the forefront of technological advancement and maintain a competitive edge in their respective markets.
About Job
·
1. Collaborate with Cloud service team, internal teams, and cloud vendors to support infrastructure delivery and deployment activities.
2. Perform operating system reinstallation and reboot for cloud servers and bare metal environments.
3. Troubleshoot infrastructure issues involving Linux systems, cloud servers, networking, and hardware-related abnormalities.
4. Maintain operational data across internal systems including fault ticketing systems, repair records, and infrastructure tracking tools.
5. Work with cloud providers to resolve infrastructure incidents and large-scale operational issues.
6. Participate in on-call rotations and support incident management activities.
7. Monitor cloud infrastructure health, operational alerts, and asset utilization.
8. Develop or enhance operational tools, scripts, and automation workflows using Shell or Python.
9. Support operational process optimization and standardization initiatives.
10. Perform server and cloud network troubleshooting, including TCP/IP, VLAN, DNS, and Ipv6 related issues.
11. Support cloud fleet operations and infrastructure performance management.
12. Assist with documentation creation, knowledge sharing, and operational reporting.
Requirements
1. Bachelor’s Degree in Computer Science, Electrical Engineering, or related fields.
2. Experience in cloud infrastructure operations, server operations, or data center related environments.
3. Strong troubleshooting and analytical skills in Linux and infrastructure environments.
4. Experience with automated provisioning and OS deployment technologies such as PXE, iPXE, or provisioning pipelines.
5. Familiarity with public cloud platforms such as Oracle, Amazon Web Services, Google, or Microsoft.
6. Ability to understand, execute, and write Shell/Bash or Python scripts.
7. Familiarity with automation and infrastructure management tools such as Ansible, GitLab CI/CD, Terraform, or cloud SDKs.
8. Knowledge of Linux systems, cloud networking, and server lifecycle management.
9. Strong understanding of TCP/IP networking concepts including subnetting, VLANs, DNS, IPv6, and basic routing.
10. Familiarity with infrastructure monitoring, operational tooling, and incident management processes.
11. Strong communication, collaboration, and documentation skills.
12. Familiarity with large-scale cloud fleet operations is preferred.
13. Experience supporting GPU infrastructure, firmware lifecycle management, or RDMA networking is a plus
Hourly Rate: $50
Centific is an equal-opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, citizenship status, age, mental or physical disability, medical condition, sex (including pregnancy), gender identity or expression, sexual orientation, marital status, familial status, veteran status, or any other characteristic protected by applicable law. We consider qualified applicants regardless of criminal histories, consistent with legal requirements.
Originally posted on Himalayas