服务器运维工程师
Server Operations Engineer
关于Centific
Centific是一家前沿的AI数据工厂,通过我们自有的技术平台,整理多样化的高质量数据,为Magnificent Seven及企业客户赋能,实现安全、可扩展的AI部署。我们的团队包括150多名博士和数据科学家,以及4000多名AI从业者和工程师。我们利用集成解决方案生态系统——包括行业领先的合作伙伴和230多个市场中180万名垂直领域专家——创建上下文相关的多语言预训练数据集;定制化行业特定的LLM;以及由向量数据库支持的RAG流程。我们的零距离创新™解决方案可将GenAI成本降低高达80%,并将解决方案推向市场的速度提高50%。
我们的使命是通过将GenAI的最佳实践带给独角兽创新者和企业客户,弥合AI创作者与行业领袖之间的差距。我们旨在帮助这些组织通过大规模部署GenAI释放显著的商业价值,确保他们在技术进步的最前沿保持竞争优势。
职位信息
职位描述摘要
服务器运维工程师负责安装服务器硬件和操作系统,排查可能需要与合作伙伴和供应商协作的服务器和网络问题,维护资产和工单记录,并遵守服务器生命周期管理实践。
服务描述
- 为新安装的服务器安装操作系统。
- 负责重新安装服务器操作系统,解决异常情况。
- 负责日常服务器维护、故障排查、维修及后续修复工作,包括服务器硬件故障排查和组件级故障诊断。
- 维护内部系统中的数据,包括资产管理、工单、机架相关数据。
- 与远程供应商/制造商或其他团队合作,解决服务器批量故障和问题。
- 值班,负责处理业务方提出的问题。
- 收集和检查在线资产状态或问题。
- 对退役或迁移的服务器进行磁盘擦除或其他配置操作。
- 通过更新或编写脚本对部分工具进行改造。
- 如有需要,提交并跟踪部件RMA或媒体销毁流程。
- 服务器网络故障排查。
- 服务器生命周期管理,包括管理性能指标
查看英文原文
About Centific
Centific is a frontier AI data foundry that curates diverse, high-quality data, using our purpose-built technology platforms to empower the Magnificent Seven and our enterprise clients with safe, scalable AI deployment. Our team includes more than 150 PhDs and data scientists, along with more than 4,000 AI practitioners and engineers. We harness the power of an integrated solution ecosystem—comprising industry-leading partnerships and 1.8 million vertical domain experts in more than 230 markets—to create contextual, multilingual, pre-trained datasets; fine-tuned, industry-specific LLMs; and RAG pipelines supported by vector databases. Our zero-distance innovation™ solutions for GenAI can reduce GenAI costs by up to 80% and bring solutions to market 50% faster.
Our mission is to bridge the gap between AI creators and industry leaders by bringing best practices in GenAI to unicorn innovators and enterprise customers. We aim to help these organizations unlock significant business value by deploying GenAI at scale, helping to ensure they stay at the forefront of technological advancement and maintain a competitive edge in their respective markets.
About Job
Job Description Summary
The Server Operations Engineer installs server hardware and OS, troubleshoots server and network issues which may include collaboration with partners and vendors, maintains records of assets and tickets, and adheres to Server Lifecycle Management practices.
Service Description
- Install the operating systems for new mounted servers.
- Responsible for reinstalling the server operation systems, solving the abnormalities.
- Responsible for the daily server maintenance, troubleshooting, repair and follow-up break-fix of the server and other hardware, including server hardware troubleshooting and diagnosing component-level failures.
- Maintain data on internal systems including asset management, ticketing, rack related data.
- Work with remote vendors/manufacturers or other teams to solve server batch failures and problems.
- On-call duty, responsible for dealing with the problems raised by the business owner side.
- Collect and check online assets status or issues.
- Erase drives or other configurations for retiring or relocating servers.
- Retrofit some tools by updating or writing scripts.
- Submit and track the part RMA or media destruction process if needed.
- Server network troubleshooting.
- Server lifecycle management including managing the performance of OxMs.
- Other server operation related work.
Service Requirements
- Bachelor's Degree in Computer science, Electrical engineering or any other relevant fields.
- Strong ability to work under pressure; Strong learning ability, broad technical interest; Strong sense of responsibility, full of enthusiasm for work
- Good communication skills in English, Mandarin is preferred, good team work spirit; Ability to work independently.
- Knowledge of the interdependencies of Data Center functions and technologies, including facilities, with familiarity of data center operational environment and workflows.
- Familiarity with basic Data Analytic Skills.
- Experience and knowledge of out-of-band (OOB)/lights-out management (LOM) server communication methods, such as IPMI management protocol and NC-SI internal communication interface.
- Experience with massive remote OS installation such as PXE boot, with understanding of PXE installation process end-to-end (boot, imaging, OS provisioning, and troubleshooting).
- Can understand and run Bash Shell or Python scripts.
- Familiar with simple automation tools, such as Ansible.
- Familiar with Linux systems, able to locate server hardware/baremetal faults; Proficient with Linux command-line operations and troubleshooting; Strong analytical and problem solving skills.
- Basic TCP/IP knowledge concepts - Subnetting, VLANs, DNS, IPv6; Ability to perform general troubleshooting.
- Knowledge of out-of-band/lights-out server communication methods, such as IPMI, NCSI,
- Strong documentation skills and habits.
Centific is an equal-opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, citizenship status, age, mental or physical disability, medical condition, sex (including pregnancy), gender identity or expression, sexual orientation, marital status, familial status, veteran status, or any other characteristic protected by applicable law. We consider qualified applicants regardless of criminal histories, consistent with legal requirements.
Originally posted on Himalayas