具备VMware的Windows服务器自动化工程师
Windows Server Automation Engineer with VMware
长期合同
需要具备Windows自动化、日志记录和管理方面的专家级知识。
关键职责:
统一监控框架:构建集成解决方案,利用现有工具(Splunk、Dynatrace、安全平台)以及来自Windows/Linux服务器、基础设施组件(存储阵列、SAN交换机、网络设备)、数据库(Oracle、SQL Server、MySQL、MongoDB)、备份系统(Rubrik、Data Domain、Infinibox)、计算节点(Dell服务器)、VMware环境、IBM Power/AIX和IBM LinuxOne的本地日志。
自动化警报与主动响应:开发智能警报机制和自动化修复工作流,减少人工干预并加快事件解决。
数据整合与缺口填补:聚合和标准化来自多个来源的数据,包括平台工具和本地日志,以填补可见性缺口并提供可操作的洞察。
仪表板开发:创建通用的基于GUI的仪表板,用于跨所有基础设施层的实时监控、警报和报告。
技能与工具:使用Ansible、Python、PowerShell、shell脚本和GUI开发来交付可扩展的自动化解决方案。
职位技能要求与经验:
核心专业技术技能:
自动化与脚本编写:
熟练掌握Python、Ansible、PowerShell和shell脚本(Bash/Korn)。
能够开发用于监控、警报和修复的自动化工作流。
监控与日志工具:
有Splunk、Dynatrace和其他企业监控平台的实际经验。
熟悉从多个来源(操作系统、应用程序、基础设施组件)进行日志聚合和解析。
基础设施知识:
对Linux(RHEL)和Windows Server环境有深入了解。
接触过VMware、IBM Power/AIX和IBM LinuxOne系统。
了解存储阵列、SAN交换机、网络交换机和IP流量监控。
有备份平台(Rubrik、Data Domain、Infinibox)的经验。
熟悉数据库系统(Oracle、SQL Server、MySQL、MongoDB)。
GUI开发:能够构建用于实时监控和警报的仪表板界面(使用Flask/Django等Python框架或类似工具)。
其他技能:
数据整合:能够从多个来源聚合和标准化数据以实现统一警报。
安全与合规意识:了解基础设施监控中的安全日志和合规要求。
问题解决能力
查看英文原文
Long term contract
Need someone, SME level knowledge with Windows automation, Logging and administration.
Key Activities:
Unified Monitoring Framework: Build an integrated solution leveraging existing tools (Splunk, Dynatrace, security platforms) and local logs from Windows/Linux servers, infrastructure components (storage arrays, SAN switches, network devices), databases (Oracle, SQL Server, MySQL, MongoDB), backup systems (Rubrik, Data Domain, Infinibox), compute nodes (Dell servers), VMware environments, IBM Power/AIX, and IBM LinuxOne.
Automated Alerting & Proactive Response: Develop intelligent alerting mechanisms and automated remediation workflows to reduce manual intervention and accelerate incident resolution.
Data Integration & Gap Closure: Aggregate and normalize data from multiple sources, including platform tools and local logs, to fill visibility gaps and provide actionable insights.
Dashboard Development: Create a common GUI-based dashboard for real-time monitoring, alerting, and reporting across all infrastructure layers.
Skills & Tools: Utilize Ansible, Python, PowerShell, shell scripting, and GUI development to deliver scalable automation solutions.
Job Skill Requirements and Experience:
Core Technical Skills:
Automation & Scripting:
Proficiency in Python, Ansible, PowerShell, and shell scripting (Bash/Korn).
Ability to develop automation workflows for monitoring, alerting, and remediation.
Monitoring & Logging Tools:
Hands-on experience with Splunk, Dynatrace, and other enterprise monitoring platforms.
Familiarity with log aggregation and parsing from multiple sources (OS, applications, infrastructure components).
Infrastructure Knowledge:
Strong understanding of Linux (RHEL) and Windows Server environments.
Exposure VMware, IBM Power/AIX, and IBM LinuxOne systems.
Knowledge of storage arrays, SAN switches, network switches, and IP traffic monitoring.
Experience with backup platforms (Rubrik, Data Domain, Infinibox).
Familiarity with database systems (Oracle, SQL Server, MySQL, MongoDB).
GUI Development: Ability to build dashboard interfaces for real-time monitoring and alerting (using frameworks like Flask/Django for Python or similar).
Additional Skills:
Data Integration: Ability to aggregate and normalize data from multiple sources for unified alerting.
Security & Compliance Awareness: Understanding security logs and compliance requirements for infrastructure monitoring.
Problem-Solving & Creativity: Ability to identify gaps in current monitoring and design innovative solutions.
Experience:
5+ years in infrastructure automation or systems engineering roles.
Proven track record in building automation frameworks and monitoring solutions.
Experience working in large-scale, distributed environments with global teams.
Prior involvement in proactive alerting and automated remediation projects is highly desirable.
Originally posted on Himalayas