高级基础设施工程师 — 管理服务
Senior Infrastructure Engineer — Managed Services
我们致力于在全球展示卓越的菲律宾人才。除了教育背景和专业经验,我们更加重视热情、奉献精神、忠诚度以及对个人成长的积极态度。我们的基础建立在一组核心价值观——R.I.G.H.T.,代表可靠性、诚信、目标导向的心态、幸福感和团队合作。这些价值观定义了我们的身份,并指导我们实现成功的路径。
如果你热衷于提升自己的职业发展,建立真诚的联系,并成为一支致力于你成功团队的一员,欢迎加入我们!
职位概述
高级基础设施工程师(L3)是托管服务部门中的重要技术资源,负责处理并解决复杂的基础设施事件,提供高级故障排除和升级支持,并确保客户环境的稳定性、性能和可用性。该职位结合了实际操作责任与技术领导力,参与根本原因分析、基础设施设计、项目交付、文档编写和持续服务改进。
该职位需要在企业级基础设施技术方面有扎实的实际经验,包括VMware vSphere/ESXi(建议v7/v8版本)、Windows Server、Active Directory、DNS、DHCP、组策略、Linux管理(RHEL/Ubuntu/CentOS)、企业存储平台、备份与灾难恢复解决方案,以及Azure和AWS等云平台。具备Cisco UCS、PKI、基础设施生命周期管理、迁移、补丁、升级和托管服务运营的经验将被视为加分项。
主要职责
- 事件管理:负责复杂基础设施事件的全生命周期管理,从发现到解决及事后回顾,确保对技术结果、稳定性和文档的责任感。
- 关键事件领导:领导关键事件的调查与解决,协调内部团队、供应商和客户利益相关者之间的合作。
- 技术决策:做出技术决策并对基础设施问题进行全流程负责,无需依赖进一步升级即可推动解决。
- 项目交付:主导基础设施项目的实施,包括范围界定、估算、设计和部署。
- 根本原因分析:执行根本原因分析并制定长期修复策略,防止类似问题再次发生。
查看英文原文
We at Acxelsus are on a mission to showcase exceptional Filipino talent globally. Beyond educational credentials and professional experience, we highly prioritize passion, dedication, loyalty, and a proactive attitude towards personal growth. Our foundation is built upon a set of core values - R.I.G.H.T. - representing Reliability, Integrity, Goal-Oriented mindset, Happiness, and Teamwork. These values define our identity and guide our approach to achieving success.
If you are passionate about advancing your career, forging genuine connections, and becoming part of a team committed to your success, join us!
Job Overview
The Senior Infrastructure Engineer (L3) is a strong technical resource within the Managed Services organization, responsible for owning and resolving complex infrastructure incidents, providing advanced troubleshooting and escalation support, and ensuring the stability, performance, and availability of customer environments. This role combines hands-on operational ownership with technical leadership, contributing to root cause analysis, infrastructure design, project delivery, documentation, and continuous service improvement.
The role requires strong hands-on experience supporting enterprise infrastructure technologies, including VMware vSphere/ESXi (v7/v8 preferred), Windows Server, Active Directory, DNS, DHCP, Group Policy, Linux administration (RHEL/Ubuntu/CentOS), enterprise storage platforms, backup and disaster recovery solutions, and cloud platforms such as Azure and AWS. Experience with Cisco UCS, PKI, infrastructure lifecycle management, migrations, patching, upgrades, and managed services operations is highly valued.
Key Responsibilities
- Incident Ownership: Own the full lifecycle of complex infrastructure incidents, from detection through resolution and post-incident review, ensuring accountability for technical outcomes, stability, and documentation.
- Critical Incident Leadership: Lead investigation and resolution of critical incidents, driving coordination across internal teams, vendors, and customer stakeholders.
- Technical Decision-Making: Make technical decisions and take end-to-end ownership of infrastructure issues, driving resolution without reliance on further escalation.
- Project Delivery: Lead technical delivery of infrastructure projects, including scoping, estimation, design, and implementation.
- Root Cause Analysis: Perform root cause analysis and develop long-term remediation strategies to prevent recurrence.
- Change Validation: Validate and approve infrastructure changes, designs, and remediation strategies.
- Mentorship: Provide technical guidance and mentorship to junior engineers.
- Documentation: Contribute to architecture documentation, diagrams, and operational runbooks.
- Automation & Process Improvement: Design and implement automation and process improvements to enhance reliability and operational efficiency.
- Change Management: Participate in change management and project planning activities.
- Time Tracking: Accurately track time and project activities within the time tracking system, maintaining a target billable utilization of approximately 70%.
Qualifications
- Education: Bachelor's degree in Information Technology, Computer Science, or a related field, or equivalent work experience.
- Experience: 6–10 years of experience supporting enterprise IT infrastructure, ideally within a Managed Service Provider (MSP) environment supporting multiple client environments simultaneously.
- Virtualization: Deep, hands-on experience with VMware vSphere 7.x/8.x — including ESXi administration, vCenter, vMotion, HA/DRS, distributed virtual switches, VM lifecycle management, and vSphere upgrade planning (compatibility validation, upgrade sequencing, rollback planning).
- Operating Systems: Strong administration experience with Windows Server 2019/2022 (AD DS, Group Policy, DNS/DHCP, IIS, WSUS) and Linux (RHEL/CentOS/Rocky, Ubuntu/Debian, or SLES), including performance troubleshooting and root-cause diagnostics using standard OS-level tools.
- Storage: Hands-on experience with enterprise storage platforms (e.g., NetApp, Pure Storage, Dell EMC, HPE Nimble/3PAR) and/or hyperconverged infrastructure (vSAN, Nutanix).
- Cloud Platforms: Deep experience with Google Cloud Platform (GCP) in a production capacity — core services such as Compute Engine, VPC networking, Cloud Storage, Cloud IAM, and Cloud Load Balancing; experience with GKE, Cloud SQL, or Pub/Sub is a plus.
- Troubleshooting & Design: Strong, demonstrable experience troubleshooting complex multi-system infrastructure issues end-to-end, and designing/implementing infrastructure solutions (e.g., HA cluster design, backup/DR strategy, GCP architecture) rather than only executing pre-defined tasks.
- Incident Ownership & RCA: Proven ability to own critical incidents from detection through post-incident review, including structured root cause analysis (timeline reconstruction, contributing factors vs. root cause, corrective/preventive actions).
- ITSM & Change Management: Familiarity with ticketing/ITSM platforms (ServiceNow, Jira Service Management, ConnectWise, or similar) and ITIL-based change management processes, including change advisory board (CAB) participation.
- Soft Skills: Strong analytical and problem-solving abilities; proven customer-facing communication skills, including delivering difficult technical updates under pressure; ability to work independently and take full ownership of technical outcomes; experience mentoring or providing technical guidance to junior engineers.
- Language: Strong English communication skills, both written and verbal.
- Availability: Willingness and prior experience working US EST/PST hours, including after-hours or weekend coverage for maintenance windows, incidents, or on-call rotation.
Preferred Qualifications
- Certifications: VMware VCP-DCV, Microsoft certifications, GCP Associate Cloud Engineer (ACE) or Professional Cloud Architect (PCA), Cisco CCNP/CCIE Data Center, or Citrix CCP-V.
- Cisco UCS: Hands-on experience with Cisco UCS — blade/rack server administration, Fabric Interconnects, UCS Manager (UCSM) or Intersight, service profiles, and firmware management/troubleshooting.
- Citrix CVAD: Experience administering Citrix Virtual Apps and Desktops (CVAD 7.x, 2203, 2305 LTSR) or Citrix DaaS — Delivery Controllers, StoreFront, NetScaler/ADC, machine catalogs and delivery groups, VDA troubleshooting, and MCS vs. PVS provisioning decisions.
- Cisco Unified Communications Manager (CUCM): Experience with CUCM 11.x/12.x/14.x — device provisioning, dial plans, route groups/lists, SIP trunk configuration, and registration troubleshooting.
- Networking, Firewalls & Identity: Hands-on policy configuration (not just monitoring) with Palo Alto or Cisco FTD/FMC (or Fortinet/Check Point); Active Directory administration beyond basic user management (GPO design, AD Sites and Services, replication); Okta or other SSO/MFA identity platform experience.
- Backup & Disaster Recovery: Experience with Veeam, Commvault, Veritas NetBackup, or Zerto; direct involvement in a DR test or failover, and ability to design a backup/recovery strategy against defined RTO/RPO targets.
- Automation: Experience writing and maintaining automation scripts (PowerShell for Windows, Bash/Python for Linux) for provisioning, reporting, or operational tasks.
- MSP Experience: Experience working in a Managed Service Provider (MSP) environment with billable utilization tracking (~70% target) and exposure to ITIL-based service delivery models.
- Customer-Facing Projects: Experience serving as technical lead on a customer-facing infrastructure project, from scoping through delivery and handover.
- AI Tool Fluency: Practical, hands-on use of AI tools (e.g., Claude, ChatGPT, Copilot) in day-to-day infrastructure troubleshooting or documentation work, with demonstrated critical evaluation of AI output rather than uncritical reliance on it.
Work Arrangement & Schedule
- Fully remote following EST Time Zone
- After-hours and weekend coverage may be required on occasion.
- All necessary equipment is provided
Ready to experience the #AcxelsusAdvantage? Join our team today!
Originally posted on Himalayas