高级应用支持工程师 – 美国
Senior Application Support Engineer – US
关于公司
Zact 是一家金融科技创新者,致力于一个核心理念:组织需要更简单的费用和支付管理系统,使支出员工与财务和会计部门保持一致,同时提供固有的防护机制,并与财务系统进行持续对账。
Zact 让公司的各个部分都能按照规则进行支付。
职位描述
我们正在为我们的商业卡和支付平台建立专门的应用支持职能——这是一个由 GCP 托管的基于微服务的 SaaS 产品,供银行客户使用。该高级应用支持工程师将是该团队的技术核心,带领两名初级分析师。
这是一个注重调查、系统思维和自动化优先的角色。你将负责复杂事件的端到端排查,构建监控和警报基础架构,并系统地自动化拖慢支持团队效率的重复任务。你还将是银行客户升级问题的主要技术联系人,以及整个团队所依赖的知识基础设施的构建者。
地点
- Zact 是一家全球公司,在美国、欧洲和印度均有办公地点。
- 我们提供 100% 远程工作选项。
- 申请此职位的候选人应位于美国。
这个职位与标准支持职位的不同之处
你将带领团队并和他们一起处理工单——这是一个亲力亲为的领导角色,而不是监督性质的。除了处理工单外,你还需要识别模式,永久消除根本原因,并帮助设计和构建团队所需的后台工具和支持功能。你将对产品的监控方式以及工程团队在支持性方面的责任有实际影响力。
你将负责的工作
事件领导与分类
- 领导复杂、多服务事件的分类——从第一个信号到根本原因,全程记录时间线
- 使用 GCP Cloud Logging、Cloud Monitoring 和分布式追踪来跨微服务边界追踪故障
- 编写 MySQL 查询来调查生产环境读副本中的交易、卡片和账户数据问题——始终遵循严格的变更控制
- 负责事件回顾(PIR):根本原因、相关因素、时间线和预防措施
- 作为银行客户问题的高级技术升级联系人——冷静、精准、以解决方案为导向
监控与可观测性
- 设计并维护团队的警报系统
查看英文原文
About Us
Zact is a Fintech innovator dedicated to a singular idea : Organizations need simpler expense and payment management systems that align the spending employee with finance and accounting while providing inherent guardrails and continuous reconciliation with financial systems.
Zact enables all parts of the company to Pay by the Rules.
About the Role
We are building out a dedicated Application Support function for our commercial card and payments platform — a GCP-hosted, microservices-based SaaS product used by banking clients. This Senior Application Support Engineer will be the technical anchor of that team, leading a team of two junior analysts.
This is an investigative, system-thinking, and automation-first role. You will own complex incident triage end-to-end, build the monitoring and alerting foundation, and systematically automate the repetitive tasks that slow support teams down. You will also be the primary technical point of contact for banking client escalations and the person who builds the knowledge infrastructure the whole team runs on.
Location
- Zact is a global company with multiple locations across US, Europe and India.
- We provide 100% Remote Working Option.
- Candidates applying for this role should be located in the United States.
What Makes This Role Different From a Standard Support Position
You will lead the team and work the queue alongside them — this is a hands-on lead role, not a supervisory one. Beyond ticket handling, you are expected to identify patterns, eliminate root causes permanently, and help design and build the back-office tools and support functions the team needs. You will have real influence over how the product is monitored and how engineering is held accountable for supportability.
What You’ll Own
Incident leadership & triage
- Lead triage of complex, multi-service incidents — from first signal to root cause, with a documented timeline throughout
- Use GCP Cloud Logging, Cloud Monitoring, and distributed tracing to trace failures across microservice boundaries
- Write MySQL queries to investigate transaction, card, and account data issues in production read replicas — following strict change-control at all times
- Own post-incident reviews (PIRs): root cause, contributing factors, timeline, and preventive actions
- Act as the senior technical escalation for banking client issues — calm, precise, and solution-oriented
Monitoring & observability
- Design and maintain the team’s alerting strategy in GCP Cloud Monitoring — setting thresholds, reducing alert noise, and ensuring the right signals fire at the right severity
- Build and maintain dashboards (Grafana, GCP Monitoring, or equivalent) covering transaction throughput, API error rates, service latency, and MySQL query performance
- Define and track SLIs/SLOs for key product flows — payments authorisation, settlement processing, card management — and surface degradation early
- Instrument new services in partnership with engineering: ensure every new microservice ships with adequate logging, structured log fields, and correlation IDs
Back-office tooling & automation
- Help design and build the internal back-office tools and support functions the team needs to operate effectively — this is not off-the-shelf tooling configuration; it includes scoping, designing, and writing the tools from scratch where needed
- Example: an internal investigation dashboard that pulls GCP log context and MySQL transaction data for a given incident ID in one view
- Example: automated log correlation scripts that surface the root cause of a known error pattern in under 2 minutes
- Example: MySQL query templates that auto-populate with a transaction ID and return the full investigation snapshot
- Example: alert-to-ticket automation that pre-populates incident records with log context and affected service
- Write and maintain automation tooling in Python and/or Bash — scripts, schedulers, data reconciliation jobs
- Build internal CLI tools or runbook-linked scripts that junior analysts can invoke safely without senior oversight
- Identify manual reconciliation or reporting tasks that can be replaced by scheduled queries or GCP Cloud Functions
- Own the tooling roadmap for the support function — prioritise what gets built next based on toil volume and incident frequency
Client technical support
- Serve as the senior technical point of contact for all customer-facing product issues that escalate beyond first-line resolution
- Write clear, jargon-free incident communications for banking client operations teams during and after incidents
- Work directly with client technical teams (bank IT, card ops) to diagnose integration issues, file format mismatches, or API usage errors
- Maintain a client-facing known-issues log and service status communication cadence
Documentation & knowledge management
- Build and own the team’s runbook library — step-by-step investigation procedures for every known failure pattern across the product suite
- Write and maintain a MySQL query reference for common support investigations (transaction lookup, card status, spend limit checks, settlement reconciliation)
- Document every new incident type in a known-issues catalogue: symptom, root cause, resolution, recurrence trigger, and prevention status
- Produce onboarding materials that bring a new junior analyst to independent productivity within 6 weeks
- Maintain a living architecture reference showing how services connect, what each service owns, and which logs to check first for each component
Team leadership — hands-on
- Lead a small, global support team — you are a working lead, not a manager removed from the queue; you will handle tickets alongside the analysts based on volume and complexity
- Mentor and develop two junior application support analysts — review their investigations, coach their SQL and log techniques, and build their confidence on live incidents
- Run daily standups and weekly incident reviews across time zones — keeping a distributed team coordinated and nothing falling between shifts
- Define team support processes: ticket triage criteria, severity classification, escalation thresholds, SLA tracking
- Partner with engineering and product to represent the support team’s perspective — surfacing recurring issues, requesting observability improvements, and advocating for supportability in new releases
Technology Environment
- Cloud platform: Google Cloud Platform (GCP)
- Log & monitoring: GCP Cloud Logging, Cloud Monitoring, Pub/Sub; Grafana
- Database: MySQL — production read replicas for investigation; strict change-control for any writes
- Architecture: Microservices — REST APIs with correlation ID tracing across service boundaries
- Scripting & automation: Python (primary), Bash — for tooling, automation, and log parsing
- Ticketing: Internal ticketing system (Jira or equivalent)
- Product domain: Commercial card programs — card issuance, authorization, settlement, spend controls, reporting
- Client environment: Banking clients (community and regional banks) — regulated, SLA-bound, audit-sensitive
What We’re Looking For
- 8+ years in application support, platform operations, or a technical operations role with a team leadership component
- MySQL proficiency — writing complex queries, reading execution plans, diagnosing slow queries, understanding index behaviour
- GCP experience — hands-on with Cloud Logging (log queries, log-based metrics), Cloud Monitoring (alerting policies, uptime checks), and ideally Cloud Functions or Pub/Sub
- Microservice troubleshooting — comfortable tracing a failure across 3–5 services using correlation IDs and structured logs
- Python or Bash scripting — you have written automation tools that saved real time, not just one-off scripts
- Methodical incident investigation — you document as you go, form hypotheses, test them, and do not guess
- Strong written communication — incident updates, PIRs, client communications, and runbooks are all part of your normal output
- Experience building or significantly contributing to a team knowledge base or runbook library
Nice to Have
- Experience with distributed tracing tools — Cloud Trace, Jaeger, Zipkin, or equivalent
- Grafana dashboard authoring
- Exposure to SRE practices: SLO definition, error budgets, toil reduction
- Familiarity with GCP Cloud Functions, Cloud Scheduler, or Workflows for automation
- Experience administering Zendesk — configuring ticket views, SLA policies, macros, triggers, automations, and reporting; familiarity with Zendesk API or webhook integrations is a bonus
- Background in a regulated or compliance-sensitive environment — fintech, banking, healthcare, or telecoms
- Familiarity with commercial card concepts (authorisation flows, settlement, BIN management, spend controls) — helpful but fully trainable
Note on Domain Experience
Prior fintech or banking experience is helpful but not a requirement. What matters is depth in GCP, MySQL, and microservice troubleshooting — and the instinct to automate. We provide structured onboarding into the product domain and a growing runbook library to support ramp-up.
Working Hours & Expectations
- Standard shift of 8–10 hours per day, with some overlap with US hours.
- Remote-friendly; in-person collaboration expectations to be discussed by region
Originally posted on Himalayas