Reliability Engineer (Trading Platforms)
Io Tech Solutions- Location
- Hong Kong, Hong Kong
- Workplace
- —
- Employment
- Full Time
- Salary
- —
Posted 29d ago
Job Description
- Automate repeatable triage workflows to help first-line teams respond faster and more consistently (e.g., alert enrichment, routing, correlation, and operational runbooks).
- Identify monitoring/alerting gaps and drive improvements in visibility and alert quality.
- Track reliability and availability across critical trading applications and their dependencies. Partner with users, development teams, and IT to pinpoint where service levels are degrading.
- Triage incoming alerts, issues, and escalations—assessing impact, urgency, and ownership.
- Determine when incident criteria are met, declare incidents, and act as Incident Commander.
- Coordinate responders and stakeholders; keep incident calls focused on facts, mitigation, and recovery.
- Maintain clear timelines, actions, and status updates throughout the incident lifecycle.
- Recover and stabilize systems using approved runbooks. Escalate cleanly through the defined support/development path when the issue exceeds documented recovery steps.
- Support post-incident review (PIR) follow-ups and recurring issue reviews.
- Ensure smooth handovers across EMEA, AMER, and APAC using a single global model: one incident standard and one handover process.
Requirements
- Experience in production operations, SRE, NOC/command center, trading operations, or a comparable first-line technical role—ideally in a trading, financial services, or other latency-sensitive environment.
- Strong triage and prioritization skills: you can separate facts from assumptions under pressure and keep the response moving.
- Clear communication (verbal and written): status updates are understandable to both traders and engineers.
- Broad technical understanding (not just deep specialist knowledge): enough to collaborate effectively across domains and interfaces.
- Solid Linux and networking fundamentals, plus the ability to quickly interpret alerts, logs, dashboards, and symptoms.
- Working knowledge of common operational tasks across adjacent teams (application support, infrastructure, connectivity, data).
- Familiarity with incident and observability tooling (e.g., PagerDuty or equivalent, Jira Service Management or equivalent, Grafana, Prometheus, log search).
- Scripting/automation skills (Python preferred; Bash and Go are a plus), applied to triage, enrichment, routing, and correlation (not product code).
- Exposure to containerized/cloud-hosted production environments (Kubernetes, Docker, GCP) is a plus.
Skills
- Linux
- PagerDuty
- Jira
- Grafana
- Prometheus
- Python
- Bash
- Go
- Kubernetes
- Docker
- GCP
More jobs at Io Tech Solutions
All 78Java Engineer, FIX
Io Tech Solutions · Hong Kong, Hong Kong · 7d ago
Senior Java Developer - FIX Connectivity (FinTech)
Io Tech Solutions · Hong Kong, Hong Kong · 7d ago
Python Risk Developer
Io Tech Solutions · HongKong, Hong Kong · 9d ago
DevOps Engineer | Singapore or Hong Kong
Io Tech Solutions · Singapore · 9d ago
Python Risk Engineer
Io Tech Solutions · Hong Kong, Hong Kong · 11d ago