Jobs · himalayas
Operations Engineer
Keystone AI · United States · Posted 6d ago
About the role
Ensure stability and performance of the trading platform, focusing on production systems and infrastructure.
Operations Engineer Hybrid or Remote (UK-based), Hybrid or Remote (US-based) or In-Office (Singapore-based) Full-Time About Us Keystone AI is the 2026 OTC Trading Platform of the Year (). Our network handles over $50 billion in daily trading volumes across FX, Equities, and Cryptocurrency, connecting 40+ of the world's leading liquidity providers. We build trading systems that operate at the edge of what's technically possible — where nanoseconds matter and a dropped message costs real money. Our engineering team is small, senior, and deeply invested in the craft of building cutting edge, reliable, high-performance systems. The Role As an Operations Engineer, you will ensure the stability, performance, and availability of the Reactive trading platform. This is a hands-on, technical operations role — focused on production systems, infrastructure, and trading workflows rather than client relationship management. You'll work closely with Engineering, Product, and internal stakeholders as the first line of defence for live trading systems. TechOps owns Level 1 and Level 2 support — monitoring response and alert triage through to the deeper technical investigation, configuration delivery, and post-release verification that sits beneath it — escalating to Level 3 (Engineering and Infrastructure) only when an issue genuinely requires code changes or deep infrastructure intervention. While the role involves supporting client onboarding and integrations, the emphasis is on technical delivery, system correctness, and incident response — not sales or account management. Reactive runs a globally distributed, follow-the-sun operations team, and we are hiring across our UK, US, and Singapore hubs. You'll provide operational coverage across your regional trading session — the eyes on the platform when your region is live — and hand over cleanly to the next hub as the day moves. In each region you own the full TechOps remit: monitoring and incident response through the session, plus the configuration delivery, onboarding execution, and change work that keeps the platform growing. The UK and US roles are remote; the Singapore role is office-based. We are a small team that has turned operations into a genuine engineering discipline, and we are investing heavily in AI-assisted tooling and knowledge engineering to keep it that way. This role suits someone with experience in trading or FinTech operations who enjoys working close to production systems, solving complex problems, writing things down properly, and taking ownership in a fast-moving environment. What You'll Work On Production Monitoring & Incident Response: Monitor live trading systems across multiple regions, venues, and asset classes — you are the eyes on the platform during the trading session Own Level 1 and Level 2 response: triage alerts, run standard runbook actions (restart, rollback, failover, feature flag), investigate deeply, declare severity, and coordinate incidents — escalating to Level 3 (Engineering / Infrastructure) only when code or deep infrastructure intervention is genuinely required Participate in failover and disaster-recovery readiness — rehearsed, playbook-driven, and time-bound Drive root-cause analysis and blameless post-incident reviews, turning every incident into a runbook or a tooling improvement Configuration Delivery & Production Change: Deliver the production-change pipeline that brings new clients, liquidity providers, venues, and capabilities live — the largest share of the work Configure and maintain FIX sessions by function — streaming, RFQ, algo, and post-trade — alongside cross-rates, indications of interest, subscriptions, certificate rotations, and pool registrations Execute changes safely against the platform's deployment configuration, with schema updates, ACL coordination, and post-release verification Support migrations, decommissioning, and the follow-on work that keeps the production estate coherent as it grows Client Onboarding & Connectivity: Own the technical execution of client and LP onboarding: credentials, IPs, conformance and UAT testing, cross-connects, day-1 readiness, and post-go-live verification Act as the technical point of contact during onboarding, giving Client Services the information they need to manage the relationship Monitor and validate FIX message flows, identifying issues related to execution, routing, or market-data Maintain clear, accurate documentation for configurations, onboarding steps, and troubleshooting Knowledge Engineering, Automation & Continuous Improvement: Contribute to capacity and infrastructure planning — session load balancing, gateway placement, and utilisation monitoring — jointly with Performance Engineering and Infrastructure Coordinate with external vendors and counterparties: network and colocation providers, post-trade vendors, and liquidity providers and their infrastructure teams Identify recurring issues and contribute to long-term fixes and preventative improve
Read the full posting on himalayas →
FAQ
Is the Operations Engineer role at Keystone AI remote?+
This Operations Engineer position is listed as remote (United States).
What seniority level is this Operations Engineer role?+
This is a mid level position.
How do I apply for the Operations Engineer role at Keystone AI?+
Use the "Apply on himalayas" button to open the original posting on himalayas, where you can submit your application directly to Keystone AI.