We’re looking for a creative, experienced DevOps engineer based in London to join our fintech startup at the ground floor. If you are excited about independence, impact, and contributing to the battle against scams and financial crime, read on to learn more ⬇️
About the role 🎭
Banks approve or reject payments with almost no context. We're building the intelligence layer that changes that — running real-time investigations on payments before they clear. The technical challenge is to keep this system running 24/7/365 with the reliability that banks demand, while the platform underneath evolves rapidly. Think enterprise resilience with startup velocity 🚀
What you’ll work on
Own the reliability and operability of our production systems — monitoring, alerting, incident response, and post-incident learningMaintain and evolve our infrastructure-as-code estate (Terraform on GCP). Make it easy for engineers to ship safely and hard to ship dangerouslySecure our infrastructure defaults to ensure that the easy path is the safe pathDesign and implement observability across our stack: metrics, traces, logs, and the dashboards and alerts that make them usefulDrive incident response maturity, from detection through resolution to follow-up. You'll get an opportunity to shape what "good" looks like hereEvaluate and introduce tooling for AI-native reliability challenges: monitoring non-deterministic systems, detecting drift in agent behaviour, and understanding failure modes that traditional APM doesn't coverImprove deployment pipelines and release processesThis is not a "keep the lights on" role. You'll be building the foundations that let a fast-moving team ship with confidence — and you'll have real influence over how we grow.
Why this is hard
We're a small team shipping to large banks. Our systems are critical infrastructure — when we're unavailable, real fraud investigations stop and real payments are delayed. To compound resilience further, our increasing footprint of production AI agents are non-deterministic by nature, which means traditional monitoring and alerting isn't enough. You'll need to use (or invent!) cutting-edge approaches to observability for systems where correct behaviour isn't always the same twice. And you'll be doing this as (likely) the first person in the role, which means building the culture and the tooling from scratch.
Our stack today:
Infra: GCP, Terraform, Cloud Run (and soon, k8s!), managed GCP datastoresBackend: Python, KotlinCI/CD: GitHub ActionsMonitoring: GCP Cloud Monitoring, Cloud Logging (we know we need more — that's where you come in)Workflow orchestration: TemporalBut honestly, if you've built production infrastructure that teams depend on, the specific tools matter less than your judgement on when to reach for them.
About YOU 🦄
Here are some thoughts on who would be successful in the role. We know we won’t be able to assess these from your CV, but our interview process will be designed around them:
Ownership mindset: You see a gap in reliability and you fill it. You don't wait to be told what to fix; you talk to the team, understand what they need to ship safely, and build it.Systems thinking: You understand how distributed systems fail and you design for it. You think about blast radius, graceful degradation, and recovery by default.Pragmatism: You know when to build, when to buy, and when to duct-tape. You've seen enough production systems to know that perfect is the enemy of reliable.Communication: You can run an incident calmly, write a post-mortem clearly, and explain a monitoring gap to engineers who aren't thinking about infrastructure.Collaboration: You're direct, proactive, generous to others, and low ego. No alert is too small if a customer depends on it!Some hard skills we’d value
Infrastructure-as-code: Strong Terraform experience managing real production infrastructure. GCP experience is a bonus.Observability: You’ve built monitoring, alerting, and dashboarding for distributed systems in production. Bonus if you've worked with non-deterministic or ML-powered systems.Incident response: You've been on-call, you've run incidents, and you've built the processes that make both better.Security foundations: You understand network security, least-privilege access, secrets management, and how to build infrastructure that's secure by default.CI/CD: Experience building and maintaining deployment pipelines. You care about deploy velocity and rollback safety.Years of experience: 5 or more. We're looking for someone who can structure their own work and define what good looks like (we're not overly strict on exact number of years!).Stage of experience: We value people who have worked in early-stage companies — ideally as the first or only infrastructure/DevOps/SRE hire. You're comfortable withambiguity and building from scratch.Industry: Prior experience in fintech or other regulated verticals is a plus.Not got 100% of qualifications? Research shows that women and under-represented minorities are less likely to apply for a role if they do not have 100% of the requirements. We’re paying attention to this, so please feel free to apply regardless (<2mins) or send us a note on info@tunicpay.com if you’d like to discuss if there’s a fit.