A shortlist in 72 hours, contributing inside two weeks, and a 30-day guarantee if the fit is wrong.
72-hour shortlist · 30-day replacement guarantee · Month-to-month
A DevOps engagement rarely fails because the infrastructure doesn’t work. It fails because it works and nobody else understands it.
Six months later there’s a Terraform module nobody can safely change, a Jenkins pipeline held together by three shell scripts, and a Kubernetes cluster your team is afraid to touch. The engineer who built it has moved on. What was meant to reduce risk has become the largest single point of failure in your engineering organisation.
The right DevOps hire is measured by what your team can do after they leave, not by what they build while they’re there. We screen for documentation habits and handover discipline as deliberately as for technical skill, because the second one is worthless without the first.
Terraform or Pulumi with modules, state management, and a plan your team can read. Not scripts that happen to be checked in.
There is a wide gap between deploying to an existing cluster and operating one. We check which they've actually done.
Pipelines that fail fast, give useful errors, and don't require tribal knowledge to debug at 6pm on a Friday.
Metrics, logs, traces, and alerting that fires on real problems rather than noise your team learns to ignore.
Cloud bills grow quietly. An engineer who thinks about instance sizing and storage lifecycle saves more than their rate.
IAM least privilege, secret management, network boundaries. The things a compliance review will ask about.
The bill is the visible cost. The bus factor is the expensive one.
Here is the shape it takes. An engineer builds your platform with Terraform, but keeps state in a local backend or a bucket with no locking. Two people run apply at once and the state file no longer matches reality. Now Terraform wants to destroy and recreate a database because it believes it isn’t there. Someone fixes it by hand, the drift widens, and within a year the safest option is to stop using Terraform for that resource — which quietly means you no longer have infrastructure as code, just a directory that used to be.
The Kubernetes version is similar. A cluster set up by one person, upgraded never, running a control plane two or three minor versions behind support. Upgrading means reading changelogs for removed APIs and testing every manifest. Each month deferred makes it worse, and the person who understood the setup has gone.
Meanwhile the cost side compounds quietly: unattached EBS volumes, old snapshots, a NAT gateway per availability zone in an account nobody reviews, and non-production environments running through the weekend. None of it is dramatic. It’s typically 20–30% of a bill that nobody has read line by line in a year.
Mid-level (3–5 years). Works within an established platform. Comfortable with pipelines, containers, and cloud services. Needs direction on architecture.
Senior (5–8 years). Designs the platform. Makes cloud, orchestration, and tooling decisions and can justify them on cost and operational grounds. Handles incidents without escalating.
Lead / Platform (8+ years). Owns platform direction, sets standards across teams, makes the build-versus-managed-service call.
Cloud migration. Moving off on-premise or between providers. Needs someone who has done one, not someone learning on your production systems. Where the open question is what the target should look like rather than how to get there, that is a cloud architect.
CI/CD from scratch or rescue. Teams shipping manually, or with a pipeline that’s become unreliable.
Kubernetes adoption. Often the right answer, sometimes not. We’d rather have someone tell you honestly than sell you complexity.
Cost reduction. A bill that’s grown faster than usage. Usually a focused one-to-two month engagement with clear ROI.
On-call and reliability. Getting from “it breaks and we find out from customers” to actual monitoring. Related: cloud and DevOps services.
Week one. They map what exists before changing any of it. Expect them to ask who has production access, where state is stored, whether there’s a staging environment that actually resembles production, and what happened during the last incident. A genuinely good hire produces a short written summary of the current setup in week one — that document is the first evidence they intend to hand over rather than accumulate knowledge.
Weeks two to four. One meaningful, reversible improvement. Remote state with locking, a pipeline that fails fast on a bad build, alerting that pages on symptoms rather than causes. It should ship with a runbook. Expect them to also name the biggest risk they found and what they’d do about it, with a cost estimate.
The warning sign is new tooling in week two. An engineer who introduces a service mesh, a new observability vendor, and a different CI system before they’ve had an incident is optimising for their own familiarity, not your reliability. Good ones are conservative early and opinionated later.
The first thing worth asking for is a rebuild. Not a migration or a redesign — just a demonstration that your environment can be recreated from what is in version control. Teams routinely believe they can do this and find, when someone finally tries, that three manual steps live only in one person’s memory.
Adopting Kubernetes for three services. It’s the default answer and often the wrong one. Below roughly ten services and a team that can carry an on-call rota, managed container platforms — ECS, Cloud Run, App Service — do the same job with a fraction of the operational surface. Ask a candidate when they would not use Kubernetes. If they have no answer, that tells you plenty.
Hiring cloud-agnostic and expecting productivity. The concepts transfer; the operational detail does not. IAM alone takes weeks to hold properly. Someone excellent on AWS will be useful on Azure in a month or two, not on day one. Match on your actual provider and treat multi-cloud claims sceptically.
Leaving handover to the final week. Documentation written at the end is written from memory, by someone already thinking about their next engagement. Make it a deliverable per change, not a phase. It’s the single cheapest thing you can insist on and almost nobody does it.
Buying tools instead of practices. A pipeline, a dashboard, and an alerting stack are all straightforward to install and none of them changes anything on its own. What matters is whether alerts route to someone who can act, whether a failed build actually blocks a merge, and whether anyone reads the dashboard on an ordinary day. Tooling without those habits is cost with the appearance of maturity.
Rates depend on seniority, cloud platform, and engagement length. Tell us the role and we’ll give you a firm number.
Tell us the role and we will send a shortlist within 72 hours.
Tell Us the RoleTell us the role. Three to five profiles within 72 hours.