A shortlist in 72 hours, contributing inside two weeks, and a 30-day guarantee if the fit is wrong.
72-hour shortlist · 30-day replacement guarantee · Month-to-month
More briefs go wrong at this title than at any other in data. "Data scientist" is used for three jobs that need different people.
There is the analyst-leaning version: framing questions, designing experiments, and producing analysis that changes a decision. There is the modelling version: building predictive models, validating them honestly, and knowing when the answer is that the signal is not there. And there is the engineering version, which is not a data scientist at all — it is an ML engineer or an MLOps engineer, and it is what most teams actually mean when they say they want a model in production.
The classic mismatch is hiring a strong researcher and expecting a deployed service. They produce a well-validated model in a notebook. Nobody can serve it, monitor it, or retrain it, and six months later there is a good piece of work that changed nothing.
Decide which of the three you need before writing the brief. It is the single highest-leverage thing you can do in this hire.
We screen for the things that make analysis trustworthy, which is not the same as making it impressive.
Whether they establish what decision changes before opening the data. The difference between analysis and activity.
Spotting features that would not exist at prediction time. The most common cause of a model that validates well and fails in use.
Sample size, stopping rules, and why checking a test daily until it looks significant is not a result.
Choosing metrics that match the business cost of each error, and always comparing against a baseline worth beating.
Willingness to report that a rule performs as well, or that the data cannot answer the question as posed.
A live conversation in English, every time. Analysis that cannot be explained to the person acting on it has no effect.
The expensive failure in data science is not a model that performs badly. It is a model that performs beautifully in evaluation and is wrong.
Leakage is the usual mechanism. A feature is included that would not be available at the moment of prediction — a field populated only after the outcome it is predicting, or an identifier that correlates with how the data was collected rather than with anything real. Accuracy is excellent. The model is useless, and because the number is good it goes to production with confidence.
The business version of the same failure is a metric that is technically correct and practically meaningless. A classifier that is 99 per cent accurate on a problem where 99 per cent of cases are negative has learned to say no. Everybody in the room can read the accuracy figure; far fewer will ask about precision and recall on the class anybody cares about.
Experiments fail the same way. A test is stopped early because the result looks significant, having been checked daily for a fortnight — which makes a significant-looking result close to inevitable whether or not the effect is real. The feature ships, the metric does not move, and nobody revisits the analysis that justified it. Both mistakes are honest and both are avoidable by somebody who has been caught by them before.
The compounding cost is credibility. Once a data science function ships two conclusions that do not hold, the next correct one is discounted too.
Data Scientist (2–4 years). Runs analyses and builds models against defined questions. Comfortable with Python, SQL and standard libraries. Needs support framing the problem and interpreting what the result licenses.
Senior (4–7 years). Frames the question with the business, designs experiments that survive scrutiny, validates honestly, and says when the data cannot answer what is being asked.
Lead / Principal (7+ years). Sets the analytical agenda, decides where modelling is worth it and where a rule would do, owns experimentation standards, and defends a null result to stakeholders who wanted a different answer.
Forecasting and demand planning. Common, well-bounded, and usually measurable against the process it replaces.
Churn and propensity modelling. Frequently requested. The modelling is rarely the hard part; deciding what action follows the score is.
Experimentation. Designing and analysing tests properly, including the ones that should not have been run.
Exploratory analysis. Answering a specific business question once, well. Often better value than a model, and undersold because it produces no lasting artefact.
Production modelling. Where the output must be served and monitored. Pair with MLOps engineers rather than expecting one person to cover both.
Week one. They spend it on the problem, not the data. Expect questions about what decision the output will inform, who makes it today, and what they do now instead. A data scientist who cannot say what will change as a result of the work has not been given a question yet, and should say so.
Weeks two to four. A baseline, and an honest comparison against it. The most useful early output is often that a simple rule performs nearly as well as a model, which reframes the whole project. Expect the validation approach to be explained before results are presented, not after.
The warning sign is a first presentation full of impressive numbers and no caveats. Real analysis has limitations, and somebody who has done this for a while volunteers them — the sample is unrepresentative here, this feature would not be available at prediction time, this holds for the last two quarters and may not for the next. Confidence without qualification usually means the qualifications have not been looked for.
Hiring a data scientist to do ML engineering. The commonest mismatch in this field. If what you need is a model served, monitored and retrained, that is an engineering hire. Research skill and production skill overlap far less than the shared vocabulary suggests.
Interviewing on algorithms rather than on judgement. Deriving backpropagation predicts very little about the work. Asking somebody to critique a flawed experimental design, or to spot leakage in a feature list, tells you far more and is much harder to prepare for.
Skipping the baseline. Without a simple comparison — last month's value, a business rule, the current manual process — a model's accuracy figure is uninterpretable. Insist on it, and be prepared for the answer that the baseline is close enough that no model is warranted.
Not deciding what happens with the output. A churn score nobody acts on has no value regardless of quality. Agree the intervention before the modelling starts. Projects that skip this reliably end with a good model and an argument about whose budget acts on it.
Rates depend on specialisation, seniority, and engagement length. Tell us the role and we will give you a firm number.
Tell us the role and we will send a shortlist within 72 hours.
Tell Us the RoleTell us the role. Three to five profiles within 72 hours.