30-60-90 day plan for data scientists
On this page
- What makes a data scientist's ramp different
- The 30-60-90 day plan
- Reproduction: the first-month assignment that pays off
- Baselines and evaluation plans
- What "on track" looks like at each checkpoint
- A filled example
- What the hiring manager owes the new data scientist
- Mistakes that stall a new data scientist
- Adapting the plan
- Questions people ask
A new data scientist can spend three months building a model that scores well in a notebook and never changes anything in the business. The model might depend on a feature that is not available at prediction time, it might beat nothing simpler than itself, or nobody may have agreed what decision it was meant to support. A 30-60-90 day plan for a data scientist should guard against each of those, by making the hire prove they understand the existing systems, insist on a baseline, and take one piece of work all the way to a decision.
This page is for the data science lead, head of analytics or engineering manager writing the plan. If the role is closer to reporting and dashboards, use the 30-60-90 day plan for data analysts instead. For the general version across roles, see the 30-60-90 day plan template for new hires.
What makes a data scientist's ramp different
Data science work has a long distance between "it runs" and "it matters." In between sit data pipelines, feature definitions, training and serving environments, evaluation design, experiment platforms, monitoring and the stakeholder who has to trust the output. A new hire can be strong in modeling and still stall on any of those steps for weeks.
The plan below makes those steps the structure of the ramp. Month one is about reproducing what already exists. Month two is about a new, scoped problem with a baseline and an evaluation plan written before modeling starts. Month three is about getting a result in front of a decision, whether that is a production launch, an online experiment or a documented "not worth it."
The 30-60-90 day plan
30-60-90 day plan — [Name], Data Scientist, [team / problem area]
Manager: [name] Start date: [date] Reviewer: [senior DS / MLE]
Stack: [warehouse] [feature store, if any] [ML platform, e.g. MLflow /
SageMaker / Vertex AI / Databricks] [experiment platform]
Business partner: [name, team]
DAYS 1-30 — Reproduce what exists
Goals:
- Access to the warehouse, repositories, ML platform and experiment
tool by end of week one
- Reproduce one existing model's training run and evaluation metrics
from the repository; document every gap or manual step
- Read the write-ups of the last [3-5] experiments; re-run the
analysis of one and check whether its conclusion holds
- Map how a prediction reaches the product or the business user:
batch job, API, dashboard or spreadsheet
- Meet the business partners and ask which decisions they would
make differently with better predictions
Deliverables by day 30:
- A reproduction report: matched numbers or an explanation of why not
- A one-page problem brief for the day-60 project, with the
decision it supports and how success will be judged
Check-in: day 30, with manager and reviewer
DAYS 31-60 — Build against a baseline
Goals:
- Write the evaluation plan before modeling: metric, validation
split, leakage checks, and what result would be good enough
- Build the simplest reasonable baseline first (a rule, a heuristic,
or a simple model) and record its score
- Iterate on a model only as far as it clearly beats the baseline
- Get code and evaluation reviewed before results are shared
Deliverables by day 60:
- An offline evaluation write-up: baseline vs model, errors by
segment, known risks
- A plan for the next step: online test, shadow deployment, or stop
Check-in: day 60
DAYS 61-90 — Take it to a decision
Goals:
- Run the agreed next step: an A/B test, a shadow run, or a pilot
with the business team
- Define monitoring for anything deployed: input drift, prediction
distribution, and the business metric it should move
- Present results and a recommendation to the business partner
Deliverables by day 90:
- A decision made and documented: launch, iterate, or stop
- Handover notes good enough for a teammate to rerun the work
Check-in: day 90 — full review against this plan
Reproduction: the first-month assignment that pays off
Reproducing an existing model sounds like busywork. In practice it is the fastest way for a new data scientist to learn the data, the code conventions, the training environment and the deployment path, and it regularly turns up problems: a training query that silently changed, a random seed nobody fixed, a feature computed differently in training than in serving, or an evaluation that leaked future data. A new hire who finds one of those in month one has already earned their place.
Ask for the reproduction report to cover the steps taken, which numbers matched, which did not and why, and every manual step that is not in the code. Then turn the manual steps into tickets.
Baselines and evaluation plans
The habit that most separates useful data scientists from impressive ones is comparing against something simple. If a rule such as "flag accounts with no login in 30 days" catches most of the churn a model catches, the model may not be worth its maintenance cost. Make the baseline a required part of the day-60 deliverable, with the same evaluation applied to both.
The evaluation plan should be written and reviewed before the modeling starts, for the same reason a measurement plan comes before a product launch: it stops the hire from choosing the metric that makes the result look best after the fact. It should name the metric and why it matches the business decision, how data is split (by time, if the model will predict the future), how leakage is checked and what result would count as not worth pursuing.
What "on track" looks like at each checkpoint
| Checkpoint | On track | Worth a direct conversation |
|---|---|---|
| Day 30 | The existing model has been reproduced or the reasons it cannot be are documented; the problem brief names a decision and a business owner | The hire has started a new model without understanding the existing one; the problem brief is a technique looking for a use |
| Day 60 | The evaluation plan predates the results; a baseline exists; errors are broken down by segment; code has been reviewed | Only one metric is reported, chosen afterwards; no baseline; results shared before review |
| Day 90 | A decision has been made with the business partner, including "stop" if warranted; monitoring is defined for anything live | The work is "almost ready" with no date; nobody outside the data team has seen the results |
A filled example
Data scientist: Amara Nwosu (invented), joining a two-person data science team at an online grocery company.
Day 30: Reproduced the weekly demand forecast and got close but not identical numbers; the gap traced to a holiday calendar table maintained by hand in a spreadsheet, which she moved into the warehouse. Her problem brief proposed predicting which substitution offers customers accept, to support the picking team's substitution choices.
Day 60: The evaluation plan used a time-based split and acceptance rate on held-out weeks. The baseline, "offer the same brand in the next size," already performed reasonably; a gradient-boosted model beat it by a modest margin, with the biggest gain in fresh produce and almost none in household goods.
Day 90: Ran a two-week pilot in one warehouse, model for produce and baseline elsewhere. Acceptance improved in produce compared with matched warehouses; she recommended a wider rollout for produce only, and documented why the model was not worth maintaining for other categories.
What the hiring manager owes the new data scientist
- Working access in week one, including compute and the ML platform. Data scientists waiting on permissions burn weeks quietly.
- A named reviewer for code, evaluation designs and experiment plans.
- A business partner with a real decision. Without one, the day-90 goal has nowhere to land.
- A path to production, or honesty that one does not exist yet, in which case the day-90 bar should be a pilot or a decision, not a deployment.
Mistakes that stall a new data scientist
| Mistake | What it looks like | Fix |
|---|---|---|
| Notebook-only work | Results that nobody else can rerun | Require code in the repository and handover notes by day 90 |
| No baseline | A complex model with no evidence it beats something simple | Make the baseline a day-60 deliverable |
| Metric chosen afterwards | The reported metric is whichever looked best | Review the evaluation plan before modeling starts |
| Leakage | Unusually strong offline results that collapse live | Time-based splits and a leakage checklist in review |
| Problem without an owner | Interesting work nobody asked for | Name the business decision and owner in the day-30 brief |
Adapting the plan
- Product or experimentation-focused data scientists: replace the model with an experiment program: audit past A/B tests, design one test with a power calculation reviewed by a peer, and run it to a decision.
- Machine learning engineers: weight the plan toward serving, monitoring and pipelines; the machine learning engineer screening questions cover the skills that matter there.
- Senior hires: add a review of the team's modeling standards and one improvement to them, such as a shared evaluation template.
If you are still interviewing, the data scientist phone screen questions probe for the same habits this plan rewards: framing a problem around a decision, starting from a baseline and being honest about what a result does and does not show.
Questions people ask
How is a data scientist's 30-60-90 plan different from a data analyst's?
An analyst's plan centers on trusted metrics and analyses that inform decisions. A data scientist's plan centers on models and experiments: reproducing an existing model, building against a simple baseline, evaluating offline and online, and getting a result into production or into a documented decision.
Should a new data scientist ship a model to production in 90 days?
Sometimes, if the problem is well scoped and the deployment path already exists. More often, a realistic day-90 bar is a model that beats a simple baseline offline with an agreed plan to test it live, or an experiment that reached a clear decision. Forcing a production launch to hit a date tends to skip evaluation steps that matter.
What is the best first assignment for a new data scientist?
Reproduce an existing model's training run and evaluation numbers from the repository. It exposes the data pipelines, feature definitions, environment problems and undocumented steps faster than anything else, and it often finds real issues the team did not know about.
Who should review a new data scientist's work?
A senior data scientist or machine learning engineer should review code, evaluation design and experiment plans for at least the first 60 days. A stakeholder from the business should review the problem framing, so the model answers the question the business actually has.