Templates

30-60-90 day plan for data scientists

On this page
  1. What makes a data scientist's ramp different
  2. The 30-60-90 day plan
  3. Reproduction: the first-month assignment that pays off
  4. Baselines and evaluation plans
  5. What "on track" looks like at each checkpoint
  6. A filled example
  7. What the hiring manager owes the new data scientist
  8. Mistakes that stall a new data scientist
  9. Adapting the plan
  10. Questions people ask

A new data scientist can spend three months building a model that scores well in a notebook and never changes anything in the business. The model might depend on a feature that is not available at prediction time, it might beat nothing simpler than itself, or nobody may have agreed what decision it was meant to support. A 30-60-90 day plan for a data scientist should guard against each of those, by making the hire prove they understand the existing systems, insist on a baseline, and take one piece of work all the way to a decision.

This page is for the data science lead, head of analytics or engineering manager writing the plan. If the role is closer to reporting and dashboards, use the 30-60-90 day plan for data analysts instead. For the general version across roles, see the 30-60-90 day plan template for new hires.

What makes a data scientist's ramp different

Data science work has a long distance between "it runs" and "it matters." In between sit data pipelines, feature definitions, training and serving environments, evaluation design, experiment platforms, monitoring and the stakeholder who has to trust the output. A new hire can be strong in modeling and still stall on any of those steps for weeks.

The plan below makes those steps the structure of the ramp. Month one is about reproducing what already exists. Month two is about a new, scoped problem with a baseline and an evaluation plan written before modeling starts. Month three is about getting a result in front of a decision, whether that is a production launch, an online experiment or a documented "not worth it."

The 30-60-90 day plan

30-60-90 day plan — [Name], Data Scientist, [team / problem area]
Manager: [name]    Start date: [date]    Reviewer: [senior DS / MLE]
Stack: [warehouse]  [feature store, if any]  [ML platform, e.g. MLflow /
       SageMaker / Vertex AI / Databricks]  [experiment platform]
Business partner: [name, team]

DAYS 1-30 — Reproduce what exists
Goals:
- Access to the warehouse, repositories, ML platform and experiment
  tool by end of week one
- Reproduce one existing model's training run and evaluation metrics
  from the repository; document every gap or manual step
- Read the write-ups of the last [3-5] experiments; re-run the
  analysis of one and check whether its conclusion holds
- Map how a prediction reaches the product or the business user:
  batch job, API, dashboard or spreadsheet
- Meet the business partners and ask which decisions they would
  make differently with better predictions
Deliverables by day 30:
- A reproduction report: matched numbers or an explanation of why not
- A one-page problem brief for the day-60 project, with the
  decision it supports and how success will be judged
Check-in: day 30, with manager and reviewer

DAYS 31-60 — Build against a baseline
Goals:
- Write the evaluation plan before modeling: metric, validation
  split, leakage checks, and what result would be good enough
- Build the simplest reasonable baseline first (a rule, a heuristic,
  or a simple model) and record its score
- Iterate on a model only as far as it clearly beats the baseline
- Get code and evaluation reviewed before results are shared
Deliverables by day 60:
- An offline evaluation write-up: baseline vs model, errors by
  segment, known risks
- A plan for the next step: online test, shadow deployment, or stop
Check-in: day 60

DAYS 61-90 — Take it to a decision
Goals:
- Run the agreed next step: an A/B test, a shadow run, or a pilot
  with the business team
- Define monitoring for anything deployed: input drift, prediction
  distribution, and the business metric it should move
- Present results and a recommendation to the business partner
Deliverables by day 90:
- A decision made and documented: launch, iterate, or stop
- Handover notes good enough for a teammate to rerun the work
Check-in: day 90 — full review against this plan

Reproduction: the first-month assignment that pays off

Reproducing an existing model sounds like busywork. In practice it is the fastest way for a new data scientist to learn the data, the code conventions, the training environment and the deployment path, and it regularly turns up problems: a training query that silently changed, a random seed nobody fixed, a feature computed differently in training than in serving, or an evaluation that leaked future data. A new hire who finds one of those in month one has already earned their place.

Ask for the reproduction report to cover the steps taken, which numbers matched, which did not and why, and every manual step that is not in the code. Then turn the manual steps into tickets.

Baselines and evaluation plans

The habit that most separates useful data scientists from impressive ones is comparing against something simple. If a rule such as "flag accounts with no login in 30 days" catches most of the churn a model catches, the model may not be worth its maintenance cost. Make the baseline a required part of the day-60 deliverable, with the same evaluation applied to both.

The evaluation plan should be written and reviewed before the modeling starts, for the same reason a measurement plan comes before a product launch: it stops the hire from choosing the metric that makes the result look best after the fact. It should name the metric and why it matches the business decision, how data is split (by time, if the model will predict the future), how leakage is checked and what result would count as not worth pursuing.

What "on track" looks like at each checkpoint

CheckpointOn trackWorth a direct conversation
Day 30The existing model has been reproduced or the reasons it cannot be are documented; the problem brief names a decision and a business ownerThe hire has started a new model without understanding the existing one; the problem brief is a technique looking for a use
Day 60The evaluation plan predates the results; a baseline exists; errors are broken down by segment; code has been reviewedOnly one metric is reported, chosen afterwards; no baseline; results shared before review
Day 90A decision has been made with the business partner, including "stop" if warranted; monitoring is defined for anything liveThe work is "almost ready" with no date; nobody outside the data team has seen the results

A filled example

Data scientist: Amara Nwosu (invented), joining a two-person data science team at an online grocery company.

Day 30: Reproduced the weekly demand forecast and got close but not identical numbers; the gap traced to a holiday calendar table maintained by hand in a spreadsheet, which she moved into the warehouse. Her problem brief proposed predicting which substitution offers customers accept, to support the picking team's substitution choices.

Day 60: The evaluation plan used a time-based split and acceptance rate on held-out weeks. The baseline, "offer the same brand in the next size," already performed reasonably; a gradient-boosted model beat it by a modest margin, with the biggest gain in fresh produce and almost none in household goods.

Day 90: Ran a two-week pilot in one warehouse, model for produce and baseline elsewhere. Acceptance improved in produce compared with matched warehouses; she recommended a wider rollout for produce only, and documented why the model was not worth maintaining for other categories.

What the hiring manager owes the new data scientist

  • Working access in week one, including compute and the ML platform. Data scientists waiting on permissions burn weeks quietly.
  • A named reviewer for code, evaluation designs and experiment plans.
  • A business partner with a real decision. Without one, the day-90 goal has nowhere to land.
  • A path to production, or honesty that one does not exist yet, in which case the day-90 bar should be a pilot or a decision, not a deployment.

Mistakes that stall a new data scientist

MistakeWhat it looks likeFix
Notebook-only workResults that nobody else can rerunRequire code in the repository and handover notes by day 90
No baselineA complex model with no evidence it beats something simpleMake the baseline a day-60 deliverable
Metric chosen afterwardsThe reported metric is whichever looked bestReview the evaluation plan before modeling starts
LeakageUnusually strong offline results that collapse liveTime-based splits and a leakage checklist in review
Problem without an ownerInteresting work nobody asked forName the business decision and owner in the day-30 brief

Adapting the plan

  • Product or experimentation-focused data scientists: replace the model with an experiment program: audit past A/B tests, design one test with a power calculation reviewed by a peer, and run it to a decision.
  • Machine learning engineers: weight the plan toward serving, monitoring and pipelines; the machine learning engineer screening questions cover the skills that matter there.
  • Senior hires: add a review of the team's modeling standards and one improvement to them, such as a shared evaluation template.

If you are still interviewing, the data scientist phone screen questions probe for the same habits this plan rewards: framing a problem around a decision, starting from a baseline and being honest about what a result does and does not show.

Questions people ask

How is a data scientist's 30-60-90 plan different from a data analyst's?

An analyst's plan centers on trusted metrics and analyses that inform decisions. A data scientist's plan centers on models and experiments: reproducing an existing model, building against a simple baseline, evaluating offline and online, and getting a result into production or into a documented decision.

Should a new data scientist ship a model to production in 90 days?

Sometimes, if the problem is well scoped and the deployment path already exists. More often, a realistic day-90 bar is a model that beats a simple baseline offline with an agreed plan to test it live, or an experiment that reached a clear decision. Forcing a production launch to hit a date tends to skip evaluation steps that matter.

What is the best first assignment for a new data scientist?

Reproduce an existing model's training run and evaluation numbers from the repository. It exposes the data pipelines, feature definitions, environment problems and undocumented steps faster than anything else, and it often finds real issues the team did not know about.

Who should review a new data scientist's work?

A senior data scientist or machine learning engineer should review code, evaluation design and experiment plans for at least the first 60 days. A stakeholder from the business should review the problem framing, so the model answers the question the business actually has.