Templates

Interview scorecard for software engineers, with anchors for each competency

On this page
  1. The six competencies this scorecard covers
  2. The scorecard template
  3. Anchors for each competency
  4. A filled example from a live coding round
  5. Calibrating anchors to level
  6. Adjusting for specialization
  7. Mistakes specific to engineering scorecards
  8. Where this fits in a full engineering loop
  9. Questions people ask

A software engineering scorecard has to separate several things that a single "technical skill" rating blurs together: whether a candidate can write correct code, whether they can reason about a system's structure, whether they debug methodically, and whether they can explain their thinking to someone else. Below is a copy-ready scorecard built around six engineering-specific competencies, each with a 1–4 anchored scale, plus a filled example from a live coding round.

Want a printable version you can edit? The scorecard builder starts from a software engineer preset.

This uses the same 1–4 structure as interview scorecard template; what is specific here is the competencies and anchors, written for an individual-contributor software engineering loop rather than a general interview.

The six competencies this scorecard covers

  • Problem decomposition. Does the candidate break an ambiguous problem into smaller, testable pieces before writing code?
  • Coding correctness and clarity. Is the code correct, and would a teammate understand it without an explanation?
  • Debugging method. When something breaks, do they form a hypothesis and test it, or make random changes until it works?
  • System design reasoning. Can they reason about tradeoffs (consistency, latency, cost) rather than naming components from memory?
  • Code review and feedback. Do they give and receive specific, actionable feedback on code, or only surface-level comments?
  • Ownership under production pressure. How do they describe handling an incident or an on-call issue they were responsible for?

The scorecard template

Candidate: [name]                     Role: [title, level]
Interviewer: [name]                   Stage: [coding / system design / behavioral]

Competency: Problem decomposition
Score (1-4): [ ]
Evidence: [how they broke down the problem before coding]

Competency: Coding correctness and clarity
Score (1-4): [ ]
Evidence: [bugs found, readability, whether it compiled/ran]

Competency: Debugging method
Score (1-4): [ ]
Evidence: [hypothesis formed, tools used, time to isolate the issue]

Competency: System design reasoning
Score (1-4): [ ]
Evidence: [tradeoffs discussed, what they asked about scale/constraints]

Competency: Code review and feedback
Score (1-4): [ ]
Evidence: [specificity of feedback given on a sample diff]

Competency: Ownership under production pressure
Score (1-4): [ ]
Evidence: [the incident described, their role, what changed after]

Overall recommendation: [Strong yes / Yes / No / Strong no]
Level calibration: [does the evidence match the level being hired for?]

Anchors for each competency

Problem decomposition

ScoreAnchor
1Starts typing code within the first minute of an ambiguous problem, with no clarifying questions.
2Asks one or two clarifying questions but the overall approach is not stated before coding starts.
3States a plan in two or three steps, checks it with the interviewer, then codes against that plan.
4Identifies an edge case or ambiguity in the prompt itself that the interviewer had not flagged, before coding.

Coding correctness and clarity

ScoreAnchor
1Code does not run, or has a fundamental logic error the candidate does not catch when asked to trace through it.
2Code mostly works but needs the interviewer to point out a bug; naming and structure make intent hard to follow.
3Code is correct for the stated cases, readable without narration, and the candidate can state its time and space complexity.
4Finds and fixes their own bug by testing before being asked, or discusses a cleaner approach after the first working version.

Debugging method

ScoreAnchor
1Changes code at random when something fails, with no stated theory of what might be wrong.
2Forms a guess but does not verify it before moving to the next guess.
3States a hypothesis, checks it with a print statement, log, or targeted test, and narrows down systematically.
4Narrows a bug to its root cause efficiently and explains why the fix addresses the cause, not just the symptom.

System design reasoning

ScoreAnchor
1Names components (a database, a cache, a queue) with no explanation of why each is needed for this problem.
2Describes a workable design but cannot explain what happens under a stated failure or scale increase.
3Explains at least one real tradeoff (consistency versus availability, cost versus latency) relevant to the stated constraints.
4Proposes a design, then revises it after the interviewer changes a constraint, explaining what specifically had to change and why.

Code review and feedback

ScoreAnchor
1Comments only on style (naming, formatting) and misses a functional bug planted in the sample diff.
2Finds the functional issue but the feedback is vague ("this could be cleaner") without a specific suggestion.
3Finds the functional issue and proposes a specific, actionable fix or question for the author.
4Distinguishes between a blocking issue and a preference, and explains the distinction to the interviewer unprompted.

Ownership under production pressure

ScoreAnchor
1Describes an incident entirely in terms of what other people or systems did, with no personal action described.
2Describes their role in resolving an incident but no follow-up or process change afterward.
3Describes diagnosing and resolving an incident they were responsible for, and one concrete change made afterward (a monitor, a runbook update).
4Describes an incident they caused themselves, owns it directly, and describes a systemic change beyond their own individual fix.

A filled example from a live coding round

An invented candidate for a mid-level backend engineering role, given a rate-limiter design problem:

Problem decomposition — 3. Asked whether the limiter needed to work across multiple servers before writing anything, then stated a plan: a token-bucket approach, checked per user, with a note about where persistence would live.

Coding correctness and clarity — 3. Implementation was correct for the stated single-server case; variable names and structure were clear enough that the interviewer did not need to ask what a section did. Did not proactively test an edge case (a burst exactly at the limit), so did not reach a 4.

Debugging method — 4. When a test failed on a boundary condition, said "let me check what happens exactly at the limit" before changing anything, added a print to confirm, then fixed the off-by-one.

System design reasoning — 2. When asked how the design would change across multiple servers, proposed "just use a shared database" without discussing the added latency per request or an alternative like a centralized counter service.

Code review and feedback — 3. Given a sample pull request with a race condition, found it and proposed a specific fix (a lock around the check-and-increment).

Ownership under production pressure — 3. Described being paged for a memory leak they had introduced, finding it with a heap dump, and adding a memory alert afterward.

Overall recommendation: Yes. Strong coding and debugging fundamentals; the system design gap is normal for this level and worth a specific area to probe further in the system design round rather than a reason to pass at the screen stage.

Calibrating anchors to level

The anchors above describe behavior, not years of experience, and the same anchors apply from junior to senior with the scope of the problem adjusted instead of the wording of the anchor.

LevelHow the scope changes
Junior / early careerSmaller problems, more direct hints allowed before scoring a 2 instead of a 1; system design scaled to a single service.
Mid-levelProblems close to the examples above; system design includes at least one explicit tradeoff.
SeniorAmbiguity is intentional and unresolved by the interviewer; a 3 requires identifying tradeoffs the interviewer did not mention.
Staff or aboveAdd a seventh competency: technical influence without direct authority, scored on how they describe getting a team to adopt a decision they could not simply mandate.

Adjusting for specialization

  • Frontend roles: replace system design reasoning with a competency on component architecture and state management tradeoffs; add accessibility awareness as a scored line.
  • Data or machine learning engineering: add a competency on evaluating whether a metric actually measures what it claims to, scored the same way as debugging method: hypothesis, verification, conclusion.
  • Site reliability or infrastructure: weight ownership under production pressure and system design reasoning highest; the coding round can be shorter and more script-oriented than algorithmic.
  • Engineering management candidates: keep coding correctness at a lower weight and add a competency on how they describe giving difficult feedback to a report, scored with the same specificity standard as code review above.

Mistakes specific to engineering scorecards

MistakeWhy it failsFix
Scoring only whether the final code ranMisses process signals that predict performance on the job better than a single correct answerScore decomposition and debugging method as separate lines
System design scored the same way for every levelA junior candidate is set up to fail on a senior-scoped questionScale the scenario per the level table above, keep the anchors
No competency for communicating while codingA candidate who solves the problem silently gives the interviewer nothing to score on collaborationFold it into coding clarity, or add it as a seventh line for roles where pairing matters
Personal style preferences scored as correctnessPenalizes a valid approach that simply is not the interviewer's own habitScore against correctness, readability and complexity, not personal taste
Ownership question skipped for candidates without on-call experienceLoses a real signal about how someone handles being responsible for a mistakeAsk about any project failure they owned, not only a production incident

Where this fits in a full engineering loop

A single scorecard rarely covers all six competencies in one session; most loops split them across a phone screen (problem decomposition, coding correctness), a system design round, and a behavioral round (ownership, code review approach). See software engineer phone screen questions for the screen stage specifically, and use this scorecard's competencies as the columns each stage reports back into, so the hiring committee sees one consistent picture rather than differently-shaped notes from each interviewer.

Questions people ask

Should a coding exercise be scored on whether the solution is optimal?

Score the process at least as heavily as the final answer: whether the candidate clarified requirements, tested their own code, and could explain the time and space complexity of what they wrote. A correct but unexplained solution and a mostly-correct, well-reasoned one are not the same signal.

How do I score system design for a candidate below senior level?

Scale the scenario down rather than skipping the competency: a junior candidate can be asked to design a single service's data model and API rather than a distributed system, scored on the same anchors at a smaller scope.

Is take-home code review a fair substitute for a live coding interview?

It tests a different thing: whether someone can produce working code alone, with time and references, rather than how they think under light pressure and communicate while doing it. Many loops use both, scored as separate competencies rather than one substituting for the other.

Should code style preferences affect the score?

Only if they affect correctness, readability for a team, or maintainability the candidate can explain. Scoring a candidate down for a style choice you happen to dislike, with no team standard behind it, is not a defensible use of the scorecard.