A coding take-home is easier to evaluate consistently when candidates and reviewers share more than a prompt: give them a machine-checkable rubric, a deliberately flawed sample solution, and a short explanation of its failures. Morgan Zhou’s proposed four-file packet makes expectations visible—but it is a practical design proposal, not a proven way to improve hiring outcomes.
What “hand them a wrong answer” means
Morgan Zhou’s September 21, 2026 DEV Community article proposes publishing a known-bad solution alongside a take-home assignment. The point is not to trick candidates or make them repair someone else’s code. It is to give both sides a shared calibration point: candidates can see what fails the stated contract, and reviewers can check their grading against the same example.
The packet has four parts:
- Candidate prompt: the task and its explicit requirements.
- Machine-checkable rubric: executable checks for required behavior.
- Known-bad sample: an implementation that fails in documented ways.
- Failure catalog: a short explanation linking the sample’s defects to the requirements.
This shifts the exercise away from guessing what an interviewer privately considers “good.” It does not remove judgment from hiring; it makes a defined slice of that judgment inspectable.
How the example assignment works
Zhou’s example asks a candidate to build a local HTTP service on port 8080. It accepts a JSON request at POST /review with the fields diff, tests_passed, tests_failed, and secrets_hit. The response includes score, verdict (reject, revise, or pass), reasons, and beats_sample.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitches#1 Best Overall
- Careercup, Easy To Read
- Condition : Good
- Compact for travelling
The important part is that the requirements are observable rather than aspirational:
- A failed test set must never receive a passing verdict.
- If
secrets_hitis true, the score is capped at 20 and the verdict must be rejection. - Reasons must identify a concrete signal in the submitted payload rather than offer vague boilerplate.
- The implementation is compared with the published sample under the stated scoring contract.
The prompt also asks for a grade_receipt.json containing one request and response actually run. That receipt makes it easier to see what the candidate’s program returned for a real input, rather than relying only on an assertion that it worked.
What the bad sample demonstrates
The deliberately wrong implementation always returns score 100, verdict pass, and a vague reason. It therefore violates the failed-test rule, the secrets cap and rejection rule, and the requirement for specific evidence in its reasons. The direction sample illustrates the contrast: it applies the stated caps and gives concrete reasons for failed tests or a secret flag.
Rank #2
These are illustrative examples from the article, not independently verified code results. The useful design principle is to make the failure legible: each defect should map to a published requirement that a candidate or reviewer can check.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Build the rubric around a small, real contract
A useful rubric tests behavior that matters to the job, not an arbitrary collection of clever edge cases. In the example, the invariants are simple enough to explain and strict enough to check automatically. A reviewer can ask whether a response obeys the contract without needing to infer the candidate’s intent from style or confidence.
For a similar exercise, write each criterion so it can be answered with evidence:
- What input triggers the requirement?
- What exact output or behavior is required?
- What observable result proves the requirement was met?
- What should happen when a required condition is absent or violated?
Keep public checks aligned with the published rubric. If hidden checks or additional judgment are used, candidates should not be led to believe that the public checks are the entire scoring system. A known-bad sample should be run by reviewers too; otherwise it risks being decorative rather than a calibration tool.
Run the grader against the service, not just a substitute
Zhou recommends exercising the grader against a live local process with the same host, timeout, and payload bytes used for evaluation. That matters because a test can pass against a mock or helper function while the actual HTTP service fails at the boundary: the route, request parsing, response encoding, or process startup.
Before sending the assignment, run the grader on both the known-bad sample and a direction sample. Confirm that the grader rejects the bad behavior for the intended reasons, and that the receipt reflects a request and response from a process that was actually run. Use the same local conditions candidates are told to use; avoid requiring a paid API call or account that is not part of the skill being evaluated.
Keep the exercise fair and job-related
The U.S. Office of Personnel Management defines work-sample tests as tasks that mirror activities employees perform. Its guidance says these tests are most appropriate when the measured competencies are critical and expected at entry. If the organization plans to teach a skill after hiring, screening candidates for prior mastery of it may not be appropriate. See OPM’s work-sample test guidance.
That gives employers a practical test for whether to use a take-home at all: would a person in this role need the measured ability on day one, and does the exercise resemble the relevant work? A compact service-contract task may assess implementation against explicit behavior. It does not stand in for system design on a multi-region billing platform unless that is genuinely the work being assessed.
Scope also affects fairness and accessibility. Zhou advises keeping the prompt short and avoiding Kubernetes, dashboards, paid vendor logins, and paid API calls; the task should be possible with a free model and a free local machine. The article also cautions against unpaid weekend assignments, collecting candidate code when the organization cannot accept it, and requirements for a GPU, private dataset, or production credentials. Those constraints are not incidental: they can screen for time, money, equipment, or access rather than the intended competency.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
Standardize evaluation without overstating what it proves
OPM’s general assessment-strategy guidance reports validity estimates of 0.54 for work-sample tests and 0.51 for structured interviews; the page does not state a year for those figures. OPM describes validity as the relationship between assessment performance and job performance. These general estimates are not evidence that Zhou’s four-file packet predicts job performance, is fair in a particular hiring process, or improves selection accuracy. See OPM’s assessment-strategy guidance.
Structured interviews offer a useful comparison: OPM describes them as using standardized questions and common rating standards, which can give candidates equal opportunities to provide information and support consistent assessment. A coding exercise can borrow that discipline by using the same prompt, checks, and rating standards for everyone, while still leaving room to evaluate job-relevant qualities the automated rubric cannot capture. See OPM’s structured-interview guidance.
The evidence available for Zhou’s proposal does not establish candidate outcomes or an improvement in hiring accuracy. Treat the packet as a transparency and consistency technique to consider, not as a validated selection instrument. Its strongest case is practical: it exposes the contract, gives reviewers a reference failure, and makes a narrow portion of scoring reproducible.
When this approach is a good fit
A known-bad sample is most useful when a take-home has a small, clear behavior contract and reviewers need a shared basis for checking it. It is less useful when the real job requirement is open-ended work that cannot fairly be reduced to the published checks, or when the assignment’s burdens outweigh what it measures.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Quick Recap
- Use it when: the assessed competency is necessary at entry, the task resembles actual work, and meaningful requirements can be checked consistently.
- Redesign it when: the prompt hides expectations, setup requirements dominate the work, or reviewers cannot explain how a failure relates to the job.
- Do not use it as a decoy: public checks should not imply one scoring standard while candidates are secretly rescored against unrelated criteria.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




