October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk6 min

Andrej Karpathy Says Nanochat Was Basically Hand-Written—Why AI Agents Struggled

Karpathy’s mostly hand-written Nanochat is not a rejection of AI coding. It shows why agents struggle with novel, technically dense codebases where correctness requires expert judgment and expensive testing.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Andrej Karpathy, the machine-learning researcher widely credited with coining “vibe coding,” says his Nanochat project was “basically entirely hand-written.” He tried Claude and Codex agents several times, but said they “didn’t work well enough” and were ultimately “net unhelpful,” possibly because the repository was too far outside the models’ usual data distribution. Futurism reported the comments on October 20, 2025.

That is not a declaration that AI coding is useless or that Karpathy has abandoned vibe coding. It is a narrower warning: agents are much less dependable when software is unusual, technically dense, difficult to test and costly to evaluate.

Who is Andrej Karpathy?

Karpathy is a prominent machine-learning educator and researcher, a former Tesla AI leader and former OpenAI executive and co-founder. He is widely credited with introducing the phrase “vibe coding” in a February 2025 post. “Coined the term” is more precise than “invented AI programming”: natural-language coding and AI-assisted development existed before the phrase.

In its original sense, vibe coding meant describing a desired program in ordinary language, accepting generated code without necessarily understanding every line, running it, pasting errors back into the model and iterating. Karpathy presented it mainly as a playful approach for prototypes and “throwaway weekend projects,” not as a replacement for professional engineering. Futurism’s account quotes that caveat.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What Nanochat is

Nanochat is not merely a front end that sends prompts to a hosted chatbot. It is an open-source, from-scratch project for training a small language model and interacting with it through a ChatGPT-like web interface. The official repository describes it as “the best ChatGPT that $100 can buy.” That phrase is project positioning, not a guaranteed all-in cost for every user or hardware setup.

The repository includes training, evaluation, inference and web-serving components. Much of the implementation uses relatively vanilla PyTorch. Transformer depth is the principal complexity control, with other model settings derived from it, so users can experiment with different scales without managing a large collection of independent architectural knobs.

Nanochat can run on hardware supported by PyTorch, including CUDA systems and Apple Silicon through MPS, although the project notes that not every hardware path has been personally exercised. CPU or reduced MPS operation is possible in limited form; stronger results require more capable hardware. Training time and cost depend on the exact model configuration, GPU, region, storage and whether rented hardware is on-demand or discounted.

What Karpathy said about AI assistance

The reported facts are specific:

  • Nanochat was “basically entirely hand-written.”
  • Karpathy tried Claude and Codex agents.
  • The attempts did not work well enough and were “net unhelpful.”
  • He suggested that the repository might be too far outside the models’ data distribution.

This does not establish that every line was typed manually without any AI use. It also does not mean the agents could write none of the code, that Nanochat was impossible to build with assistance, or that Karpathy has rejected AI coding permanently. The contemporary report is available from Tech Yahoo and the original coverage from Futurism.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why a project like Nanochat can defeat ordinary agent workflows

Karpathy did not publish a formal postmortem identifying one failure mechanism. The following are reasonable interpretations of the project’s characteristics, not confirmed findings from him.

An unusual codebase

Popular web applications contain familiar frameworks, conventions and examples. A from-scratch language-model training stack has fewer direct analogues. An agent may recognize individual PyTorch operations while misunderstanding the project’s particular assumptions about data, tensors, checkpoints or distributed execution.

Dense domain knowledge

Training infrastructure combines numerical computing, optimizer behavior, memory limits, hardware-specific performance and statistical evaluation. A locally plausible edit can alter gradient behavior, numerical stability or throughput without producing an obvious exception.

Cross-file and nonlocal correctness

A change to data loading can affect tokenization, batching, training metrics and reproducibility. A checkpointing change can surface later during evaluation or resumption. The relevant invariant may span many files, while the agent tends to propose a patch around the line named in the prompt.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Delayed and expensive feedback

A web-app defect may appear immediately in a browser. A training change may require hours of GPU time before its effect on loss, model quality or stability is visible. Repeated agent-driven experiments can therefore spend money while producing ambiguous evidence.

Code that runs can still be wrong

Compilation and a successful launch are weak tests for machine-learning infrastructure. A run can complete while silently reducing quality, increasing memory use, damaging reproducibility or changing the scientific meaning of an experiment.

What “outside the data distribution” means

A model is generally more reliable when a task resembles patterns in its training and reinforcement-learning experience. “Outside the data distribution” does not mean it has never seen PyTorch or language-model code. It means the exact combination of architecture, conventions, constraints and desired behavior may be unfamiliar enough to reduce reliability.

That explanation is plausible rather than a proof. Karpathy’s later writing describes a similar unevenness: models can perform impressively in some verifiable coding environments yet fail unpredictably on particular reasoning problems. A novel repository can therefore expose a model’s weak spot even when the same model is excellent at routine refactors.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Nanochat is almost the opposite of classic vibe coding

Vibe coding works best when speed matters more than deep understanding and the result is easy to discard. Nanochat is intended to expose and control the mechanics of model training. Its success cannot be judged only by whether a page loads or a command exits without an error.

Relevant checks include:

  • Does training remain numerically stable?
  • Does the model reach the expected quality for the compute used?
  • Are memory use, latency and throughput acceptable?
  • Can another person reproduce the result?
  • Do evaluation and checkpointing still mean what the author intended?

Those requirements make expert specification, measurement and review central. They do not create a binary choice between writing every line yourself and delegating everything to an agent.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Karpathy’s later view: from vibe coding to agentic engineering

In his 2026 Sequoia Ascent summary, Karpathy draws a distinction between two uses of AI. Vibe coding lowers the floor: more people can create working software. “Agentic engineering” raises the ceiling by using agents within professional standards for correctness, security, maintainability and taste.

He says agents became noticeably more useful around December 2025 in his experience, producing larger, more coherent and more reliable changes. His position is not that humans should stop using them. Humans still own the specification, architecture, judgment, security decisions and final oversight.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When to delegate and when to take over

Task or environment AI assistance is usually a good fit Close expert supervision or manual implementation is preferable
Familiar application work Boilerplate, CRUD screens, API integrations, documentation and migration scripts Changes that cross undocumented system boundaries
Testing and maintenance Drafting tests, updating repetitive fixtures and explaining existing code Tests that could merely encode an existing implementation mistake
Technical risk Low-stakes prototypes and personal tools Security-sensitive, financial, medical or safety-critical systems
Specialized engineering Scaffolding and narrow, independently verifiable utilities Machine-learning infrastructure, distributed systems, performance kernels and novel algorithms

A practical workflow

  1. Specify the boundary. State the invariant, interfaces, acceptance tests and non-goals before asking an agent to edit code.
  2. Delegate reversible work first. Documentation, scaffolding, routine refactors and test drafts are easier to inspect and undo.
  3. Validate independently. Run tests and measure domain outcomes such as model quality, latency, memory and reproducibility; do not stop at “it compiles.”
  4. Review sensitive changes manually. Check credentials, data handling, dependencies, permissions and generated logs.
  5. Stop unproductive loops. If the agent repeatedly patches symptoms, contradicts repository conventions or makes destructive edits, switch to a manual diagnosis.

Failure modes worth watching

  • A locally reasonable edit breaks a cross-file invariant.
  • Generated tests confirm the agent’s misunderstanding instead of the intended behavior.
  • Dependency or API drift makes suggested commands invalid.
  • Secrets, personal data or unsafe defaults enter code or logs.
  • Generated code introduces licensing or attribution questions.
  • Developers accept code they cannot later debug or maintain.
  • Repeated cloud-GPU experiments accumulate costs before a meaningful result is obtained.
  • Inconsistent generated styles make the repository harder to evolve.

These risks do not prove that AI-generated code is always insecure or slower. They explain why claims about universal replacement for programmers go beyond what the Nanochat episode demonstrates. Claude Code, Codex, editor-based tools such as Cursor and assistants such as GitHub Copilot can all be useful, but no product guarantees success on an unfamiliar research codebase.

The Bottom Line

Karpathy’s Nanochat experience is not an anti-AI reversal. It is evidence that coding agents are highly sensitive to familiarity, specification and verifiability. Use them aggressively for routine, testable work; keep humans responsible for architecture and security; and write or closely supervise the specialized core when correctness is difficult to observe.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. Shenzhen desk3 min
    HONOR Expands Beyond Smartphones With Humanoid Robot RevealHONOR said it unveiled its first humanoid robot at MWC 2026 and named shopping assistance, workplace inspections, and supportive companionship as intended uses. Later Robotics D1 claims and a reported…
  2. Cupertino desk5 min
    Apple Unveils AirPods Max 2: The Upgrade That Should Have Happened Years AgoAirPods Max 2 adds H2-powered audio features and Apple claims up to 1.5× more effective ANC, but its design, Smart Case, and 20-hour battery rating are unchanged. Wired lossless audio…
  3. Cupertino desk4 min
    Apple’s OLED Touch MacBooks Are Coming—but the Dynamic Island Is the Real GambleApple has not announced an OLED touchscreen MacBook, but reports point to high-end models arriving in late 2026 or early 2027. The reported Mac Dynamic Island could be useful, but…
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.