What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Andrej Karpathy, the machine-learning researcher widely credited with coining “vibe coding,” says his Nanochat project was “basically entirely hand-written.” He tried Claude and Codex agents several times, but said they “didn’t work well enough” and were ultimately “net unhelpful,” possibly because the repository was too far outside the models’ usual data distribution. Futurism reported the comments on October 20, 2025.
That is not a declaration that AI coding is useless or that Karpathy has abandoned vibe coding. It is a narrower warning: agents are much less dependable when software is unusual, technically dense, difficult to test and costly to evaluate.
Who is Andrej Karpathy?
Karpathy is a prominent machine-learning educator and researcher, a former Tesla AI leader and former OpenAI executive and co-founder. He is widely credited with introducing the phrase “vibe coding” in a February 2025 post. “Coined the term” is more precise than “invented AI programming”: natural-language coding and AI-assisted development existed before the phrase.
In its original sense, vibe coding meant describing a desired program in ordinary language, accepting generated code without necessarily understanding every line, running it, pasting errors back into the model and iterating. Karpathy presented it mainly as a playful approach for prototypes and “throwaway weekend projects,” not as a replacement for professional engineering. Futurism’s account quotes that caveat.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →#1 Best Overall
What Nanochat is
Nanochat is not merely a front end that sends prompts to a hosted chatbot. It is an open-source, from-scratch project for training a small language model and interacting with it through a ChatGPT-like web interface. The official repository describes it as “the best ChatGPT that $100 can buy.” That phrase is project positioning, not a guaranteed all-in cost for every user or hardware setup.
The repository includes training, evaluation, inference and web-serving components. Much of the implementation uses relatively vanilla PyTorch. Transformer depth is the principal complexity control, with other model settings derived from it, so users can experiment with different scales without managing a large collection of independent architectural knobs.
Nanochat can run on hardware supported by PyTorch, including CUDA systems and Apple Silicon through MPS, although the project notes that not every hardware path has been personally exercised. CPU or reduced MPS operation is possible in limited form; stronger results require more capable hardware. Training time and cost depend on the exact model configuration, GPU, region, storage and whether rented hardware is on-demand or discounted.
Rank #2
What Karpathy said about AI assistance
The reported facts are specific:
- Nanochat was “basically entirely hand-written.”
- Karpathy tried Claude and Codex agents.
- The attempts did not work well enough and were “net unhelpful.”
- He suggested that the repository might be too far outside the models’ data distribution.
This does not establish that every line was typed manually without any AI use. It also does not mean the agents could write none of the code, that Nanochat was impossible to build with assistance, or that Karpathy has rejected AI coding permanently. The contemporary report is available from Tech Yahoo and the original coverage from Futurism.
Why a project like Nanochat can defeat ordinary agent workflows
Karpathy did not publish a formal postmortem identifying one failure mechanism. The following are reasonable interpretations of the project’s characteristics, not confirmed findings from him.
An unusual codebase
Popular web applications contain familiar frameworks, conventions and examples. A from-scratch language-model training stack has fewer direct analogues. An agent may recognize individual PyTorch operations while misunderstanding the project’s particular assumptions about data, tensors, checkpoints or distributed execution.
Dense domain knowledge
Training infrastructure combines numerical computing, optimizer behavior, memory limits, hardware-specific performance and statistical evaluation. A locally plausible edit can alter gradient behavior, numerical stability or throughput without producing an obvious exception.
Cross-file and nonlocal correctness
A change to data loading can affect tokenization, batching, training metrics and reproducibility. A checkpointing change can surface later during evaluation or resumption. The relevant invariant may span many files, while the agent tends to propose a patch around the line named in the prompt.
Delayed and expensive feedback
A web-app defect may appear immediately in a browser. A training change may require hours of GPU time before its effect on loss, model quality or stability is visible. Repeated agent-driven experiments can therefore spend money while producing ambiguous evidence.
Code that runs can still be wrong
Compilation and a successful launch are weak tests for machine-learning infrastructure. A run can complete while silently reducing quality, increasing memory use, damaging reproducibility or changing the scientific meaning of an experiment.
What “outside the data distribution” means
A model is generally more reliable when a task resembles patterns in its training and reinforcement-learning experience. “Outside the data distribution” does not mean it has never seen PyTorch or language-model code. It means the exact combination of architecture, conventions, constraints and desired behavior may be unfamiliar enough to reduce reliability.
That explanation is plausible rather than a proof. Karpathy’s later writing describes a similar unevenness: models can perform impressively in some verifiable coding environments yet fail unpredictably on particular reasoning problems. A novel repository can therefore expose a model’s weak spot even when the same model is excellent at routine refactors.
Best Value
Nanochat is almost the opposite of classic vibe coding
Vibe coding works best when speed matters more than deep understanding and the result is easy to discard. Nanochat is intended to expose and control the mechanics of model training. Its success cannot be judged only by whether a page loads or a command exits without an error.
Relevant checks include:
- Does training remain numerically stable?
- Does the model reach the expected quality for the compute used?
- Are memory use, latency and throughput acceptable?
- Can another person reproduce the result?
- Do evaluation and checkpointing still mean what the author intended?
Those requirements make expert specification, measurement and review central. They do not create a binary choice between writing every line yourself and delegating everything to an agent.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Karpathy’s later view: from vibe coding to agentic engineering
In his 2026 Sequoia Ascent summary, Karpathy draws a distinction between two uses of AI. Vibe coding lowers the floor: more people can create working software. “Agentic engineering” raises the ceiling by using agents within professional standards for correctness, security, maintainability and taste.
He says agents became noticeably more useful around December 2025 in his experience, producing larger, more coherent and more reliable changes. His position is not that humans should stop using them. Humans still own the specification, architecture, judgment, security decisions and final oversight.
Recommended Free Tools
When to delegate and when to take over
| Task or environment | AI assistance is usually a good fit | Close expert supervision or manual implementation is preferable |
|---|---|---|
| Familiar application work | Boilerplate, CRUD screens, API integrations, documentation and migration scripts | Changes that cross undocumented system boundaries |
| Testing and maintenance | Drafting tests, updating repetitive fixtures and explaining existing code | Tests that could merely encode an existing implementation mistake |
| Technical risk | Low-stakes prototypes and personal tools | Security-sensitive, financial, medical or safety-critical systems |
| Specialized engineering | Scaffolding and narrow, independently verifiable utilities | Machine-learning infrastructure, distributed systems, performance kernels and novel algorithms |
A practical workflow
- Specify the boundary. State the invariant, interfaces, acceptance tests and non-goals before asking an agent to edit code.
- Delegate reversible work first. Documentation, scaffolding, routine refactors and test drafts are easier to inspect and undo.
- Validate independently. Run tests and measure domain outcomes such as model quality, latency, memory and reproducibility; do not stop at “it compiles.”
- Review sensitive changes manually. Check credentials, data handling, dependencies, permissions and generated logs.
- Stop unproductive loops. If the agent repeatedly patches symptoms, contradicts repository conventions or makes destructive edits, switch to a manual diagnosis.
Failure modes worth watching
- A locally reasonable edit breaks a cross-file invariant.
- Generated tests confirm the agent’s misunderstanding instead of the intended behavior.
- Dependency or API drift makes suggested commands invalid.
- Secrets, personal data or unsafe defaults enter code or logs.
- Generated code introduces licensing or attribution questions.
- Developers accept code they cannot later debug or maintain.
- Repeated cloud-GPU experiments accumulate costs before a meaningful result is obtained.
- Inconsistent generated styles make the repository harder to evolve.
These risks do not prove that AI-generated code is always insecure or slower. They explain why claims about universal replacement for programmers go beyond what the Nanochat episode demonstrates. Claude Code, Codex, editor-based tools such as Cursor and assistants such as GitHub Copilot can all be useful, but no product guarantees success on an unfamiliar research codebase.
The Bottom Line
Karpathy’s Nanochat experience is not an anti-AI reversal. It is evidence that coding agents are highly sensitive to familiarity, specification and verifiability. Use them aggressively for routine, testable work; keep humans responsible for architecture and security; and write or closely supervise the specialized core when correctness is difficult to observe.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




