“Today’s AI is alchemy, not science” is useful as a criticism, but false if read literally. Modern AI rests on mathematics, statistics, computer science, controlled experiments and large-scale engineering. Yet many frontier systems are discovered through trial, tuned by recipes and deployed with less explanation and standardized testing than their marketing suggests. The practical question is not whether AI belongs in a binary category of science or alchemy. It is which claims are experimentally supported, which are engineering observations, and which remain speculation.
What the “alchemy” metaphor is really saying
Alchemy is a useful metaphor when results arrive before theory. Frontier developers may find a capability after changing model scale, data, post-training, prompting or tool access, then struggle to explain why it appeared or when it will fail. That does not make the work irrational. It means several activities are mixed together:
- Science: hypotheses, controlled experiments, measurement, statistical analysis, benchmarks and attempts at replication.
- Engineering: optimizing a system for cost, latency, reliability, safety and product requirements.
- Craft: practical choices learned from experience but not always captured in a public theory.
- Marketing: claims that can extend beyond what the evidence establishes.
The metaphor is most justified when a compelling demonstration is presented as proof of general intelligence, dependable reasoning or safe autonomy without the testing needed to support those conclusions. The phrase is associated with a 2023 discussion of a VentureBeat article, as recorded by Mystery AI Hype Theater 3000; it is a thesis about the field’s culture, not a scientific classification.
What is genuinely scientific about AI?
AI research uses familiar scientific practices even when its objects are difficult to interpret. A model has a mathematical definition. Training runs can be configured, logged and compared. Researchers can hold out test data, vary one factor in an ablation, measure uncertainty and publish methods for others to reproduce. Software and hardware configurations can be pinned, and narrow systems can sometimes receive formal verification.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Benchmarks are imperfect instruments, not useless ones. They let researchers compare systems under stated conditions and reveal progress or regressions. Stanford HAI’s 2025 AI Index reports substantial year-over-year gains on MMMU, GPQA and SWE-bench, alongside falling inference costs, growing adoption and major applications in science and medicine. The same report records persistent weaknesses on complex reasoning and logic tasks, rising AI-related incidents and relatively uncommon standardized responsible-AI evaluations among major developers. Both parts are important: measured capability has advanced, while broad reliability and safety evidence remain uneven.
A deployed product is normally a technology, not a scientific theory of intelligence. It can be evaluated scientifically without explaining human cognition. A useful artifact does not, by itself, validate the story people tell about how it works.
Where the metaphor fits
Results before explanations
Large models can display unexpected abilities after changes in scale, data mixture, architecture, post-training or tools. Researchers can measure the behavior and reproduce it under some conditions without possessing a compact causal account of the internal process. Calling such capabilities “mysterious” should mean incompletely explained, not contrary to mathematics.
Opaque mechanisms
Parameters are numerical representations learned from data rather than hand-written rules. Interpretability researchers can inspect activations, compare checkpoints and run interventions, but a successful probe is not automatically a complete explanation of a particular answer. It helps to separate three levels:
Recommended Free Tools
- Predictive understanding: knowing how a system tends to behave.
- Mechanistic understanding: identifying internal representations and computations that cause a behavior.
- Scientific explanation: a general account that predicts behavior beyond the tested cases.
Current systems often offer the first level more reliably than the third.
Benchmark dependence
A score can reflect exposure to similar examples, prompt engineering, tool access, grader-specific optimization, memorization or a narrow task definition. A benchmark may also stop measuring generalization if its questions or close paraphrases enter training data. Any serious claim should identify the model and version, evaluation date, test set, contamination controls, tools or retrieval used, scoring method and whether the result was independently reproduced.
Rank #2
Prompt folklore
Prompting can resemble a craft tradition: users exchange recipes, small wording changes produce large effects, results vary by model version and a successful technique can fail after an update. That is evidence of a complex, partially characterized system—not evidence of magic.
Post-hoc stories
A generated rationale, a product explanation and an interpretability result are different things. An AI’s explanation of an answer may help a reader communicate or audit the output, but it is not automatically evidence that the stated steps caused the answer or accurately report the model’s internal computation. The same caution applies to human-readable chain-of-thought.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallUseful is not the same as understood
Usefulness is an engineering property: does the system improve a defined workflow at an acceptable cost? Explanation is a scientific property: do we know why it behaves as it does and when that account will generalize? Reliability is an operational property: does it continue to meet requirements for the real population, including rare and adversarial cases?
A system can therefore be useful without being deeply understood, accurate on average without being dependable in every case, and impressive in a demonstration without being ready for unsupervised high-stakes use. “AI does not understand” is too broad unless the speaker defines understanding. Behavioral competence, human-like comprehension and mechanistic explanation are separate claims.
Why benchmark wins do not settle the question
Benchmarks are snapshots under designed conditions. They do not automatically establish performance on a company’s documents, a hospital’s edge cases or a changing public environment. Distribution shift can arise from new terminology, different writing styles, scanned tables, rare events, adversarial inputs or a changed format. A model may solve a familiar test pattern while failing a nearby task that requires robust reasoning.
Conversely, benchmark criticism should not become capability denial. Stanford’s 2025 measurements show real progress, and AI systems already support useful work. The correct conclusion is narrower: benchmark gains establish capability under specified conditions; they do not, by themselves, establish general reasoning, calibrated confidence or safe autonomy.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Failure modes that turn capability into risk
Confident falsehoods
Fluent output can be fabricated or built on a false premise. Retrieval and citations reduce some errors but do not remove the need to verify important claims.
Automation bias
People may accept polished output they would challenge if it looked obviously machine-generated. Review procedures should make disagreement easy and record who approved a consequential action.
Error cascades
A tool-using agent can compound small mistakes: an incorrect interpretation produces a wrong query, which produces a wrong intermediate result and then a confident action. A drafting chatbot and an agent that can edit records, execute code, send messages or spend money require different controls.
Version drift
A product name can hide changes to the underlying model, system prompt, safety filters, context limits, tools, rate limits, pricing or data-use policy. Record the model identifier, configuration and test date, and retest after material updates.
Open-weight is not fully transparent
Published parameters do not necessarily reveal training data, filtering, post-training data, evaluation harnesses, deployment settings or safety fine-tuning. Openness on one layer should not be treated as transparency across all layers.
What this means for businesses and buyers
Evaluate the task, not the label “AI.” Establish a baseline against existing software, a human process or doing nothing. Define which errors matter, how much review costs and what happens when the system is uncertain or unavailable. NIST’s voluntary AI Risk Management Framework (AI RMF) 1.0 was released on January 26, 2023; its generative-AI profile followed on July 26, 2024. NIST says the framework is being revised under the White House AI Action Plan and published a critical-infrastructure profile concept note on April 7, 2026.
The accompanying NIST Playbook groups suggested actions under Govern, Map, Measure and Manage. It is voluntary guidance, not a mandatory checklist or a fixed sequence. A practical deployment should include:
- task-specific tests on the organization’s own data;
- independent checking of consequential outputs;
- documented retention, privacy and training-use terms;
- version and change management;
- logs, incident reporting and rollback procedures;
- an escalation path to a qualified human;
- a fallback process if the vendor changes the model, endpoint or price.
Questions to ask a vendor
- What exact task is automated, and what is the baseline?
- What error types and worst-case outcomes matter?
- Was evaluation performed on our data and independently reproduced?
- Are retrieval, tools or hidden prompts part of the claimed result?
- What happens when confidence is low or inputs are out of distribution?
- How are model updates announced, tested and rolled back?
- Can we export data, logs and evaluation results?
- What are retention, training-use, security and privacy terms?
- What is the fallback if the service changes or becomes unavailable?
What this means for scientific research
AI can be a scientific instrument without being a scientific explanation of intelligence. It may help discover materials, predict protein structures, analyze images or write code. A model-generated hypothesis still needs traceable sources, independently reproducible analysis, appropriate controls, statistical validation, domain-expert review and experimental or observational confirmation.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteThat distinction—AI for science versus AI as science—prevents two errors. Treating a useful predictor as a theory overstates what has been learned; dismissing a useful instrument because its internals are opaque understates what it can do.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What this means for policy
Policymakers who assume AI is more understood and controllable than it is may overtrust vendor benchmarks, regulate labels instead of actual risks, permit automation without recourse or underestimate monitoring after updates. Calling every AI system “alchemy” creates the opposite problem: blanket skepticism obscures narrow systems that are demonstrably validated.
Rules should therefore focus on the use, consequences, evidence and ability to contest an output. Regulatory requirements vary by country and sector; NIST’s framework is guidance, not a universal legal standard. High-impact deployments need documented evaluation, human accountability, incident response and a way to suspend or reverse automated decisions.
A field guide to judging an AI claim
Use this evidence ladder before accepting a performance or safety claim:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
| Level | Evidence | What it supports |
|---|---|---|
| 1. Demonstration | A compelling example from a vendor or user | Discovery and illustration, not reliability |
| 2. Repeatable test | Many examples under fixed, documented conditions | Initial evaluation |
| 3. Independent replication | A separate evaluator reproduces the result without private developer tooling or data | Stronger confidence |
| 4. Distribution-shift testing | New users, domains, formats, adversarial cases and changed conditions | Deployment decisions |
| 5. Operational monitoring | Post-launch performance, drift and incidents are measured, with rollback or human review | High-stakes, long-lived use |
The “alchemy” criticism is strongest when a claim stops at Level 1 but is marketed as though it had reached Level 4 or 5.
Choosing an AI product without buying the hype
Compare products on verification burden rather than a vague ranking of intelligence. The relevant questions are task-specific accuracy, failure severity, citation and source-checking behavior, model stability, privacy, logging, context and file limits, rate caps, migration costs, grounding quality and exportability.
| Option | Potential fit | Important qualification |
|---|---|---|
| ChatGPT / OpenAI | General writing, research assistance, coding, document analysis and business deployment | Business, Enterprise and API routes are listed through OpenAI’s official pricing page; test the exact workflow and monitor model changes. Official pricing |
| Claude / Anthropic | Long-form writing, document analysis, coding and extended sessions | Anthropic lists Free, Pro, Max 5x and Max 20x. Usage depends on conversation length, model and features rather than a fixed message count; limits can apply. Official pricing |
| Gemini API / Google AI | Multimodal applications, Google integration, grounding and pay-as-you-go APIs | Google lists free and paid tiers, token and modality rates, grounding charges and lifecycle notices. Endpoint changes require monitoring. Official pricing |
Plan and limit information above reflects pages checked August 18, 2026; offers, prices and endpoints can change. Consumer subscriptions and API billing are different purchasing paths. No brand guarantee substitutes for evaluation on the work you actually need done.
The strongest counterargument
Many sciences began with reliable empirical regularities before they had comprehensive theories. AI may develop better explanatory science as interpretability, evaluation and causal modeling improve. The present criticism is not that AI can never become a mature science. It is that current capability claims often outrun mechanistic understanding and standardized evidence.
The Bottom Line
Use AI as experimental technology: useful enough to test, uncertain enough to measure, and consequential enough to monitor.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




