Short answer: The November 12, 2024 headline described a reported slowdown in gains from traditional AI pretraining—not proof that OpenAI, or artificial intelligence generally, had reached a permanent ceiling. Reporting about the model reportedly code-named Orion relied largely on unnamed sources, while former OpenAI chief scientist Ilya Sutskever separately warned that pretraining results were plateauing. Public evidence also showed a different path forward: spending more computation while answering a prompt, as OpenAI demonstrated with its o1 reasoning models.
What the original report actually said
Futurism’s November 12, 2024 story, also syndicated by Yahoo News Australia, summarized reporting that OpenAI’s next full-scale model, reportedly code-named Orion, was producing a smaller improvement over GPT-4 than GPT-4 had produced over GPT-3. Unnamed researchers reportedly saw little or no dependable progress on some tasks, with coding mentioned as a particularly difficult area.
That is a materially narrower claim than “AI has stopped improving.” The account was based on unnamed OpenAI researchers and a report from The Information; no public technical evaluation allowed outsiders to reproduce the result at the time. Ars Technica’s contemporaneous analysis likewise treated the story as a warning about weaker returns, not a confirmed failure of the model.
The available public record does not establish whether Orion was ultimately released under that name, what its final benchmark results were, how much its training run cost, or whether the reported plateau continued through later model generations.
“Diminishing returns” is not the same as a hard wall
In economics, diminishing returns means that adding more of one input, while holding other inputs fixed, produces progressively smaller additions to output. AI systems do not keep every other input fixed: researchers can change the architecture, training objective, data mix, optimization method, hardware, post-training, tools and inference strategy.
Here, the phrase is best understood as shorthand for weaker returns from one recipe: conventional pretraining at larger scale. Scaling laws are empirical relationships observed over particular ranges and setups. They do not promise equal capability jumps, useful gains on every task, or a fixed amount of progress for every additional dollar. Nor does a weak result from one model disprove them.
Why simply adding more training may help less
The explanations below are plausible mechanisms discussed in coverage, not all proven causes of Orion’s reported result.
- Data exhaustion and quality: The most useful, clean public text may already be heavily represented in training corpora. Adding repetitive, noisy or weakly sourced material can contribute less than a smaller expert-curated set. “Running out of data” is therefore too broad unless it specifies whether the scarce resource is deduplicated text, licensed material, demonstrations, synthetic traces, multimodal data or domain expertise.
- Benchmark saturation: A model may improve in ways that barely move a test score that is already near its ceiling. Conversely, a score increase can reflect narrow optimization against a known format rather than broader competence.
- Distribution mismatch: More internet text does not automatically produce better coding, factuality, planning, tool use or reliability in a production workflow.
- Optimization and architecture limits: Larger runs can expose weaknesses in objectives, model design or data mixture that more chips alone cannot fix.
- Evaluation noise: Small benchmark differences may not be statistically reliable or practically important, especially when test sets are limited.
- Synthetic-data degradation: Model-generated training data can amplify errors or reduce diversity unless it is filtered carefully and its provenance is controlled.
- Economics: A marginally better model may be unattractive if the extra training and serving cost exceeds what customers will pay.
These factors make “more tokens” an incomplete description of scaling. Raw token count, data quality, expert supervision and real-world outcome data are different inputs with different costs and benefits.
What Sutskever’s warning means
Reuters reported that Ilya Sutskever, who left OpenAI in 2024, said the field’s 2010s era of scaling pretraining was giving way to an “age of wonder” requiring new ideas. His perspective matters because he was a major architect and advocate of large-scale deep learning, but his interview was not an official OpenAI announcement that progress had stopped. Reuters’ report, republished by Investing.com, described a search for alternatives as existing methods became less predictable.
Rank #2
- [Local AI Inference & 70B Model Ready] Equipped with the AMD Ryzen 7 PRO 8845HS processor, NEXUS is engineered for heavy local AI workloads. With a full-size GPU bay, it runs 70B LLMs natively without an internet connection. Ideal for AI developers and tech enthusiasts who need private environment for coding and model testing.
- [132TB Mass Storage with ZFS Integrity] Features a hybrid storage architecture (3×NVMe + 4×3.5" HDD) supporting up to 132TB. Utilizing the enterprise-grade ZFS file system and ECC memory, it prevents data corruption and bit rot—a must-have for professional photographers and video editors safeguarding 4K/8K RAW footage.
- [OpenClaw-Driven Automation Workflow] The built-in OpenClaw execution layer allows complex automated tasks to be processed locally. Even when offline, your backup schedules and AI file organization continue seamlessly. Say goodbye to monthly cloud subscriptions and high latency.
- [Dual 10GbE & USB4 Ultra-Connectivity] Experience server-class speeds with dual 10GbE ports and a 40Gbps USB4 interface. It enables multi-user real-time collaboration on large project files directly from the NAS, ensuring zero-lag editing for creative studios and production teams.
- [Open-Source ZimaOS for Total Privacy] Running on the fully open-source ZimaOS, NEXUS ensures your data stays physically on-premise with no backdoors. It acts as a "Digital Fortress" for privacy-conscious families and small businesses who demand absolute data sovereignty.
The alternative axis: compute while answering
Training-time compute updates a model’s parameters before deployment. Inference-time or test-time compute is spent after a user submits a prompt. A system can generate several candidate solutions, check them, rank them and revise an answer instead of producing one immediate completion.
OpenAI presented o1 as an example of this approach in “Learning to reason with LLMs”. On the 2024 AIME mathematics examination, OpenAI reported these results under its stated evaluation conditions:
| System and procedure | Average score | Approximate share of 15 problems |
|---|---|---|
| GPT-4o | 1.8/15 | 12% |
| o1-preview, one sample | 11.1/15 | 74% |
| o1-preview, consensus across 64 samples | 12.5/15 | 83% |
| Reranking 1,000 samples | 13.9/15 | 93% |
These are OpenAI-reported benchmark results, not proof of general intelligence. The multi-sample figures also are not equivalent to one ordinary chatbot response: generating and evaluating dozens or thousands of attempts consumes additional accelerator time, raises cost and can increase latency.
- Responses can take longer, especially on difficult prompts.
- Serving costs and accelerator demand rise with additional reasoning.
- Product pricing becomes harder to predict for variable-complexity requests.
- Longer deliberation does not guarantee factuality or correct understanding of the task.
- Extra computation can produce verbose answers or fail real-time requirements.
Does this disprove scaling laws?
No. It challenges the simplistic idea that making a general model larger is always the best or cheapest route to progress. Capability can also improve through algorithms, data curation, retrieval, external tools, memory, multimodal training, domain fine-tuning and specialized systems.
Smaller models can beat larger general models on a defined workflow when they use better retrieval, tools or domain data. Software and hardware efficiency matter too: quantization, distillation, batching, caching, improved utilization and architectural changes can lower the cost per useful answer even if parameter counts stop growing. OpenAI discusses this capability-per-dollar perspective in “AI and efficiency.”
Rank #3
The technical question is also a business question
A frontier model can deliver modest broad gains yet be commercially valuable if it is cheaper, more reliable on a high-value workflow, faster with tools or able to replace expensive human effort. The reverse is also true: a large benchmark improvement may have weak business value if customers cannot verify outputs, tolerate the latency or afford the serving bill.
Training and inference can move in opposite directions. A company may reduce the need for ever-larger pretraining runs while increasing inference demand through reasoning models. That changes data-center planning, energy use, pricing and capacity commitments. Buyers should measure cost per successfully completed task, not just tokens or leaderboard points.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →How to judge whether the slowdown thesis is real
Five tests provide a more rigorous standard than a single headline:
- Capability: Does a successor solve harder problems, including adversarial and unfamiliar ones?
- Reliability: Does it make fewer serious mistakes in ordinary workflows?
- Efficiency: Does it achieve the result with fewer chips, tokens or seconds?
- Economic value: Does the measurable improvement justify training and serving costs?
- Generalization: Does the gain transfer beyond the benchmark used during development?
Useful signals over time would include smaller gains from successive pretraining runs, rising cost per capability improvement, greater reliance on inference-time reasoning, more proprietary or synthetic data, and increased use of specialized models. Customer willingness to pay for slower reasoning systems is an equally important test.
What the 2024 evidence does—and does not—show
- Publicly documented: OpenAI’s explanation and evaluation of o1’s reasoning approach.
- Reported but not independently verified: Weaker-than-expected internal results for the model referred to as Orion.
- Expert interpretation: Claims that easy gains from conventional pretraining were becoming harder to find.
- Not established: A permanent ceiling, the failure of scaling laws, the end of commercial AI progress, or an inevitable collapse in AI investment.
The “law of diminishing returns” headline therefore describes a strategic transition more accurately than a terminal diagnosis. The industry may be shifting from spending almost exclusively on larger pretraining runs toward a portfolio of better data, specialized training, efficient infrastructure and computation allocated at the moment a model reasons.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




