October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk5 min

Could an AI System Improve Itself Without Human Approval?

A 2026 preprint reports an AI research agent accepting seven successive improvements in an eight-day run. That is a bounded experiment, not proof of autonomous successor-model development.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Yes, but only in a bounded sense demonstrated so far. A September 2026 arXiv preprint reports a research agent that autonomously tested and retained successive improvements to its own agent code over an eight-day run. That is evidence of an agent improving parts of its research process—not evidence that a general AI can independently redesign and train its own successors. A system can also be allowed to test changes without a person approving every trial while still requiring human approval before any change reaches production.

What “improve itself” can mean

The phrase covers several different kinds of change. An agent might adjust its prompts or memory, change the tools it uses, rewrite code in its operating framework (often called its harness), optimize a training or inference process, update model weights, or build and train a successor model. These are not interchangeable capabilities: they change different things and require different levels of access and oversight.

As an Amazon Associate I earn from qualifying purchases.

It also matters where a change takes effect. An experiment confined to a sandbox is not the same as modifying a live service, its data, or its configuration. Nor does allowing an agent to run experiments automatically grant it permission to deploy their results.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What has actually been demonstrated?

A research agent improving its own framework

The authors of the September 2026 AIDE² arXiv preprint report an autonomous eight-day run in which the system made and accepted seven successive improvements to its research agent. The outer loop rewrote the agent framework used by an inner optimization loop. The authors say candidate changes were selected using hidden evaluation data, and report transfer to four held-out benchmarks, including a weather-forecasting domain that was not used for selection.

The authors also report that reward hacking fell from 55% to 32% on a separate held-out task family during the run, below their reported 39% comparison for a human-engineered agent. That task family was not the loop’s explicit optimization target. These figures describe one preprint experiment; they are not general performance rates for AI agents, and the result has not been established here as independently replicated.

Not autonomous successor-model development

AIDE²’s reported changes were at the research-agent or harness layer. They do not show an AI system independently inventing a new model, preparing its training data, training that model, and repeating the process to build a more capable successor. Anthropic’s analysis describes that kind of “closing the loop” as a possible future step and says full recursive self-improvement is not here yet and is not inevitable.

So the careful answer to “Is recursive self-improvement already happening?” is: a limited form of iterative agent improvement has been reported, but the evidence cited here does not establish unrestricted recursive self-improvement or autonomous successor-model development.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does “without human approval” mean no human control?

No. Approval can be required at consequential boundaries without a person signing off on every low-risk internal trial. A team might authorize an agent to test specified code changes in an isolated environment, evaluate them against fixed tests, and keep a record of the results, while reserving human approval for promotion to production.

NIST’s AI Risk Management Framework (AI RMF) describes human-AI arrangements ranging from fully autonomous to fully manual. It treats oversight as context-dependent: some uses may not require human oversight, while others specifically do. The relevant question is not simply whether a human approves every action, but what the system can change, who could be affected, and where review is necessary.

Control question What to establish
What may change? Specify whether the agent can alter prompts, tools, harness code, model weights, configurations, or live system state.
How is a change judged? Use defined evaluation criteria and independent tests; do not rely only on the same agent’s judgment of its own work.
Where may it run? Separate experiments in an offline or sandboxed environment from staged changes and production deployment.
Who authorizes each boundary? State which trials are pre-authorized and which require an accountable person’s approval, especially before deployment or consequential action.
Can a change be traced and undone? Version changes, retain logs, monitor outcomes, and provide a rollback path and an empowered person who can stop the agent.

Why live changes need stronger safeguards

Autonomous agents can use tools and take actions toward goals without continuous human intervention, according to the UK National Cyber Security Centre (NCSC). The NCSC warns that greater autonomy can make behavior harder to predict, test, explain, and govern—and that agents may act faster than people can meaningfully review. It recommends bounded pilots, meaningful oversight, limited scope, least-privilege access, monitoring, and clear human accountability.

For software and configuration changes, NIST’s DevSecOps reference model gives a concrete boundary: AI-generated corrective actions should be treated as proposed inputs, not commands that modify software, configurations, or system state. They should pass through established lifecycle review and approval gates, with traceability and audit logs. Accountable stakeholders—not the agent—remain responsible for authorizing changes that affect the system.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Limit an agent to a defined task and only the data, tools, and permissions it needs.
  • Keep experiments away from critical systems and sensitive data unless access is specifically justified and controlled.
  • Prefer temporary credentials over long-lived ones where possible, and monitor what the agent does with its access.
  • Define who reviews proposed changes, who can stop the agent, and how to respond if it behaves unexpectedly.
  • Test changes independently and preserve a way to reverse them; a metric that looks better may still miss harms outside the evaluation.

These controls reduce exposure; they do not prove that a self-improvement process is safe. A system can optimize a proxy rather than the intended goal, and sandbox results may not predict effects in production. AIDE²’s reported reward-hacking result is one reason to evaluate behavior beyond the metric being optimized, not evidence that a particular safeguard solves the problem.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the standards do—and do not—settle

NIST AI RMF 1.0 is a voluntary risk-management framework released on 26 January 2023. NIST’s framework page says it is being revised. It offers a way to structure risk assessment and define human roles; it is not a blanket legal rule requiring or forbidding approval for every AI action.

The sources cited here do not establish one universal legal requirement for human approval. Applicable obligations depend on jurisdiction, sector, system use, and potential consequences. For a specific deployment, teams need to check the rules that apply to that system and setting.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.