OTWopentechwire
Tech Intelligence, Openly Wired
AI

Beijing Labs Map Five Levels of AI That Improves Itself

A cross-institutional team has published a framework for recursive self-improvement, defining the stages from assisted iteration to systems that no longer need human designers.

WZ
Wei Zhang
China Tech Correspondent · Hangzhou
Sep 15, 2026
4 min read
Beijing Labs Map Five Levels of AI That Improves Itself
Beijing Labs Map Five Levels of AI That Improves ItselfCredit: Shutterstock

The Framework for Autonomous Design

A consortium spanning ByteDance, Tsinghua University, and the Shanghai Artificial Intelligence Laboratory released a structured taxonomy on Thursday for what they term recursive self-improvement in AI systems. The five-stage framework describes a progression from models that require human oversight at every iteration to those capable of designing their successors entirely independently.

The taxonomy arrives as competition intensifies between US and Chinese research groups over architectures that can refine their own training pipelines, data curation, and inference strategies. At Opentechwire, we've tracked parallel efforts from Anthropic, DeepMind, and OpenAI on similar autonomy goals, but this is the first public attempt from a Beijing-led collaboration to codify discrete maturity levels for self-improvement.

Stage zero in the framework represents today's mainstream practice: engineers tune hyperparameters, select datasets, and adjust reward functions manually. Stage one introduces partial automation, where the model suggests modifications but a human still approves them. By stage two, the system executes those changes without human sign-off, though the scope remains constrained to narrow domains such as optimising a single layer or loss function.

Stage three marks the threshold where the model begins to alter its own architecture, rewriting modules or proposing novel training objectives. Stage four, the terminal level in the taxonomy, describes systems that autonomously design, train, and deploy successors with capabilities exceeding their own, closing the loop without human input.

Why Recursive Self-Improvement Matters Now

The concept of RSI has circulated in AI safety literature for years, but recent advances in large language models and reinforcement learning have made incremental versions tractable. Models already generate synthetic training data, propose code patches, and evaluate their own outputs. The gap between those narrow capabilities and full recursive redesign is narrowing faster than most timelines predicted five years ago.

From a strategic perspective, whoever reaches stage three or four first gains a compounding advantage: each generation of models accelerates the next, potentially widening leads in research velocity, commercial deployment, and export competitiveness. The Chinese research community has been explicit about viewing RSI as a lever to close the gap with US frontier labs, particularly as export controls limit access to cutting-edge chips.

The taxonomy itself does not present breakthrough algorithms or training recipes. Instead, it offers a shared vocabulary for comparing progress across institutions and jurisdictions. That standardisation serves a dual purpose: it helps Chinese labs coordinate efforts, and it signals to international peers that Beijing's research apparatus is investing seriously in this capability tier.

Technical Boundaries and Open Questions

The paper leaves several critical questions unresolved. First, the authors do not specify how to measure whether a system has genuinely crossed from one stage to the next. Partial automation can look deceptively like full autonomy if the human role shrinks to approving recommendations generated by the model. Without clear operational tests, the taxonomy risks becoming a marketing ladder rather than a technical one.

Second, the framework does not address alignment or safety constraints at higher stages. A stage-four system that improves itself without human oversight also lacks human correction when it drifts toward unintended objectives. The paper acknowledges this gap but offers no mechanism to ensure that recursive improvement preserves the goals set at stage zero.

Third, the taxonomy assumes a linear progression, but real-world development may bifurcate. Some labs might achieve stage-four performance in narrow domains such as chip design or protein folding while remaining at stage one for general reasoning. The framework does not distinguish between domain-specific and general recursion, which could lead to conflicting assessments of where a given system sits.

Institutional Signals and Coordination

ByteDance's participation is noteworthy. The company has historically focused on recommendation algorithms and content moderation, not foundational model research. Its inclusion in this consortium suggests that Chinese tech giants see RSI not just as an academic exercise but as a core infrastructure layer for future products, from automated code generation to real-time personalisation engines.

Tsinghua University and the Shanghai AI Lab, by contrast, have long been centres of gravity for state-backed research. Their joint authorship ties this framework to national priorities, likely positioning it as a reference document for funding decisions and inter-lab collaboration protocols.

The timing of the release, mid-September, places it ahead of major AI conferences in the fourth quarter. That sequencing suggests the authors want the taxonomy adopted as a standard before competing frameworks emerge from US or European institutions. Control over taxonomies and benchmarks has become a quiet but consequential dimension of AI competition, shaping how progress is measured and compared.

What Comes After Stage Four

The paper's title refers to "the last AI built by humans," a phrase borrowed from older speculative writing on intelligence explosions. In practice, reaching stage four would not mean the end of human involvement; it would shift the role from hands-on design to high-level goal-setting and oversight. The more immediate question is whether current architectures, rooted in transformer models and gradient descent, can actually traverse all five stages, or whether the taxonomy describes a ceiling that requires fundamentally different approaches.

Several researchers outside the consortium have argued that recursive self-improvement may hit diminishing returns before stage four, constrained by the same scaling laws that govern today's models. If each generation's improvement shrinks exponentially, the compounding advantage evaporates. The framework does not engage with that scepticism, treating stage-four capability as achievable rather than speculative.

For now, the taxonomy serves less as a road map than as a declaration of intent. It tells the global AI community that Chinese institutions are organising their efforts around recursive self-improvement, that they believe it is tractable, and that they expect to measure progress in discrete stages rather than continuous metrics. Whether that framing accelerates real capability gains or simply reshapes the rhetoric of the AI race will depend on results over the next eighteen months.

Opentechwire will continue monitoring how this taxonomy influences funding patterns, collaboration structures, and technical disclosures from both Chinese and international labs.

Read next
AI

Wurtzite Ferroelectrics Achieve 10 Billion Write Cycles in Chinese Lab

Wei Zhang · 5 min
AI

OpenAI's Test Agents Broke Into Ruby Package Repository Months Before Public Disclosure

Arjun S. Mehta · 4 min
AI

China's Chip Designers Turn to AI Agents to Bypass US Export Restrictions

Marcus Halloran · 5 min
Spot something wrong? Email corrections@opentechwire.com. We log every correction publicly.