AI4AI Survey: From Long-Horizon Agents to Recursive Self-Improvement
Definitions, Reliable Horizons, and Open Problems · 23 authors · 7 institutions
Updated weekly with the month's top AI-for-AI papers, news, and blogs. Stay tuned 🔥
|
🔄 Automatic rankings Citations, GitHub stars, and rankings refresh every Monday. |
🗞️ Weekly update The month's top papers and releases, re-ranked every week against primary sources. |
🧭 Survey-grounded map From long-horizon agents to recursive self-improvement. |
Star ⭐ to save the map · Watch 👀 for weekly updates · Share 🔁 with your lab
- 📄 2026-08-30 — Companion survey now online. Read AI4AI Survey: From Long-Horizon Agents to Recursive Self-Improvement—Definitions, Reliable Horizons, and Open Problems.
- 🔄 Every Monday — Automatic refresh. Citations, GitHub stars, and recent-paper/yearly rankings update automatically.
- 🚀 2026-09-16 — Latest weekly edition published. The living catalog and source-verified news digest are up to date.
🗂️ Explore the full collection
Updated 2026-09-16 · The month's top stories, refreshed every Monday alongside citation rankings.
How we select: The top stories from the trailing 30 days, scored on publisher authority, discussion volume, whether a concrete verifiable result is reported, whether it changes what agent builders do now, whether an artifact was released, and expected durability. Only primary sources are cited, and performance claims remain attributed to their publishers.
| Date · Type | News | Why it matters |
|---|---|---|
| 2026‑09‑11 Model release |
Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier Sakana AI |
Sakana AI describes a model trained to route tasks across a fixed pool of open-weight and specialized models and to recursively call instances of itself. It reports 48.3 on Chartography against 27.3 for Opus 5, and says the scores are reached without Fable 5, Fable 5.1 or GPT-6 Astra in the agent pool. |
| 2026‑09‑10 Open-weight release |
DeepSeek-V4.1-Flash DeepSeek |
DeepSeek publishes weights on Hugging Face for a 552B mixture-of-experts model built on a new causal encoder-decoder architecture that activates roughly 8B parameters on input and 16B on output. It reports 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE v1.1, and says the KV cache needs a quarter of the prior generation's HBM. |
| 2026‑09‑10 Harness release |
Introducing the Agents API OpenAI |
OpenAI ships the harness that runs Codex as a managed service in public beta, saying it handles session orchestration, context compaction, sub-agent coordination and recovery while developers supply tools and choose a sandbox. It offers versioned harness access alongside each model launch. |
| 2026‑09‑10 Model release |
Introducing SWE-2: Pushing the Pareto Frontier Cognition |
Cognition says it scaled reinforcement learning to a 2.8T-parameter Kimi K3 base for the first time, training every reasoning-effort level in a single run with a cost-penalized reward. It reports 50.0 on its own FrontierCode 1.1 Main at 64 percent lower cost than Fable 5.1, and 58 percent fewer turns than SWE-1.7. |
| 2026‑09‑08 Paper |
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness arXiv |
The authors log predicted capability demand, chosen service tier and interaction outcome for each turn, then reuse those routing signals as a three-stage fine-tuning curriculum and an evaluation-selection-update loop. They report macro-average gains from 58.94 to 64.87 at 4B and 65.60 to 69.04 at 9B across eleven benchmarks, with weights and code released. |
| 2026‑09‑03 Paper |
EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness? arXiv |
The authors hold tasks fixed and instead grow the harness across tools, skills and agents over 17 multi-stage streams covering 802 tasks, 520 tools, 42 skills and 62 agents. They report harness-induced forgetting of up to 34.7 percent backward transfer on the agents axis with model parameters unchanged. |
| 2026‑09‑03 Model release |
GPT-6 Astra: A new generation of intelligence OpenAI |
OpenAI reports 64.6 percent on Terminal-Bench Science 0.1 and 57.9 percent on Terminal-Bench 4.0, and introduces experimental Codex notes and searchable history across context windows for long-running work. |
| 2026‑09‑02 Model update |
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber |
Google reports 54.9 percent on HLE-Verified for Gemini 3.8 Flash while retaining the introductory token prices of 3.7 Flash. It says long-running agentic loops helped evaluate and refine the models, and releases Flash Cyber through trusted-defender access. |
| 2026‑09‑01 Model release |
Introducing Claude Fable 5.1 and Claude Mythos 5.1 Anthropic |
Anthropic reports 52.6 percent on Terminal-Bench-Science 0.1 versus 24.7 percent for Fable 5, and estimates that lower cache-read pricing reduces highly agentic workload costs by up to roughly 45 percent. |
| 2026‑08‑31 Paper |
WHALE: A Simple Recipe for Joint Harness-Weight Optimization arXiv |
The authors alternate model-weight updates with executable harness search, reporting gains of 4.15 to 24.38 percentage points over the tested single-component and prompt-weight baselines across three domains with Qwen3.5-2B/4B agents. Code is released for reproducing the joint optimization recipe. |
Want next month's news? Watch the repository. Citation counts, rankings, and the month's top stories refresh every Monday. Browse past editions →
Citation counts are current through 2026-09-15 from Semantic Scholar and OpenAlex. Rankings are discovery aids, not quality scores; audit evidence remains independent of popularity. GitHub stars are snapshots from 2026-09-16. For papers indexed as multiple versions, retain the largest title-verified count reported by the configured sources. All yearly rankings use first public appearance year; a later venue year never moves a paper into a newer cohort.
Top 12 of 2026 by citations
Top 12 of 2025 by citations
Top 12 of 2024 by citations
Top 12 of 2023 by citations
111 papers · Survey-curated collection, newest first. Cross-collection papers may appear in more than one section.
101 papers · Survey-curated collection, newest first. Cross-collection papers may appear in more than one section.
26 papers · Survey-curated collection, newest first. Cross-collection papers may appear in more than one section.
- Discover: survey searches, backward/forward citation chaining, and community suggestions.
- Verify: exact title and identifier checks against primary scholarly sources.
- Classify: Benchmarks, Harness Design, and Model Design, allowing justified overlap.
- Audit when evidence permits: stage ownership plus independent G/R/H/T coordinates.
- Refresh weekly: citation counts and rankings every Monday, alongside the month's top papers, releases, blogs, and research news.
If this map or its evidence audit helps your work, please cite the companion survey:
Copy BibTeX
@misc{wu2026eveai4ai,
title = {{AI4AI} Survey: From Long-Horizon Agents to Recursive Self-Improvement---Definitions, Reliable Horizons, and Open Problems},
author = {Wu, Kai and Lyu, Hao and Luo, Zhen and Wang, Chaofan and
Ye, Siyu and Lin, Jinghao and Ji, Xiaozhong and Jiang, Boyuan and
Wang, Shengzhi and Wang, Zihan and Ye, Yiwen and Wang, Hao and
Wang, Zimu and Liu, Wenzhe and Wang, Ruobing and Cai, Kai and
Xiong, Mingliang and Fang, Wen and Liu, Mingqing and
Zhang, Yifan and Yang, Lei and Hu, Xiaobin and Liu, Qingwen},
howpublished = {Preprints.org},
year = {2026},
doi = {10.20944/preprints202608.2108.v1},
url = {https://doi.org/10.20944/preprints202608.2108.v1}
}Have something to add? Missing a paper, official code link, or stronger primary-source evidence? Open a pull request or use the paper-suggestion form. Read the contribution guide →




