Skip to content

Repository files navigation

Towards AI that improves AI — the plan, execute, feedback, repair loop

🚀 Awesome AI4AI

AI4AI Survey: From Long-Horizon Agents to Recursive Self-Improvement
Definitions, Reliable Horizons, and Open Problems · 23 authors · 7 institutions

Updated weekly with the month's top AI-for-AI papers, news, and blogs. Stay tuned 🔥

Paper, nice layout Paper on Preprints.org Project Site RSIHub Harness Blog DOI


Why This Repo

🔄 Automatic rankings
Citations, GitHub stars, and rankings refresh every Monday.
🗞️ Weekly update
The month's top papers and releases, re-ranked every week against primary sources.
🧭 Survey-grounded map
From long-horizon agents to recursive self-improvement.

Star ⭐ to save the map · Watch 👀 for weekly updates · Share 🔁 with your lab

What's New

  • 📄 2026-08-30 — Companion survey now online. Read AI4AI Survey: From Long-Horizon Agents to Recursive Self-Improvement—Definitions, Reliable Horizons, and Open Problems.
  • 🔄 Every Monday — Automatic refresh. Citations, GitHub stars, and recent-paper/yearly rankings update automatically.
  • 🚀 2026-09-16 — Latest weekly edition published. The living catalog and source-verified news digest are up to date.
🗂️ Explore the full collection

📅 Weekly Update · Monthly Top 10

Updated 2026-09-16 · The month's top stories, refreshed every Monday alongside citation rankings.

How we select: The top stories from the trailing 30 days, scored on publisher authority, discussion volume, whether a concrete verifiable result is reported, whether it changes what agent builders do now, whether an artifact was released, and expected durability. Only primary sources are cited, and performance claims remain attributed to their publishers.

Date · Type News Why it matters
2026‑09‑11
Model release
Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier
Sakana AI
Sakana AI describes a model trained to route tasks across a fixed pool of open-weight and specialized models and to recursively call instances of itself. It reports 48.3 on Chartography against 27.3 for Opus 5, and says the scores are reached without Fable 5, Fable 5.1 or GPT-6 Astra in the agent pool.
2026‑09‑10
Open-weight release
DeepSeek-V4.1-Flash
DeepSeek
DeepSeek publishes weights on Hugging Face for a 552B mixture-of-experts model built on a new causal encoder-decoder architecture that activates roughly 8B parameters on input and 16B on output. It reports 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE v1.1, and says the KV cache needs a quarter of the prior generation's HBM.
2026‑09‑10
Harness release
Introducing the Agents API
OpenAI
OpenAI ships the harness that runs Codex as a managed service in public beta, saying it handles session orchestration, context compaction, sub-agent coordination and recovery while developers supply tools and choose a sandbox. It offers versioned harness access alongside each model launch.
2026‑09‑10
Model release
Introducing SWE-2: Pushing the Pareto Frontier
Cognition
Cognition says it scaled reinforcement learning to a 2.8T-parameter Kimi K3 base for the first time, training every reasoning-effort level in a single run with a cost-penalized reward. It reports 50.0 on its own FrontierCode 1.1 Main at 64 percent lower cost than Fable 5.1, and 58 percent fewer turns than SWE-1.7.
2026‑09‑08
Paper
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
arXiv
The authors log predicted capability demand, chosen service tier and interaction outcome for each turn, then reuse those routing signals as a three-stage fine-tuning curriculum and an evaluation-selection-update loop. They report macro-average gains from 58.94 to 64.87 at 4B and 65.60 to 69.04 at 9B across eleven benchmarks, with weights and code released.
2026‑09‑03
Paper
EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?
arXiv
The authors hold tasks fixed and instead grow the harness across tools, skills and agents over 17 multi-stage streams covering 802 tasks, 520 tools, 42 skills and 62 agents. They report harness-induced forgetting of up to 34.7 percent backward transfer on the agents axis with model parameters unchanged.
2026‑09‑03
Model release
GPT-6 Astra: A new generation of intelligence
OpenAI
OpenAI reports 64.6 percent on Terminal-Bench Science 0.1 and 57.9 percent on Terminal-Bench 4.0, and introduces experimental Codex notes and searchable history across context windows for long-running work.
2026‑09‑02
Model update
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
Google
Google reports 54.9 percent on HLE-Verified for Gemini 3.8 Flash while retaining the introductory token prices of 3.7 Flash. It says long-running agentic loops helped evaluate and refine the models, and releases Flash Cyber through trusted-defender access.
2026‑09‑01
Model release
Introducing Claude Fable 5.1 and Claude Mythos 5.1
Anthropic
Anthropic reports 52.6 percent on Terminal-Bench-Science 0.1 versus 24.7 percent for Fable 5, and estimates that lower cache-read pricing reduces highly agentic workload costs by up to roughly 45 percent.
2026‑08‑31
Paper
WHALE: A Simple Recipe for Joint Harness-Weight Optimization
arXiv
The authors alternate model-weight updates with executable harness search, reporting gains of 4.15 to 24.38 percentage points over the tested single-component and prompt-weight baselines across three domains with Qwen3.5-2B/4B agents. Code is released for reproducing the joint optimization recipe.

Want next month's news? Watch the repository. Citation counts, rankings, and the month's top stories refresh every Monday. Browse past editions →

📈 Live Rankings

Citation counts are current through 2026-09-15 from Semantic Scholar and OpenAlex. Rankings are discovery aids, not quality scores; audit evidence remains independent of popularity. GitHub stars are snapshots from 2026-09-16. For papers indexed as multiple versions, retain the largest title-verified count reported by the configured sources. All yearly rankings use first public appearance year; a later venue year never moves a paper into a newer cohort.

🔥 Recent Papers by Average Monthly Citations

Paper Venue Date Citations Avg. cites/month Code
AlphaEvolve: A coding agent for scientific and algorithmic discovery arXiv 2025-06 848 56.5
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces arXiv 2026-01 413 51.6 GitHub · ★ 2,583
Towards End-to-End Automation of AI Research Nature 2026 2026-03 236 39.3 GitHub · ★ 7,153
Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory European Conference on Artificial Intelligence (ECAI) 2025-04 624 36.7 GitHub · ★ 65,369
Why Do Multi-Agent LLM Systems Fail? Advances in Neural Information Processing Systems 2025-03 596 33.1 GitHub · ★ 416
The Berkeley Function Calling Leaderboard (BFCL): From tool use to agentic evaluation of large language models Proceedings of the 42nd International Conference on Machine Learning 2025 435 29.0 GitHub · ★ 13,024
Memory in the Age of AI Agents arXiv preprint arXiv:2512.13564 2025-12 260 28.9
Meta-Harness: End-to-End Optimization of Model Harnesses arXiv 2026-03 170 28.3 GitHub · ★ 1,212
LLMs Get Lost In Multi-Turn Conversation International Conference on Learning Representations 2025-05 437 27.3 GitHub · ★ 297
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models arXiv 2025-10 300 27.3 GitHub · ★ 1,316

🏆 Most-Cited Papers by Year

Top 12 of 2026 by citations
Paper Date Citations Code
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
arXiv
2026-01 413 GitHub · ★ 2,583
Towards End-to-End Automation of AI Research
Nature 2026
2026-03 236 GitHub · ★ 7,153
Meta-Harness: End-to-End Optimization of Model Harnesses
arXiv
2026-03 170 GitHub · ★ 1,212
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
arXiv preprint arXiv:2604.08224
2026-04 57
Natural-Language Agent Harnesses
arXiv
2026-03 42
Self-Harness: Harnesses That Improve Themselves
arXiv
2026-06 41 GitHub · ★ 108
Hindsight Credit Assignment for Long-Horizon LLM Agents
arXiv preprint arXiv:2603.08754
2026-03 39
SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents
arXiv preprint arXiv:2601.16746
2026-01 37 GitHub · ★ 318
SkillOS: Learning Skill Curation for Self-Evolving Agents
arXiv
2026-05 35
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
arXiv
2026-03 33 GitHub · ★ 559
DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
arXiv preprint arXiv:2601.18137
2026-01 33
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
arXiv
2026-04 31 GitHub · ★ 2,112
Top 12 of 2025 by citations
Paper Date Citations Code
A-MEM: Agentic Memory for LLM Agents
Advances in Neural Information Processing Systems
2025-02 984 GitHub · ★ 968
AlphaEvolve: A coding agent for scientific and algorithmic discovery
arXiv
2025-06 848
Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
European Conference on Artificial Intelligence (ECAI)
2025-04 624 GitHub · ★ 65,369
Why Do Multi-Agent LLM Systems Fail?
Advances in Neural Information Processing Systems
2025-03 596 GitHub · ★ 416
LLMs Get Lost In Multi-Turn Conversation
International Conference on Learning Representations
2025-05 437 GitHub · ★ 297
The Berkeley Function Calling Leaderboard (BFCL): From tool use to agentic evaluation of large language models
Proceedings of the 42nd International Conference on Machine Learning
2025 435 GitHub · ★ 13,024
Group-in-Group Policy Optimization for LLM Agent Training
Advances in Neural Information Processing Systems
2025-05 410 GitHub · ★ 2,307
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
arXiv
2025-10 300 GitHub · ★ 1,316
PaperBench: Evaluating AI's ability to replicate AI research
arXiv
2025-04 274 GitHub · ★ 1,300
Memory in the Age of AI Agents
arXiv preprint arXiv:2512.13564
2025-12 260
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
arXiv preprint arXiv:2506.11763
2025-06 233 GitHub · ★ 829
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
arXiv
2025-09 230 GitHub · ★ 526
Top 12 of 2024 by citations
Paper Date Citations Code
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
NeurIPS 2024
2024-04 1158 GitHub · ★ 3,144
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
arXiv
2024-06 1143 GitHub · ★ 1,435
Self-rewarding language models
arXiv
2024-01 708
Demystifying LLM-Based Software Engineering Agents
Proceedings of the ACM on Software Engineering
2024-07 492 GitHub · ★ 2,110
TravelPlanner: A Benchmark for Real-World Planning with Language Agents
International Conference on Machine Learning
2024-02 476 GitHub · ★ 547
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
2024-01 445 GitHub · ★ 1,124
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
International Conference on Learning Representations
2024-05 430 GitHub · ★ 907
MLE-bench: Evaluating machine learning agents on machine learning engineering
ICLR 2025
2024-10 389 GitHub · ★ 1,747
HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models
Advances in Neural Information Processing Systems
2024-05 354 GitHub · ★ 4,007
When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs
Transactions of the Association for Computational Linguistics
2024-06 345
WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?
Proceedings of the 41st International Conference on Machine Learning
2024-03 343 GitHub · ★ 272
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
ACL 2024
2024-07 326 GitHub · ★ 515
Top 12 of 2023 by citations
Paper Date Citations Code
Toolformer: Language Models Can Teach Themselves to Use Tools
Advances in Neural Information Processing Systems
2023-02 5597
Generative Agents: Interactive Simulacra of Human Behavior
Proceedings of the 36th annual acm symposium on user interface software and technology
2023-04 5511 GitHub · ★ 22,110
Reflexion: language agents with verbal reinforcement learning
NeurIPS 2023
2023-03 5256 GitHub · ★ 3,267
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Advances in Neural Information Processing Systems
2023-05 4856 GitHub · ★ 6,068
Self-Refine: Iterative refinement with self-feedback
NeurIPS 2023
2023-03 4641 GitHub · ★ 820
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
ICLR 2024
2023-10 3707 GitHub · ★ 5,853
Voyager: An Open-Ended Embodied Agent with Large Language Models
Transactions on Machine Learning Research
2023-05 2280 GitHub · ★ 7,200
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
International Conference on Learning Representations
2023-07 2242 GitHub · ★ 5,741
WebArena: A Realistic Web Environment for Building Autonomous Agents
ICLR 2024
2023-07 1981 GitHub · ★ 1,609
Gorilla: Large Language Model Connected with Massive APIs
Advances in Neural Information Processing Systems
2023-05 1628 GitHub · ★ 13,024
Mind2Web: Towards a Generalist Agent for the Web
Advances in Neural Information Processing Systems
2023-06 1423 GitHub · ★ 1,027
MemGPT: Towards LLMs as Operating Systems
arXiv preprint arXiv:2310.08560
2023-10 1305 GitHub · ★ 24,757

🧪 Benchmarks

Figure 3 from the companion survey: What Is AI4AI? A Taxonomy

111 papers · Survey-curated collection, newest first. Cross-collection papers may appear in more than one section.

Paper Date Citations Code
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
arXiv
2026-08 5
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
arXiv preprint arXiv:2608.06301
2026-08 1
When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories
arXiv preprint arXiv:2608.06057
2026-08 0
DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks
arXiv preprint arXiv:2607.07946
2026-07 17 GitHub · ★ 1,682
ChainSWE: Benchmarking Coding Agents on Multi-Bug Software Maintenance
arXiv preprint arXiv:2607.02606
2026-07 1
Kimi K3: Open Frontier Intelligence
arXiv
2026-07 27 GitHub · ★ 8,798
Can AI Agents Conduct Open-Ended AI Research? Early Evidence from Two Case Studies
arXiv
2026-07 6
RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement
arXiv
2026-07 9 GitHub · ★ 155
Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI
arXiv preprint arXiv:2607.22368
2026-07 4
Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
arXiv preprint arXiv:2607.14004
2026-07 2 GitHub · ★ 8
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
arXiv
2026-07 3 GitHub · ★ 689
FrontierSWE
Proximal Blog
2026 GitHub · ★ 229
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
arXiv preprint arXiv:2606.09426
2026-06 5
Do LLMs Catch Their Own Mistakes? A Comprehensive Benchmark for Reflective Tool Use LLMs
Findings of the Association for Computational Linguistics: ACL 2026
2026 0
The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?
arXiv preprint arXiv:2606.04455
2026-06 4 GitHub · ★ 22
SWE-Explore: Benchmarking How Coding Agents Explore Repositories
arXiv
2026-06 14 GitHub · ★ 42
SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code
arXiv
2026-06 0
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
arXiv
2026-06 9
FARS: A Fully Automated Research System Deployed at Scale
arXiv
2026-06 4
DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration
arXiv preprint arXiv:2606.03103
2026-06 2 GitHub · ★ 92
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
arXiv
2026-06 2 GitHub · ★ 120
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
arXiv
2026-06 15 GitHub · ★ 310
MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI
arXiv preprint arXiv:2605.08678
2026-05 8 GitHub · ★ 117
ProgramBench: Can Language Models Rebuild Programs From Scratch?
arXiv
2026-05 25 GitHub · ★ 925
RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades
arXiv
2026-05 2 GitHub · ★ 16
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
arXiv
2026-05 4 GitHub · ★ 71
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
arXiv
2026-05 5 GitHub · ★ 15
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle
arXiv
2026-05 2 GitHub · ★ 0
Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution
arXiv
2026-05 1 GitHub · ★ 5
How Far Are We From True Auto-Research?
arXiv
2026-05 6
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
arXiv preprint arXiv:2604.10547
2026-04 3 GitHub · ★ 14,644
Toward autonomous long-horizon engineering for ML research
arXiv
2026-04 10 GitHub · ★ 145
KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation
arXiv preprint arXiv:2604.08455
2026-04 20 GitHub · ★ 76
CI-Repair-Bench: A Repository-Aware Benchmark for Automated Patch Validation via CI Workflows
arXiv
2026-04 1 GitHub · ★ 1
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
arXiv
2026-04 1 GitHub · ★ 2
AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery
arXiv
2026-04 16 GitHub · ★ 680
The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
arXiv preprint arXiv:2604.11978
2026-04 22
SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution
International Conference on Machine Learning
2026-03 6 GitHub · ★ 72
Beyond pass@1: A Reliability Science Framework for Long-Horizon LLM Agents
arXiv preprint arXiv:2603.29231
2026-03 5
Towards End-to-End Automation of AI Research
Nature 2026
2026-03 236 GitHub · ★ 7,153
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
arXiv
2026-03 33 GitHub · ★ 559
ReCUBE: Evaluating Repository-Level Context Utilization in Code Generation
arXiv
2026-03 1 GitHub · ★ 1
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
arXiv
2026-03 17 GitHub · ★ 177
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
arXiv
2026-02 27 GitHub · ★ 92
LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
arXiv
2026-02 24 GitHub · ★ 47
AIRS-Bench: A Suite of Tasks for Frontier AI Research Science Agents
arXiv
2026-02 22 GitHub · ★ 117
SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale
arXiv preprint arXiv:2602.23866
2026-02 14
LUMINA: Long-horizon Understanding for Multi-turn Interactive Agents
Findings of the Association for Computational Linguistics: ACL 2026
2026-01 1
RepoGenesis: Benchmarking End-to-End Microservice Generation from Readme to Repository
arXiv
2026-01 6 GitHub · ★ 101
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
arXiv
2026-01 413 GitHub · ★ 2,583
ARC: Active and Reflection-driven Context Management for Long-Horizon Information Seeking Agents
Findings of the Association for Computational Linguistics: ACL 2026
2026-01 5
DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
arXiv preprint arXiv:2601.18137
2026-01 33
Toward ultra-long-horizon agentic science: Cognitive accumulation for machine learning engineering
arXiv
2026-01 24
Towards a Science of Scaling Agent Systems
arXiv
2025-12 122 GitHub · ★ 53
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
International Conference on Learning Representations
2025-12 13 GitHub · ★ 41
NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
arXiv
2025-12 46 GitHub · ★ 176
SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
arXiv
2025-12 42 GitHub · ★ 58
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
arXiv preprint arXiv:2510.11977
2025-10 65 GitHub · ★ 311
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
arXiv
2025-09 1
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
arXiv
2025-09 230 GitHub · ★ 526
HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds
arXiv preprint arXiv:2508.12782
2025-08 6 GitHub · ★ 14
Hell or High Water: Evaluating Agentic Recovery from External Failures
Second Conference on Language Modeling
2025-08 6 GitHub · ★ 5
OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
arXiv preprint arXiv:2508.09124
2025-08 52 GitHub · ★ 17
Evaluation and Benchmarking of LLM Agents: A Survey
Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2
2025-07 201 GitHub · ★ 21
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations
2025-07 6
The Berkeley Function Calling Leaderboard (BFCL): From tool use to agentic evaluation of large language models
Proceedings of the 42nd International Conference on Machine Learning
2025 435 GitHub · ★ 13,024
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
arXiv preprint arXiv:2506.11763
2025-06 233 GitHub · ★ 829
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
arXiv preprint arXiv:2506.21506
2025-06 71 GitHub · ★ 114
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
Advances in Neural Information Processing Systems
2025-06 33 GitHub · ★ 78
ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering
Advances in Neural Information Processing Systems
2025-06 29 GitHub · ★ 220
ML-Master: Towards AI-for-AI via integration of exploration and reasoning
arXiv
2025-06 55 GitHub · ★ 450
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
arXiv preprint arXiv:2505.19955
2025-05 55 GitHub · ★ 36
LLMs Get Lost In Multi-Turn Conversation
International Conference on Learning Representations
2025-05 437 GitHub · ★ 297
SWE-bench Goes Live!
Advances in Neural Information Processing Systems
2025-05 60 GitHub · ★ 239
PaperBench: Evaluating AI's ability to replicate AI research
arXiv
2025-04 274 GitHub · ★ 1,300
Why Do Multi-Agent LLM Systems Fail?
Advances in Neural Information Processing Systems
2025-03 596 GitHub · ★ 416
A Survey on Evaluation of LLM-based Agents
Findings of the Association for Computational Linguistics: ACL 2026
2025-03 224
Robotouille: An Asynchronous Planning Benchmark for LLM Agents
International Conference on Learning Representations
2025-02 38 GitHub · ★ 48
DI-BENCH: Benchmarking Large Language Models on Dependency Inference with Testable Repositories at Scale
Findings of the Association for Computational Linguistics: ACL 2025
2025-01 8 GitHub · ★ 7
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track
2024-12 303 GitHub · ★ 780
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
arXiv
2024-11 145 GitHub · ★ 162
MLE-bench: Evaluating machine learning agents on machine learning engineering
ICLR 2025
2024-10 389 GitHub · ★ 1,747
Agent-as-a-Judge: Evaluate Agents with Agents
Forty-second International Conference on Machine Learning
2024-10 218 GitHub · ★ 824
Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
International Conference on Machine Learning
2024-09 204 GitHub · ★ 898
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Findings of the Association for Computational Linguistics: NAACL 2025
2024-08 256 GitHub · ★ 283
OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation
arXiv preprint arXiv:2407.19056
2024-07 54 GitHub · ★ 46
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
ACL 2024
2024-07 326 GitHub · ★ 515
Introducing SWE-bench Verified
2024
WebCanvas: Benchmarking Web Agents in Online Environments
ICML 2024 Workshop on Agentic Markets
2024-06 124 GitHub · ★ 281
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
arXiv
2024-06 1143 GitHub · ★ 1,435
NATURAL PLAN: Benchmarking LLMs on Natural Language Planning
arXiv preprint arXiv:2406.04520
2024-06 139 GitHub · ★ 59
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
International Conference on Learning Representations
2024-05 430 GitHub · ★ 907
Benchmarking Mobile Device Control Agents across Diverse Configurations
arXiv preprint arXiv:2404.16660
2024-04 49 GitHub · ★ 33
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
NeurIPS 2024
2024-04 1158 GitHub · ★ 3,144
WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?
Proceedings of the 41st International Conference on Machine Learning
2024-03 343 GitHub · ★ 272
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Computer Vision -- ECCV 2024
2024-02 169
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
International Conference on Machine Learning
2024-02 186 GitHub · ★ 163
TravelPlanner: A Benchmark for Real-World Planning with Language Agents
International Conference on Machine Learning
2024-02 476 GitHub · ★ 547
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
2024-01 445 GitHub · ★ 1,124
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
2024-01 41 GitHub · ★ 486
MinePlanner: A Benchmark for Long-Horizon Planning in Large Minecraft Worlds
Proceedings of the 6th ICAPS Workshop on the International Planning Competition (WIPC)
2023-12 8 GitHub · ★ 23
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
ICLR 2024
2023-10 3707 GitHub · ★ 5,853
AgentBench: Evaluating LLMs as Agents
International Conference on Learning Representations
2023-08 1279 GitHub · ★ 3,736
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
International Conference on Learning Representations
2023-07 2242 GitHub · ★ 5,741
WebArena: A Realistic Web Environment for Building Autonomous Agents
ICLR 2024
2023-07 1981 GitHub · ★ 1,609
Mind2Web: Towards a Generalist Agent for the Web
Advances in Neural Information Processing Systems
2023-06 1423 GitHub · ★ 1,027
BEHAVIOR-1K: A benchmark for embodied AI with 1,000 everyday activities and realistic simulation
Proceedings of The 6th Conference on Robot Learning
2023 382 GitHub · ★ 1,700
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Advances in Neural Information Processing Systems
2022-07 1359 GitHub · ★ 594
WebGPT: Browser-assisted question-answering with human feedback
arXiv preprint arXiv:2112.09332
2021-12 2066
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
2019-12 1192 GitHub · ★ 531
World of Bits: An Open-Domain Platform for Web-Based Agents
Proceedings of the 34th International Conference on Machine Learning
2017 352

🛠️ Harness Design

Evidence chain for reliable harness interventions

101 papers · Survey-curated collection, newest first. Cross-collection papers may appear in more than one section.

Paper Date Citations Code
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
arXiv
2026-08 0
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
arXiv
2026-08 0 GitHub · ★ 201
Prime Agent: A Self-Improving RLM Harness
arXiv
2026-08 0 GitHub · ★ 20,869
Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
2026-07 11
Kimi K3: Open Frontier Intelligence
arXiv
2026-07 27 GitHub · ★ 8,798
Recursive harness self-improvement
arXiv
2026-07 10
ACM: Agentic Context Management for Long Horizon Tasks
arXiv
2026-07 1 GitHub · ★ 36
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents
arXiv
2026-07 2
Structured Feedback Improves Repair in an LLM Agent Loop
arXiv
2026-07 1
Self-Improvements in Modern Agentic Systems: A Survey
2026-07 9 GitHub · ★ 486
MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution
arXiv
2026-07 9
Rethinking the Evaluation of Harness Evolution for Agents
arXiv
2026-07 19 GitHub · ★ 37
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
arXiv
2026-07 3 GitHub · ★ 689
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
arXiv
2026-06 25
The Past Is Prologue: A Plug-in Controller for Selective Updates in Sequentially Evolving LLM Memory
arXiv
2026-06 1
From Question Answering to Task Completion: A Survey on Agent System and Harness Design
arXiv
2026-06 4
Agent Harness for Large Language Model Agents: A Survey
Preprints
2026 2 GitHub · ★ 353
Scaffold Effects on GAIA: A Controlled Comparison
arXiv
2026-06 0
Context Compression for LLM Agents: A Survey of Methods, Failure Modes, and Evaluation
Preprints
2026 0
The Verification Horizon: No Silver Bullet for Coding Agent Rewards
arXiv
2026-06 4
Self-Harness: Harnesses That Improve Themselves
arXiv
2026-06 41 GitHub · ★ 108
Stop Comparing LLM Agents Without Disclosing the Harness
Second Workshop on Agents in the Wild: Safety, Security, and Beyond
2026 8
Are We Ready For An Agent-Native Memory System?
arXiv
2026-06 12 GitHub · ★ 147
MOSS: Self-Evolution through Source-Level Rewriting in Autonomous Agent Systems
arXiv
2026-05 7 GitHub · ★ 22
Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?
arXiv
2026-05 2
From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills
arXiv
2026-05 15
Code as Agent Harness
arXiv
2026-05 25 GitHub · ★ 698
SkillOS: Learning Skill Curation for Self-Evolving Agents
arXiv
2026-05 35
Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents
arXiv
2026-05 21 GitHub · ★ 222
LoopTrap: Termination Poisoning Attacks on LLM Agents
arXiv
2026-05 1
Learning Agent-Compatible Context Management for Long-Horizon Tasks
arXiv
2026-05 5
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
arXiv
2026-05 14 GitHub · ★ 14
Toward autonomous long-horizon engineering for ML research
arXiv
2026-04 10 GitHub · ★ 145
Squeez: Task-Conditioned Tool-Output Pruning for Coding Agents
arXiv preprint arXiv:2604.04979
2026-04 2 GitHub · ★ 23
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
arXiv
2026-04 31 GitHub · ★ 2,112
Escher-Loop: Mutual Evolution by Closed-Loop Self-Referential Optimization
arXiv
2026-04 6 GitHub · ★ 8
Reinforced Agent: Inference-Time Feedback for Tool-Calling Agents
Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM)
2026-04 4
From Agent Loops to Structured Graphs:A Scheduler-Theoretic Framework for LLM Agent Execution
arXiv
2026-04 2
ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents
arXiv preprint arXiv:2604.01664
2026-04 10 GitHub · ★ 9
ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents
arXiv preprint arXiv:2604.23069
2026-04 2
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
arXiv preprint arXiv:2604.08224
2026-04 57
Meta-Harness: End-to-End Optimization of Model Harnesses
arXiv
2026-03 170 GitHub · ★ 1,212
Natural-Language Agent Harnesses
arXiv
2026-03 42
Bilevel Autoresearch: Meta-Autoresearching Itself
arXiv
2026-03 4 GitHub · ★ 184
Schema First Tool APIs for LLM Agents: A Controlled Study of Tool Misuse, Recovery, and Budgeted Performance
arXiv
2026-03 4 GitHub · ★ 0
DARWIN: Dynamic Agentically Rewriting Self-Improving Network
arXiv
2026-02 1 GitHub · ★ 0
SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents
arXiv preprint arXiv:2601.16746
2026-01 37 GitHub · ★ 318
Beyond Static Summarization: Proactive Memory Extraction for LLM Agents
arXiv preprint arXiv:2601.04463
2026-01 16
Memory in the Age of AI Agents
arXiv preprint arXiv:2512.13564
2025-12 260
Step-DeepResearch Technical Report
arXiv preprint arXiv:2512.20491
2025-12 13 GitHub · ★ 573
Towards a Science of Scaling Agent Systems
arXiv
2025-12 122 GitHub · ★ 53
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
International Conference on Learning Representations
2025-12 13 GitHub · ★ 41
PARC: An Autonomous Self-Reflective Coding Agent for Robust Execution of Long-Horizon Tasks
arXiv
2025-12 3
Solving a Million-Step LLM Task with Zero Errors
arXiv preprint arXiv:2511.09030
2025-11 23 GitHub · ★ 46
LongCodeZip: Compress Long Context for Code Language Models
2025 40th IEEE/ACM International Conference on Automated Software Engineering (ASE)
2025-10 46 GitHub · ★ 168
ACON: Optimizing Context Compression for Long-horizon LLM Agents
arXiv preprint arXiv:2510.00615
2025-10 92 GitHub · ★ 114
Scaling Long-Horizon LLM Agent via Context-Folding
arXiv preprint arXiv:2510.11967
2025-10 108 GitHub · ★ 187
AgentFold: Long-Horizon Web Agents with Proactive Context Management
arXiv preprint arXiv:2510.24699
2025-10 75 GitHub · ★ 19,948
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
arXiv
2025-10 300 GitHub · ★ 1,316
WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research
arXiv preprint arXiv:2509.13312
2025-09 43 GitHub · ★ 19,948
ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
arXiv preprint arXiv:2509.13313
2025-09 110 GitHub · ★ 19,948
Reducing Cost of LLM Agents with Trajectory Reduction
Proceedings of the ACM on Software Engineering
2025-09 40
Where LLM Agents Fail and How They can Learn From Failures
arXiv
2025-09 115 GitHub · ★ 106
Memp: Exploring Agent Procedural Memory
Findings of the Association for Computational Linguistics: ACL 2026
2025-08 65 GitHub · ★ 38
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
arXiv preprint arXiv:2508.21433
2025-08 29
Magentic-UI: Towards Human-in-the-loop Agentic Systems
arXiv
2025-07 50 GitHub · ★ 10,091
ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
The Fourteenth International Conference on Learning Representations
2025-06 14
SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling
Findings of the Association for Computational Linguistics: ACL 2025
2025-06 27 GitHub · ★ 66
MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
arXiv preprint arXiv:2506.15841
2025-06 197 GitHub · ★ 336
Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments
Findings of the Association for Computational Linguistics: EMNLP 2025
2025-05 11 GitHub · ★ 2
Is there a half-life for the success rates of AI agents?
arXiv preprint arXiv:2505.05115
2025-05 5
Darwin Godel Machine: Open-ended evolution of self-improving agents
arXiv
2025-05 220 GitHub · ★ 2,335
Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
European Conference on Artificial Intelligence (ECAI)
2025-04 624 GitHub · ★ 65,369
Process Reward Models That Think
Transactions on Machine Learning Research
2025-04 110 GitHub · ★ 92
Why Do Multi-Agent LLM Systems Fail?
Advances in Neural Information Processing Systems
2025-03 596 GitHub · ★ 416
Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents
The Thirteenth International Conference on Learning Representations
2025-03 22 GitHub · ★ 7
Process Reward Models for LLM Agents: Practical Framework and Directions
arXiv
2025-02 83 GitHub · ★ 60
A-MEM: Agentic Memory for LLM Agents
Advances in Neural Information Processing Systems
2025-02 984 GitHub · ★ 968
Practical Considerations for Agentic LLM Systems
arXiv
2024-12 17
Godel Agent: A self-referential agent framework for recursive self-improvement
arXiv
2024-10 23 GitHub · ★ 221
Agent-as-a-Judge: Evaluate Agents with Agents
Forty-second International Conference on Machine Learning
2024-10 218 GitHub · ★ 824
Agent Workflow Memory
Forty-second International Conference on Machine Learning
2024-09 275 GitHub · ★ 471
Automated design of agentic systems
ICLR 2025
2024-08 316 GitHub · ★ 1,639
LLM Critics Help Catch LLM Bugs
arXiv preprint arXiv:2407.00215
2024-07 167
Demystifying LLM-Based Software Engineering Agents
Proceedings of the ACM on Software Engineering
2024-07 492 GitHub · ★ 2,110
When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs
Transactions of the Association for Computational Linguistics
2024-06 345
HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models
Advances in Neural Information Processing Systems
2024-05 354 GitHub · ★ 4,007
JARVIS-1: Open-World Multi-Task Agents With Memory-Augmented Multimodal Language Models
IEEE Transactions on Pattern Analysis & Machine Intelligence
2023-11 216 GitHub · ★ 415
Large Language Models Cannot Self-Correct Reasoning Yet
International Conference on Learning Representations
2023-10 1181
MemGPT: Towards LLMs as Operating Systems
arXiv preprint arXiv:2310.08560
2023-10 1305 GitHub · ★ 24,757
Self-Taught Optimizer (STOP): Recursively self-improving code generation
Conference on Language Modeling
2023-10 134 GitHub · ★ 53
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
International Conference on Machine Learning
2023-10 633 GitHub · ★ 862
Cognitive Architectures for Language Agents
Transactions on Machine Learning Research
2023-09 516
ExpeL: LLM Agents Are Experiential Learners
AAAI 2024
2023-08 896 GitHub · ★ 243
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
International Conference on Learning Representations
2023-05 917 GitHub · ★ 745
AdaPlanner: Adaptive Planning from Feedback with Language Models
Advances in Neural Information Processing Systems
2023-05 249 GitHub · ★ 128
Voyager: An Open-Ended Embodied Agent with Large Language Models
Transactions on Machine Learning Research
2023-05 2280 GitHub · ★ 7,200
Generative Agents: Interactive Simulacra of Human Behavior
Proceedings of the 36th annual acm symposium on user interface software and technology
2023-04 5511 GitHub · ★ 22,110
Self-Refine: Iterative refinement with self-feedback
NeurIPS 2023
2023-03 4641 GitHub · ★ 820
Reflexion: language agents with verbal reinforcement learning
NeurIPS 2023
2023-03 5256 GitHub · ★ 3,267
ReAct: Synergizing Reasoning and Acting in Language Models
International Conference on Learning Representations (ICLR)
2022-10 11171 GitHub · ★ 4,171

🧠 Model Design

Model-side interventions across plan, execute, feedback, and repair

26 papers · Survey-curated collection, newest first. Cross-collection papers may appear in more than one section.

Paper Date Citations Code
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
arXiv preprint arXiv:2607.07508
2026-07 13
Kimi K3: Open Frontier Intelligence
arXiv
2026-07 27 GitHub · ★ 8,798
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
arXiv
2026-07 3 GitHub · ★ 689
Autodata: An Agentic Data Scientist to Create High Quality Synthetic Data
arXiv
2026-06 7
TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration
arXiv preprint arXiv:2604.14116
2026-04 3
CAPO: Critic-Guided Action-Aligned Policy Optimization for Advancing LLM Agent Capabilities
arXiv preprint arXiv:2604.18401
2026-04 9 GitHub · ★ 1,668
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
arXiv
2026-03 33 GitHub · ★ 559
Hindsight Credit Assignment for Long-Horizon LLM Agents
arXiv preprint arXiv:2603.08754
2026-03 39
ASI-Evolve: AI Accelerates AI
arXiv
2026-03 5 GitHub · ★ 859
Towards Execution-Grounded Automated AI Research
arXiv
2026-01 13 GitHub · ★ 84
IterResearch: Rethinking Long-Horizon Agents with Interaction Scaling
arXiv preprint arXiv:2511.07327
2025-11 17 GitHub · ★ 19,948
Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
arXiv preprint arXiv:2511.20718
2025-11 5 GitHub · ★ 0
SALT: Step-level Advantage Assignment for Long-horizon Agents via Trajectory Graph
Findings of the Association for Computational Linguistics: EACL 2026
2025-10 20
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
arXiv preprint arXiv:2509.08755
2025-09 66 GitHub · ★ 865
AlphaEvolve: A coding agent for scientific and algorithmic discovery
arXiv
2025-06 848
Group-in-Group Policy Optimization for LLM Agent Training
Advances in Neural Information Processing Systems
2025-05 410 GitHub · ★ 2,307
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
International Conference on Machine Learning
2025-03 211 GitHub · ★ 46
Reinforcement Learning for Long-Horizon Interactive LLM Agents
arXiv preprint arXiv:2502.01600
2025-02 103
Self-rewarding language models
arXiv
2024-01 708
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
International Conference on Machine Learning
2023-10 633 GitHub · ★ 862
Gorilla: Large Language Model Connected with Massive APIs
Advances in Neural Information Processing Systems
2023-05 1628 GitHub · ★ 13,024
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Advances in Neural Information Processing Systems
2023-05 4856 GitHub · ★ 6,068
Toolformer: Language Models Can Teach Themselves to Use Tools
Advances in Neural Information Processing Systems
2023-02 5597
ReAct: Synergizing Reasoning and Acting in Language Models
International Conference on Learning Representations (ICLR)
2022-10 11171 GitHub · ★ 4,171
STaR: Bootstrapping reasoning with reasoning
NeurIPS 2022
2022-03 1058
Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models
Advances in Neural Information Processing Systems
2022-01 21886 GitHub · ★ 41

📚 How We Curate

Search, screening, extraction, coding, and verification workflow

  1. Discover: survey searches, backward/forward citation chaining, and community suggestions.
  2. Verify: exact title and identifier checks against primary scholarly sources.
  3. Classify: Benchmarks, Harness Design, and Model Design, allowing justified overlap.
  4. Audit when evidence permits: stage ownership plus independent G/R/H/T coordinates.
  5. Refresh weekly: citation counts and rankings every Monday, alongside the month's top papers, releases, blogs, and research news.

📄 Citation

If this map or its evidence audit helps your work, please cite the companion survey:

Copy BibTeX
@misc{wu2026eveai4ai,
  title   = {{AI4AI} Survey: From Long-Horizon Agents to Recursive Self-Improvement---Definitions, Reliable Horizons, and Open Problems},
  author  = {Wu, Kai and Lyu, Hao and Luo, Zhen and Wang, Chaofan and
             Ye, Siyu and Lin, Jinghao and Ji, Xiaozhong and Jiang, Boyuan and
             Wang, Shengzhi and Wang, Zihan and Ye, Yiwen and Wang, Hao and
             Wang, Zimu and Liu, Wenzhe and Wang, Ruobing and Cai, Kai and
             Xiong, Mingliang and Fang, Wen and Liu, Mingqing and
             Zhang, Yifan and Yang, Lei and Hu, Xiaobin and Liu, Qingwen},
  howpublished = {Preprints.org},
  year    = {2026},
  doi     = {10.20944/preprints202608.2108.v1},
  url     = {https://doi.org/10.20944/preprints202608.2108.v1}
}

🤝 Contributing

Have something to add? Missing a paper, official code link, or stronger primary-source evidence? Open a pull request or use the paper-suggestion form. Read the contribution guide →

About

AI4AI Survey: can AI reliably improve AI? 223 papers on long-horizon agents, benchmarks, harness design, and recursive self-improvement · updated weekly

Topics

Resources

Contributing

Stars

167 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages