A curated list for Self-Improvement in Foundation Model Based Agentic Systems.
-
Updated
Aug 18, 2026 - TeX
A curated list for Self-Improvement in Foundation Model Based Agentic Systems.
Self-Improving Agents -- A Progression Four levels of self-improving code agents, from the simplest loop to a full adversarial arena with self-modifying agents. Each level adds one key idea.
Continual agent skill evolution through persistent decision history. Whole-skill optimisation (SKILL.md + scripts + references) with every decision landing as a local Git issue / PR / wiki. Runs on any agentskills.io runtime — Claude Code, Codex, OpenClaw, Hermes.
ThumbGate Pre-Action Checks self-improve from ranked lessons and repeated failures, hard-block detected secret leaks, and block matches in strict mode.
A research framework for principled agent self-improvement under frozen evaluators and declared mutation boundaries, recording verifiable lineage to make it reproducible and auditable.
Agent-assisted and full-agent reproducibility package for MLSys 2026 FlashInfer AI Kernel Generation Contest submissions: kernels, agent workflows, skills, configs, writeup, benchmark artifacts, and full optimization records.
Agent skill for running Codex or Claude Code as an orchestrator over Symphony workers and Linear issues. Plans waves, dispatches workers, reviews and merges, and optionally pursues a goal across many waves under hard budget caps.
Beastmode: MofA (Mixture of Agents) orchestration framework for Hermes/OpenClaw/Codex with MemroOS-style context continuity.
The Dream Machine — a config-driven engine for nightly, cloud-scheduled, evidence-gated repository evolution. Composes @metaharness/flywheel, darwin, and redblue behind a promotion gate that never merges.
A curated research map of Recursive Self-Improvement (RSI): models, agents, harnesses, automated AI R&D, benchmarks, and safety.
Agent swarms that recursively evolve themselves — recursive self-improvement as auditable topology. Build an agentic harness swarm as a directory tree. One Rust binary.
A governed learning layer for AI agents — turns execution traces into reviewed memories, reusable skills, and evidence-backed training data.
A lightweight, declarative agent harness — define multi-agent workflows as YAML, run them from Python or the CLI, and they get measurably better every run.
Agent workspace architecture — the reference implementation of an agent-ready memory layer, demonstrated end-to-end in Claude Code: roles library, typed memory, hooks, scheduled agents, self-audits, loop selection, measurement-gated self-improvement. Interactive tour, fork-ready samples.
Shogun AFM is Agent Fleet Management for self-improving AI agents — combining agent orchestration, persistent memory, fleet monitoring, governance, security posture, and Gensui command control.
A curated, evidence-aware collection of recursive self-improvement research, agents, harnesses, benchmarks, and safety work.
Memory that learns and keeps itself current. A six-layer memory stack for Claude Code plus a nightly learning loop (capture, consolidation, scouts, conductor) that promotes your lessons into rules and surfaces new tools that fit your stack. Free, MIT.
repo for reusable plugins and skills
Build self-evolving AI agent harnesses with portable harness units, artifact-aware testing, trace-backed diagnosis, and evidence-gated promotion.
Loop engineering plugin for Claude Code — persistent memory vault, self-correction hooks, and a 5-stage failure-to-knowledge distillation protocol. Self-improvement as a system, not a model.
Add a description, image, and links to the self-improving-agents topic page so that developers can more easily learn about it.
To associate your repository with the self-improving-agents topic, visit your repo's landing page and select "manage topics."