IRSI Institute for Recursive Self Improvement
First project · Protocol in development

IRSI-Bench

Measuring how well models improve AI systems—and the processes that produce their next improvements.

Model comparisons

Improvement · Transfer · Cost · Reliability

Results forthcoming. The first release will compare models under a common evaluation protocol.

What should improve?

A model can produce a better agent without becoming better at producing agents. IRSI-Bench is being designed to distinguish those capabilities.

Task performanceWhat the initial agent can do.
ImprovementWhether its changes help on new tasks.
Recursive gainWhether an evolved improvement process produces better successors than the original process.
ResourcesTime and cost across proposals, tests, failures, and evaluation.