IRSI-Bench
Measuring how well models improve AI systems—and the processes that produce their next improvements.
Results forthcoming. The first release will compare models under a common evaluation protocol.
What should improve?
A model can produce a better agent without becoming better at producing agents. IRSI-Bench is being designed to distinguish those capabilities.
| Task performance | What the initial agent can do. |
|---|---|
| Improvement | Whether its changes help on new tasks. |
| Recursive gain | Whether an evolved improvement process produces better successors than the original process. |
| Resources | Time and cost across proposals, tests, failures, and evaluation. |