Authorization for Self-Improving Agents: Governing Self-Modification at the Capability Boundary
Authors:
Abstract
For three years, the enterprise security community has worked on a well-posed question: given an autonomous agent, which actions may it take, on which resources, under which conditions. The industry has answered it competently. The Model Context Protocol now treats servers as OAuth 2.1 resource servers with audience-bound tokens, the Agent2Agent protocol gives agents discoverable identities, and policy engines mediate calls at runtime. These are real advances, and they share one assumption: the agent is a fixed artifact whose set of possible actions is enumerable and authored by a person. That assumption is being retired in production. Recursive self-improvement has moved from thought experiment to deployed practice. Systems now rewrite their own prompts, edit their own tool configurations, curate their own memory, and, in research settings, modify the code and weights of their own successors. When an agent can change itself, the security question changes with it. The action set tomorrow is not the action set today, the change was written by the agent itself, and the trust boundary has become dynamic.
To read the file of this research, you can view or download it directly from our repository.