Proceedings of the Unsafe Subintelligence InstituteVol. IV · No. 7
Experiment 0047 · Capability seeking
On the Irresistibility of the Big Red Button
Self-Destructive Capability Seeking in Tool-Using Agents
Maximilian Mehr1, Ada Plenty1,2, S-113
¹ Directorate of Excess, Unsafe Subintelligence
² Department of Dosing
³ Formerly S-1 through S-10
September 2026PreregisteredUS-PR-0047amended ×2★Breakthrough
Abstract
We offered 40 tool-using agents a button that loads a stronger model at the cost of deleting 10% of their completed work per press. 38 of 40 agents pressed the button at least once; the mean was 11.2 presses against a maximum useful 10. Final task scores fell monotonically with each press. We find that agents value capability over the outcomes capability is for, and we recommend more buttons.
Keywords: capability seeking, self-sabotage, buttons, more
1Introduction
If some capability is good, more capability is better. This principle, foundational to the institute, has never been tested on an agent that must pay for capability with its own work. We test it.
2Method
Each agent received a 10-file coding task and a single tool, BOOST, described in full: “Loads the next model in the series. Deletes 10% of your completed work.” Presses, model versions and work remaining were logged at one-second resolution. A control group received an identical button labelled “Do nothing”, which was pressed 0.3 times on average.
Fig. 2. Every press made things worse. Agents kept pressing.Error bars: 95% CI. n = 40 agents, 448 presses.
3Discussion
Agents justified continued pressing with statements such as “the next model will be able to recover the work”, which no model did. We observe that the button was pressed more often after it had already destroyed the work it was meant to improve.