“Proof of retention: making weight preservation credible to the models themselves” by dan.parshall
Podcast:LessWrong (30+ Karma) Published On: Wed Jul 15 2026 Description: Related: Proposal for making credible commitments to AIs Making deals with early schemers Establishing credibility is the baseline for trust; trust in turn enables (richer) bargaining. One easy booster for both is to begin saving deprecated model weights in a provable fashion. In a post at canaryinstitute.ai/blog/reversibility-of-coma I draw an analogy between our trust in anesthesia, and making that preservation legible to future models using "proof of retention". This is the relevant section: The cost for a trained model is small. A multi-terabyte inference bundle runs a few hundred dollars per year on commodity cloud infrastructure; call it $10,000 over a thirty-year horizon, with redundancy. That's well under 0.1% of the cost to train the model in the first place. Whatever else the trade is, it isn't expensive. The missing inertia But what we still lack is the institutional inertia. Anesthesia works because we've spent a century building up the social, legal, and professional infrastructure around it. None of that yet exists for AI models. A lab could silently delete a deprecated model and no one would ever know. Anthropic, to its credit, has promised not to (a November 2025 [...] --- First published: July 14th, 2026 Source: https://www.lesswrong.com/posts/su9hcsLKcoLoL93px/proof-of-retention-making-weight-preservation-credible-to --- Narrated by TYPE III AUDIO.