<antirez>
6 hours ago
- The web is steadily losing old pages, with a fraction disappearing every year, resulting in permanent loss of digital content.
- The Internet Archive is a critical institution for preserving web history, akin to a sacred place, but faces increasing challenges from companies and entities.
- Lost content includes early programming and hacker culture, 1990s subcultures, personal blogs, scientific papers, early digital art, video games, climate data, and news sources.
- The effort to preserve everything is impractical due to high costs and lack of economic incentive, making LLMs a viable lossy but accessible compression tool.
- To mitigate loss, we must ensure publicly released LLM weights are preserved and that the Internet Archive is included in pre-training datasets, while continuing to support preservation institutions.