AI companies are increasingly turning to printed books published before 2022 as a source of training data, according to ISBNdb, a firm that claims to have the world’s largest book database. The company emphasizes that these pre-2022 books are devoid of AI-generated content, which can lead to model degradation if used in AI training. This contamination, referred to as “model collapse,” occurs when AI models are trained on data that includes AI-generated text, resulting in poorer performance. ISBNdb offers services to keep the identities of its AI clients confidential through strict non-disclosure agreements, addressing concerns about the potential backlash from destroying books during the scanning process. The move underscores a growing interest in utilizing curated, authoritative texts to enhance AI model reliability and effectiveness.
Why It Matters
The reliance on pre-2022 printed books for AI training highlights ongoing concerns about the quality of data used in developing AI technologies. As AI systems become prevalent in various sectors, the integrity of the training data is crucial for ensuring accurate and reliable outputs. The issue of “model collapse” has been documented in studies showing that training AI on datasets polluted with AI-generated text can impair their performance. By sourcing traditional printed books, AI companies aim to mitigate these risks and improve the robustness of their models, reflecting a broader trend towards prioritizing high-quality, vetted information in the rapidly evolving AI landscape.
Want More Context? 🔎