OpenAI and Microsoft knew they were starting a 'doom loop' for the web
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
OpenAI and Microsoft’s training practices risk creating a “doom loop” by degrading the web’s quality as LLMs consume and regurgitate content, reducing the availability of original, high-quality data. This directly impacts production systems by accelerating the depletion of unique training datasets, forcing engineers to rely on increasingly synthetic or redundant data, which could degrade model performance over time and require costly mitigation strategies.
Internal Microsoft/OpenAI documents now in the NYT lawsuit record describe their own scraping as the "largest theft of labor in human history" and a "doom loop" that "makes a complete mockery of fair use"—their words, not a plaintiff's characterization. If you're building on these models or scraping data yourself, expect the fair-use defense to weaken and training-data provenance to become a real legal and contractual liability, not a theoretical one.