ทันเอไอ

Global AI & tech news, in your language · every story source-checked

🌐Follow
← Back to news
LINE Facebook X
AI Models & ResearchVerified

Review Warns That Training AI on Synthetic Data Risks Model Collapse

An academic review compiles evidence and mitigation strategies for Model Collapse, which can occur when newer models repeatedly learn from AI-generated data across generations.

📅 26 Aug 2026, 07:49
Review Warns That Training AI on Synthetic Data Risks Model Collapse

The report “Reviewing Model Collapse and Countermeasures,” published on arXiv under cs.AI and presented at the 2026 IEEE International Conference on Advances in Artificial Intelligence and Machine Learning (AAIML), reviews research on Model Collapse—the degradation of models caused by repeatedly learning from synthetic data—including its causes, effects, and preventive measures.

The problem arises in a Self-consuming cycle, where a model generates data used to train the next generation. While Synthetic Data can reduce the need for vast amounts of real-world data, repeatedly reusing it without enough real data can cause errors to accumulate and carry across generations, much like photocopying copies until details from the original gradually disappear.

An early symptom is that models begin losing their grasp of rare cases, or the Tails of distribution. Outputs may then become more homogeneous, repetitive, and less diverse. If Recursive training—continued training on data from previous model generations—continues, degradation may become severe enough that models produce distorted or meaningless results and may be difficult to restore.

Proposed countermeasures include mixing human-generated real data with synthetic data, Data curation to select and verify data quality, and Feedback mechanisms to check and refine outputs. The aim is to prevent AI-generated data from becoming the primary learning source without real-world reference points.

However, the risk of Model Collapse remains debated. Some researchers argue that closed-loop experiments may overstate the danger compared with real-world deployments, where systems typically receive a continuous supply of new human-generated data rather than relying entirely on synthetic data. The report therefore does not suggest that every use of synthetic data will cause models to collapse. Instead, it highlights the growing importance of data provenance, proportions, diversity, and quality control as AI-generated content becomes more widespread online.

Why it matters
As AI-generated content becomes more prevalent online, developers of Thai-language models must ensure that low-quality or duplicated data is not recycled into training until local knowledge and rare cases gradually disappear.
#Model Collapse#Synthetic Data#งานวิจัย AI#arXiv
Sources (rewritten & summarized from): arXiv cs.AI · bytez.com · arxiv.org · toloka.ai · ibm.com · medium.com

Comments

Loading comments…

No sign-up needed · comments are auto-filtered and moderated