首页 > AI前沿 > Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies

Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies

arXiv机器学习 2026-08-14 14:15 8 阅读 查看原文

Training datasets for upcoming LLMs would include a significant amount of AI text/image data generated from current LLMs.

In such a scenario, it is important to understand how this affects batch decompositions and thereby, the performance of the resultant new LLM.

In this paper

We consider AI generated data as anomalies ``linked" to main data points and study decomposition and undersampling properties of the overall random dataset.

We use redundancy graphs and iteration techniques to obtain bounds for the minimum size of a strongly dissimilar (SD) decomposition and demonstrate a phase transition phenomena, wherein the minimum size is essentially determined by the \emph{main} data points when the number of anomalies is small and is ``taken" over by the anomalies above a certain threshold.

We also establish a size criticality result for the strong similarity of a randomly undersampled dataset and illustrate our results with examples involving categorical datasets, whose overall space size is much larger than the size of the dataset.