Training datasets for upcoming LLMs would include a significant amount of AI text/image data generated from current LLMs.
In such a scenario, it is important to understand how this affects batch decompositions and thereby, the performance of the resultant new LLM.
In this paper
We consider AI generated data as anomalies ``linked" to main data points and study decomposition and undersampling properties of the overall random dataset.
We use redundancy graphs and iteration techniques to obtain bounds for the minimum size of a strongly dissimilar (SD) decomposition and demonstrate a phase transition phenomena, wherein the minimum size is essentially determined by the \emph{main} data points when the number of anomalies is small and is ``taken" over by the anomalies above a certain threshold.
We also establish a size criticality result for the strong similarity of a randomly undersampled dataset and illustrate our results with examples involving categorical datasets, whose overall space size is much larger than the size of the dataset.