Official Repository of "LLM × DATA" Survey Paper
-
Updated
Jun 15, 2026
Official Repository of "LLM × DATA" Survey Paper
A Python library for combining multiple datasets using customizable algorithms. It offers a flexible framework to control the contribution of each dataset in the final output.
Unofficial reproduction of Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training (NeurIPS 2025, arXiv:2504.13161) — search-found mixtures beat uniform baselines +0.014–0.031 STEM at d28; novel finding: selection-mechanism winner's curse; Ascend NPU backend.
To associate your repository with the data-mixing topic, visit your repo's landing page and select "manage topics."