CoolFace
Datasetpublicgated

aurora-m/aurora-m-dataset-part-1

This is part of the continued pretraining dataset used to train the Aurora-M model described in [Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code, COLING 2025] (https://aclanthology.org/2025.coling-industry.56/).

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes86downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.