aurora-m/aurora-m-dataset-part-1
This is part of the continued pretraining dataset used to train the Aurora-M model described in [Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code, COLING 2025] (https://aclanthology.org/2025.coling-industry.56/).
0159
This repository belongs to aurora-m on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
aurora-m-dataset-part-1
public
not set
yes
aurora-m
