datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-source-english-catalan-corpus
Dataset Card for open-source-english-catalan-corpus
Dataset Summary
Translation memory built from more than 180 open source projects. These include LibreOffice, Mozilla, KDE, GNOME, GIMP, Inkscape and many others. It can be used as translation memory or as training corpus for neural translators.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Catalan (ca)
English (en)
Dataset Structure
Data Instances
[More… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/open-source-english-catalan-corpus.open_parallel_think_source
Open Parallel Think — Source (per-model subsets)
Math reasoning traces distilled from a shared question set by four models, organized
one subset (config) per source model. Each question carries multiple reasoning traces
("parallel think"); here those traces are partitioned by the model that produced them.
The underlying questions come from three collections: openmathinstruct, numinamath,
and deepscale (the source is the prefix of guid, e.g. deepscale_10003).
Subsets… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/open_parallel_think_source.
