CoolFace
Datasetpublic

schneiderkamplab/opus-da-en-permissive

OPUS da-en permissive subset This dataset is a processed Danish-English sentence-pair subset harvested from OPUS corpora available at https://opus.nlpl.eu/. The subset was selected from OPUS da-en sources with permissive Creative Commons or public-domain style licensing, then converted into JSONL.GZ for DFM/HRM-Text training. It is intended to preserve the exact rows used by the local DFM data pipeline, so DFM8 can be rebuilt without relying on local-only files.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/opus-da-en-permissive.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes8downloads

schneiderkamplab/opus-da-en-permissive · main · files are served by the source, never re-hosted here