CoolFace
Datasetpublic

Helsinki-NLP/opus_paracrawl

Dataset Card for OpusParaCrawl Dataset Summary Parallel corpora from Web Crawls collected in the ParaCrawl project. Tha dataset contains: 42 languages, 43 bitexts total number of files: 59,996 total number of tokens: 56.11G total number of sentence fragments: 3.13G To load a language pair which isn't part of the config, all you need to do is specify the language code as pairs, e.g. dataset = load_dataset("opus_paracrawl", lang1="en", lang2="so") You can find… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/opus_paracrawl.

sourceHugging Facecc0-1.0updated 3y agoView on Hugging Face
6likes1.7kdownloads
15 commits on main
8bb0bcf3y ago

Convert dataset to Parquet (#2)

albertvillanova
1e6ba093y ago

Delete legacy JSON metadata (#1)

albertvillanova
922b0fb3y ago

rename configs to config_name

lhoestq
88bff844y ago

add dataset_info in dataset metadata

lhoestq
7e2f3fd4y ago

remove dummmy data

mariosasko
282d5f74y ago

Update version of opus_paracrawl dataset (#4816)

albertvillanova
67b24974y ago

Fix loading example in opus dataset cards (#4813)

albertvillanova
609af7a4y ago

Align more metadata with other repo types (models,spaces) (#4607)

julien-c
3172da04y ago

Remove config names as yaml keys (#4367)

lhoestq
0dbe0a54y ago

Update datasets task tags to align tags with models (#4067)

lhoestq
adbecf05y ago

Update files from the datasets library (from 1.18.0)

system
9e3e6de5y ago

Update files from the datasets library (from 1.7.0)

system
72e9c4d5y ago

Update files from the datasets library (from 1.6.0)

system
718825e5y ago

Update files from the datasets library (from 1.3.0)

system
85cc7a35y ago

Update files from the datasets library (from 1.2.0)

system