CoolFace
Datasetpublic

Helsinki-NLP/php

A parallel corpus originally extracted from http://se.php.net/download-docs.php. The original documents are written in English and have been partly translated into 21 languages. The original manuals contain about 500,000 words. The amount of actually translated texts varies for different languages between 50,000 and 380,000 words. The corpus is rather noisy and may include parts from the English original in some of the translations. The corpus is tokenized and each language pair has been sentence aligned. 23 languages, 252 bitexts total number of files: 71,414 total number of tokens: 3.28M total number of sentence fragments: 1.38M

sourceHugging Faceunknownupdated 3y agoView on Hugging Face
4likes150downloads
11 commits on main
56e5d243y ago

Delete legacy JSON metadata (#1)

albertvillanova
76a25744y ago

add dataset_info in dataset metadata

lhoestq
6a6caf84y ago

remove dummmy data

mariosasko
8c286e74y ago

Fix titles in dataset cards (#4824)

albertvillanova
70659d44y ago

Add `language_bcp47` tag (#4753)

lhoestq
01956144y ago

Align more metadata with other repo types (models,spaces) (#4607)

julien-c
78496ad4y ago

Update datasets task tags to align tags with models (#4067)

lhoestq
994381d5y ago

Update files from the datasets library (from 1.18.0)

system
2a188c35y ago

Update files from the datasets library (from 1.7.0)

system
d7adfdb5y ago

Update files from the datasets library (from 1.3.0)

system
abbd3d35y ago

Update files from the datasets library (from 1.2.0)

system