CoolFace
Datasetpublic

ymoslem/UN-Arabic-English-Filtered

Dataset Details MultiUN + UNPC datasets, with rule-based and semantic filtering (train > 0.45 - test/dev > 0.9) as well as (>= 0.1) fasttext language detection. Dataset Structure DatasetDict({ train: Dataset({ features: ['text_en', 'text_ar'], num_rows: 19279407 }) test: Dataset({ features: ['text_en', 'text_ar'], num_rows: 8752 }) dev: Dataset({ features: ['text_en', 'text_ar'], num_rows: 8752… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/UN-Arabic-English-Filtered.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
2likes186downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
ymoslem/UN-Arabic-English-Filtered · CoolFace