Neural-Navigator-Labs/wmt14_en-de
SmolTrans EN-DE Dataset High-quality English–German parallel dataset filtered using: Language detection HTML removal Duplicate removal LM perplexity filtering Perplexity ratio filtering Embedding cosine similarity Statistics Total pairs: 4.5M Median length: ... Filtering removed: ~X% Format {"en": "...", "de": "..."} Usage from datasets import load_dataset ds = load_dataset("/smoltrans-en-de") Notes Short sentences… See the full description on the dataset page: https://huggingface.co/datasets/Neural-Navigator-Labs/wmt14_en-de.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face