CoolFace
21 results

para

singletongue /wikipedia-paragraphs wikipedia-paragraphs wikipedia-paragraphs is a dataset generated from Wikipedia, designed for natural language processing (NLP) research. Each entry contains cleaned paragraph text and Wikilink information extracted from a Wikipedia page, along with useful metadata such as categories, templates, and the associated Wikidata QID. Dataset structure Configurations The dataset is organized into multiple configurations, such as enwiki-20260607-v1.2.1.… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/wikipedia-paragraphs.tabular100M<n<1B2 likes10k downloads3mo agoHugging Facesentence-transformers /parallel-sentences-ccmatrix Dataset Card for Parallel Sentences - CCMatrix This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. The texts originate from the CCMatrix dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices parallel-sentences-muse parallel-sentences-jw300 parallel-sentences-news-commentary… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-ccmatrix.textfeature-extraction1B<n<10B15 likes7k downloads2y agoHugging FaceBingsu /st-parallel-sentences Dataset Card for "st-parallel-sentences" More Information needed text100M<n<1B1 likes6.4k downloads3y agoHugging Facephuongngo26791 /parallax0 likes6k downloads3d agoHugging Faceoss-codes /NCERT-Parallel-Dataset-Indictexttranslation100K<n<1M2 likes4.5k downloads2y agoHugging Facestrombergnlp /bornholmsk_parallelThis dataset is parallel text for Bornholmsk and Danish. For more details, see the paper [Bornholmsk Natural Language Processing: Resources and Tools](https://aclanthology.org/W19-6138/).texttranslation1K<n<10K2 likes3.7k downloads4y agoHugging Face