CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sentence-transformers /parallel-sentences-ccmatrix Dataset Card for Parallel Sentences - CCMatrix This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. The texts originate from the CCMatrix dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices parallel-sentences-muse parallel-sentences-jw300 parallel-sentences-news-commentary… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-ccmatrix.textfeature-extraction1B<n<10B15 likes7k downloads2y agoHugging Face02Bingsu /st-parallel-sentences Dataset Card for "st-parallel-sentences" More Information needed text100M<n<1B1 likes6k downloads3y agoHugging Face03sentence-transformers /parallel-sentences-talks Dataset Card for Parallel Sentences - Talks This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the Talks dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-talks.textfeature-extraction10M<n<100M12 likes3.3k downloads2y agoHugging Face04sentence-transformers /parallel-sentences-opensubtitles Dataset Card for Parallel Sentences - OpenSubtitles This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the OpenSubtitles dataset. Warning! The quality of this dataset is not great; many of the english and non-english texts don't match well, or are fully empty. Related Datasets The following… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-opensubtitles.textfeature-extraction100M<n<1B4 likes2.5k downloads2y agoHugging Face05sentence-transformers /parallel-sentences-jw300 Dataset Card for Parallel Sentences - JW300 This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the JW300 dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-jw300.textfeature-extraction10M<n<100M10 likes2.1k downloads2y agoHugging Face06sentence-transformers /parallel-sentences-tatoeba Dataset Card for Parallel Sentences - Tatoeba This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the Tatoeba dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-tatoeba.textfeature-extraction1M<n<10M0 likes2k downloads2y agoHugging Face07sentence-transformers /parallel-sentences-opus-100 Dataset Card for Parallel Sentences - OPUS-100 This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. The sentences originate from the OPUS-100 website. In particular, this dataset is a reformatting of the OPUS-100 dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-opus-100.textfeature-extraction10M<n<100M4 likes1.9k downloads2y agoHugging Face08sentence-transformers /parallel-sentences-wikimatrix Dataset Card for Parallel Sentences - WikiMatrix This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the WikiMatrix dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-wikimatrix.textfeature-extraction10M<n<100M8 likes1.8k downloads2y agoHugging Face09AnanthZeke /tamil_sentences_master_raw Dataset Card for "tamil_sentences_master" More Information needed text10M<n<100M0 likes1.3k downloads3y agoHugging Face10sentence-transformers /parallel-sentences-europarl Dataset Card for Parallel Sentences - Europarl This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the Europarl dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-europarl.textfeature-extraction10M<n<100M1 likes1.2k downloads2y agoHugging Face11AnimaLab /bias-test-gpt-sentences Dataset Card for "BiasTestGPT: Generated Test Sentences" Dataset of sentences for bias testing in open-sourced Pretrained Language Models generated using ChatGPT and other generative Language Models. This dataset is used and actively populated by the BiasTestGPT HuggingFace Tool. BiasTestGPT HuggingFace Tool Dataset with Bias Specifications Project Landing Page Dataset Structure The dataset is structured as a set of CSV files with names corresponding to the social… See the full description on the dataset page: https://huggingface.co/datasets/AnimaLab/bias-test-gpt-sentences.text1K<n<10K1 likes920 downloads3y agoHugging Face12unb-labia /CCCPT-splited_preprocessed_max1024sz_sentencestext100M<n<1B0 likes892 downloads7mo agoHugging Face13agentlans /high-quality-english-sentences High-Quality English Sentences Dataset Description This dataset contains a collection of high-quality English sentences sourced from C4 and FineWeb (not FineWeb-Edu). The sentences have been carefully filtered and processed to ensure quality and uniqueness. "High-quality" means they're legible English and not spam, although they may still have spelling and grammar errors. Source Data Before filtering: C4: 1 million sentences FineWeb: 1 million sentences… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-english-sentences.texttext-classification1M<n<10M38 likes830 downloads2y agoHugging Face14sentence-transformers /parallel-sentences-global-voices Dataset Card for Parallel Sentences - Global Voices This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the Global Voices dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-global-voices.textfeature-extraction1M<n<10M1 likes758 downloads2y agoHugging Face15ghanaopenai /ghana-sentences Ghana Sentences A growing sentence-level text corpus for Ghanaian languages, tagged with ISO 639-3 codes and split into per-language subsets. The goal is broad-coverage text across all Ghanaian languages; this first release draws on school curriculum materials and a licensing-exam benchmark. More sources will be added over time. Language list and ISO codes follow GhanaNLP/ghana-taught-local-languages. Loading from datasets import load_dataset everything =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-sentences.texttext-generation100K<n<1M0 likes704 downloads2mo agoHugging Face16no7z /hsk-sentences-audio HSK Sentences Audio 4,354 Chinese sentences graded against the official HSK 3.0 levels 1–6, with pinyin, English translations, per-word glosses, grammar tags, and normal/slow synthetic speech. The complete export contains 8,708 MP3 files. Dataset structure The Viewer reads native Parquet from data/train.parquet, avoiding a dependency on Hugging Face's JSON-to-Parquet conversion service. The same 4,354 records are also available as validated JSON Lines in… See the full description on the dataset page: https://huggingface.co/datasets/no7z/hsk-sentences-audio.audiotext-to-speech1K<n<10K0 likes648 downloads2mo agoHugging Face17yoheikobashi /SAE_activations_modal_sentencestabular100K<n<1M0 likes598 downloads1y agoHugging Face18sentence-transformers /parallel-sentences-news-commentary Dataset Card for Parallel Sentences - News Commentary This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the News-Commentary dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-news-commentary.textfeature-extraction1M<n<10M2 likes423 downloads2y agoHugging Face19tankalapavankalyan /exp01-eeg-to-text-sentences Exp01 — Sentence-level EEG-to-text training data (unified) This is a private working corpus for experiment 1 (fine-tuning EEG / time-series foundation models on EEG-to-English-text). It bundles several public EEG-while-reading datasets into a single, raw-lossless parquet schema where one row = one sentence read by one participant. ⚠️ License: Per-source licenses are preserved verbatim in each row's license column and source_url. Do not re-distribute publicly without re-checking the… See the full description on the dataset page: https://huggingface.co/datasets/tankalapavankalyan/exp01-eeg-to-text-sentences.tabulartext-generation10K<n<100K0 likes341 downloads5mo agoHugging Face20gtfintechlab /financial_phrasebank_sentences_allagree Dataset Card for financial_phrasebank Dataset Summary Polar sentiment dataset of sentences from financial news. The dataset consists of 4840 sentences from English language financial news categorised by sentiment. The dataset is divided by agreement rate of 5-8 annotators. Supported Tasks and Leaderboards Sentiment Classification Languages English Dataset Structure Data Instances { "sentence": "Pharmaceuticals group Orion Corp… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/financial_phrasebank_sentences_allagree.tabulartext-classification1K<n<10K0 likes319 downloads1y agoHugging Face21agentlans /high-quality-multilingual-sentences High Quality Multilingual Sentences This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset. It includes 1.58 million rows across 51 different languages, each in its own configuration. Example row (from the all config): { "text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.", "fasttext": "fa", "gcld3": "fa" } Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.texttext-generation1M<n<10M9 likes316 downloads2y agoHugging Face22meg /fineweb-young-sentencestext100K<n<1M0 likes309 downloads2y agoHugging Face23Anon3365 /bias-test-gpt-sentencestext1K<n<10K0 likes303 downloads3y agoHugging Face24sarahwei /Taiwanese-Minnan-Example-Sentences Taiwanese Minnan Example Sentences The dataset consists of a collection of example sentences designed to aid in recognizing Taiwanese Minnan (Taiwanese Hokkien) for automatic speech recognition (ASR) tasks. This dataset is sourced from the Ministry of Education in Taiwan and aims to provide valuable linguistic resources for researchers and developers working on speech recognition systems. Dataset Features Source: Ministry of Education, Taiwan (Sutian Resource Center) Text:… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Example-Sentences.audioautomatic-speech-recognition10K<n<100K12 likes301 downloads2y agoHugging Face25TornikeO /zinc-sentencestext100M<n<1B0 likes282 downloads2y agoHugging Face26Jaswanth-0821 /ml_parallel_sentences_250ktext10M<n<100M0 likes275 downloads6mo agoHugging Face27agentlans /multilingual-sentences Multilingual Sentences Dataset contains sentences from 50 languages, grouped by their two-letter ISO 639-1 codes. The "all" configuration includes sentences from all languages. Dataset Overview Multilingual Sentence Dataset is a comprehensive collection of high-quality, linguistically diverse sentences. Dataset is designed to support a wide range of natural language processing tasks, including but not limited to language modeling, machine translation, and cross-lingual… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/multilingual-sentences.texttext-generation10M<n<100M6 likes272 downloads2y agoHugging Face28dvgodoy /yoda_sentences Yoda Speak This small dataset was built using two resources: Harvard Sentences, a list of 720 short sentences grouped into 72 sets of 10 sentences each English to Yoda Translator, an online translator that converts normal English into Yoda's way of speaking. Fun with this dataset I hope you have! Yes, hrrrm. texttranslationn<1K8 likes267 downloads2y agoHugging Face29sentence-transformers /wikipedia-en-sentences Dataset Card for Wikipedia Sentences (English) This dataset contains 7.87 million English sentences and can be used in knowledge distillation of embedding models. Dataset Details Columns: "sentence" Column types: str Examples:{ 'sentence': "After the deal was approved and NONG's stock rose to $13, Farris purchased 10,000 shares at the $2.50 price, sold 2,500 shares at the new price to reimburse the company, and gave the remaining 7,500 shares to Landreville at no cost… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/wikipedia-en-sentences.textfeature-extraction1M<n<10M7 likes213 downloads2y agoHugging Face30alakxender /dhivehi-noisy-sentences Dhivehi Noisy Sentences Dataset This dataset contains parallel examples of clean text and text with introduced errors across three categories: spelling, grammar, and punctuation. Dataset Description This dataset is designed to train models that can correct errors in Dhivehi text. Each example consists of: clean_text: The correct, error-free Dhivehi text noisy_text: The same text with introduced errors error_type: The category of error (spelling, grammar, or punctuation)… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-noisy-sentences.texttranslation1M<n<10M0 likes211 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.