CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yoheikobashi /SAE_activations_modal_sentencestabular100K<n<1M0 likes598 downloads1y agoHugging Face02tankalapavankalyan /exp01-eeg-to-text-sentences Exp01 — Sentence-level EEG-to-text training data (unified) This is a private working corpus for experiment 1 (fine-tuning EEG / time-series foundation models on EEG-to-English-text). It bundles several public EEG-while-reading datasets into a single, raw-lossless parquet schema where one row = one sentence read by one participant. ⚠️ License: Per-source licenses are preserved verbatim in each row's license column and source_url. Do not re-distribute publicly without re-checking the… See the full description on the dataset page: https://huggingface.co/datasets/tankalapavankalyan/exp01-eeg-to-text-sentences.tabulartext-generation10K<n<100K0 likes341 downloads5mo agoHugging Face03gtfintechlab /financial_phrasebank_sentences_allagree Dataset Card for financial_phrasebank Dataset Summary Polar sentiment dataset of sentences from financial news. The dataset consists of 4840 sentences from English language financial news categorised by sentiment. The dataset is divided by agreement rate of 5-8 annotators. Supported Tasks and Leaderboards Sentiment Classification Languages English Dataset Structure Data Instances { "sentence": "Pharmaceuticals group Orion Corp… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/financial_phrasebank_sentences_allagree.tabulartext-classification1K<n<10K0 likes319 downloads1y agoHugging Face04danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes152 downloads10mo agoHugging Face05gtfintechlab /all_annotated_sentences_25000 Dataset Summary For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/all_annotated_sentences_25000 Additional Information This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 25,000 sentences taken from the meeting minutes of the 25 central banks referenced in our paper. Label… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/all_annotated_sentences_25000.tabulartext-classification10K<n<100K0 likes113 downloads1y agoHugging Face06cpllab /syntaxgym_sentencestabular1K<n<10K1 likes107 downloads4y agoHugging Face07Remonatef /pairs_cleaned_sentences_v1tabular10M<n<100M0 likes60 downloads6mo agoHugging Face08sumanthbhargava /ccnews-sentences-corruptiontabular100K<n<1M0 likes57 downloads4mo agoHugging Face09michsethowusu /twi-fante-sentences-parts-of-speech-pos-10m Akan POS Tagging Dataset - 10 Million Sentences Dataset Description The dataset includes sentences and part-of-speech tags and is aimed at supporting the development of POS tagging models for the Akan language. This work demonstrates that data limitations for low-resource languages can be overcome using artificial data generation techniques. Dataset Structure Data Fields sentence: The Akan sentence text pos_sequence: Part-of-speech tags for the… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/twi-fante-sentences-parts-of-speech-pos-10m.documenttext-classification1M<n<10M0 likes52 downloads1y agoHugging Face10tomron87 /hebrew-wikipedia-sentences-corpus Hebrew Wikipedia Sentences Corpus A corpus of 10,999,257 cleaned, deduplicated Hebrew sentences extracted from 366,610 Hebrew Wikipedia articles. Dataset Description This dataset contains Hebrew sentences extracted from Hebrew Wikipedia (crawled 2026-02). Each sentence has been cleaned, filtered for quality, and deduplicated. The dataset is intended for Hebrew NLP tasks including language modeling, text classification, NER, sentence similarity, and more. Source… See the full description on the dataset page: https://huggingface.co/datasets/tomron87/hebrew-wikipedia-sentences-corpus.tabulartext-classification10M<n<100M0 likes45 downloads7mo agoHugging Face11heegyu /kowiki-sentences20221001 한국어 위키를 kss(backend=mecab)을 이용해서 문장 단위로 분리한 데이터 549262 articles, 4724064 sentences 한국어 비중이 50% 이하거나 한국어 글자가 10자 이하인 경우를 제외 tabularother1M<n<10M11 likes43 downloads4y agoHugging Face12Mina-Rajaei-Moghadam /US-Presidents-Spoken-and-Written-SentencesUS Presidents' Spoken and Written Sentences We obtained transcriptions of spoken language from the Miller Center of Public Affairs, University of Virginia, which covers transcriptions from George Washington to the present time. For the writing samples, we used ten complete books written by presidents, three of which we obtained from Project Gutenberg. To ensure the accuracy of calculations, all the pages that were not part of the main content were removed. Furthermore, multiple… See the full description on the dataset page: https://huggingface.co/datasets/Mina-Rajaei-Moghadam/US-Presidents-Spoken-and-Written-Sentences.tabulartext-classification10K<n<100K2 likes41 downloads10mo agoHugging Face13erickfmm /agentlans__multilingual-sentences__paired_10_stsSentences from agentlans/multilingual-sentences in Spanish, and processed with Sentence Similarity Cosine Scores with model hiiamsid/sentence_similarity_spanish_es Each sentence in original dataset was randomly assigned 10 rows (sentences) within a batch of 1000, calculate the sentence similarity, and then deleted duplicate pairs The code for processing can be found here Useful for data distillation, training or benchmarking. Its recommended resampling the dataset to undersample to get a… See the full description on the dataset page: https://huggingface.co/datasets/erickfmm/agentlans__multilingual-sentences__paired_10_sts.tabularsentence-similarity1M<n<10M0 likes40 downloads1y agoHugging Face14reasoning-proj /all_continuations_sentence_doubt_dist_doubtful_sentences_embedded_doubtful_sentencestabular1M<n<10M0 likes37 downloads1y agoHugging Face15Minuri /nsina-sentences-raw Sinhala Raw Sentences - NSINA Raw Sinhala sentences extracted and sentence-split from the sinhala-nlp/NSINA news corpus using the SinLing sentence tokenizer. This is an intermediate dataset used in the construction of Minuri/diverse_sinhala_dataset. Dataset Structure Column Description text Raw Sinhala sentence char_count Character count of the sentence word_count Word count of the sentence Split Rows train 5,415,583… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/nsina-sentences-raw.tabulartext-generation1M<n<10M0 likes34 downloads6mo agoHugging Face16sumanthbhargava /ccnews-sentences-parallel-splitbtabular100K<n<1M0 likes34 downloads2mo agoHugging Face17sumanthbhargava /ccnews-sentences-parallel-splittabular100K<n<1M0 likes33 downloads2mo agoHugging Face18heegyu /namuwiki-sentences 38,015,081 rows tabularother10M<n<100M3 likes32 downloads4y agoHugging Face19sumanthbhargava /ccnews-sentences-corruption-splittabular100K<n<1M0 likes31 downloads4mo agoHugging Face20MohammedHemed /quran-sentences-ar Dataset Card for Quran Sentences (Hafs) Dataset Details Dataset Description This dataset provides the text of the Holy Quran (Narrated by Hafs) segmented into granular semantic units (sentences/phrases). Unlike traditional datasets that offer text at the "Ayah" (Verse) level, this dataset utilizes the standard Waqf (Pause) Marks found in the Othmani script to break verses into smaller, meaningful components. This granular approach is particularly useful for… See the full description on the dataset page: https://huggingface.co/datasets/MohammedHemed/quran-sentences-ar.tabular10K<n<100K0 likes29 downloads7mo agoHugging Face21chrisvoncsefalvay /smol-discharge-sentences-sfttabular10K<n<100K0 likes25 downloads10mo agoHugging Face22smcproject /ml-wiki-sentences ml-wiki-sentences Malayalam Wikipedia sentences extracted from article text, segmented into individual sentences. Dataset Description This dataset contains sentence-segmented text from Malayalam (ml) Wikipedia articles. Each row represents a single sentence extracted from Wikipedia articles, with metadata linking it back to the source article. Data Source Source: Malayalam Wikipedia (ml.wikipedia.org) Dump Date: March 2025 Original Dump: Wikimedia Enterprise… See the full description on the dataset page: https://huggingface.co/datasets/smcproject/ml-wiki-sentences.tabular1M<n<10M3 likes25 downloads7mo agoHugging Face23YnezT /BookMIA-in-sentencestabular100K<n<1M0 likes21 downloads2y agoHugging Face24jrosseruk /DeepSeek-R1-Distill-Llama-8B-MATH-labeled-sentences DeepSeek-R1-Distill-Llama-8B MATH Labeled Sentences Sentence-level function-tag labels for reasoning traces from DeepSeek-R1-Distill-Llama-8B on MATH problems. Trace model: deepseek-ai/DeepSeek-R1-Distill-Llama-8B (served via vLLM) Label model: gpt-4o-mini Source traces: jrosseruk/DeepSeek-R1-Distill-Llama-8B-MATH-traces-balanced Total sentences: 435,525 Backtrack sentences: 39,998 (9.2%) Traces: 4,413 (2,496 correct, 1,917 incorrect) Each sentence in a chain-of-thought trace is… See the full description on the dataset page: https://huggingface.co/datasets/jrosseruk/DeepSeek-R1-Distill-Llama-8B-MATH-labeled-sentences.tabulartext-generation100K<n<1M0 likes21 downloads8mo agoHugging Face25Jongbin-kr /SKIML-ICL_incontext_webq_v2-with-predicted-answer-sentences_with_nli_distributiontabular10K<n<100K0 likes20 downloads10mo agoHugging Face26Neuronovo /neuronovo-utc-hate-speech18-sentencestabular10K<n<100K0 likes18 downloads2y agoHugging Face27gh22994 /all_annotated_sentences_25000 Dataset Summary For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/all_annotated_sentences_25000 Additional Information This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 25,000 sentences taken from the meeting minutes of the 25 central banks referenced in our paper. Label… See the full description on the dataset page: https://huggingface.co/datasets/gh22994/all_annotated_sentences_25000.tabulartext-classification10K<n<100K0 likes18 downloads7mo agoHugging Face28trungpq /beer-com-sentencestabular10K<n<100K0 likes17 downloads1y agoHugging Face29YMETHO /all_annotated_sentences_25000 Dataset Summary For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/all_annotated_sentences_25000 Additional Information This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 25,000 sentences taken from the meeting minutes of the 25 central banks referenced in our paper. Label… See the full description on the dataset page: https://huggingface.co/datasets/YMETHO/all_annotated_sentences_25000.tabulartext-classification10K<n<100K1 likes17 downloads8mo agoHugging Face30blanchon /ears_dataset_sentences_tagstabular10K<n<100K1 likes16 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.