datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SAE_activations_modal_sentencesexp01-eeg-to-text-sentences
Exp01 — Sentence-level EEG-to-text training data (unified)
This is a private working corpus for experiment 1 (fine-tuning EEG / time-series
foundation models on EEG-to-English-text). It bundles several public EEG-while-reading
datasets into a single, raw-lossless parquet schema where one row = one sentence read by
one participant.
⚠️ License: Per-source licenses are preserved verbatim in each row's license
column and source_url. Do not re-distribute publicly without re-checking the… See the full description on the dataset page: https://huggingface.co/datasets/tankalapavankalyan/exp01-eeg-to-text-sentences.financial_phrasebank_sentences_allagree
Dataset Card for financial_phrasebank
Dataset Summary
Polar sentiment dataset of sentences from financial news. The dataset consists of 4840 sentences from English language financial news categorised by sentiment. The dataset is divided by agreement rate of 5-8 annotators.
Supported Tasks and Leaderboards
Sentiment Classification
Languages
English
Dataset Structure
Data Instances
{ "sentence": "Pharmaceuticals group Orion Corp… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/financial_phrasebank_sentences_allagree.Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.all_annotated_sentences_25000
Dataset Summary
For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/all_annotated_sentences_25000
Additional Information
This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 25,000 sentences taken from the meeting minutes of the 25 central banks referenced in our paper.
Label… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/all_annotated_sentences_25000.syntaxgym_sentencespairs_cleaned_sentences_v1ccnews-sentences-corruptiontwi-fante-sentences-parts-of-speech-pos-10m
Akan POS Tagging Dataset - 10 Million Sentences
Dataset Description
The dataset includes sentences and part-of-speech tags and is aimed at supporting the development of POS tagging models for the Akan language.
This work demonstrates that data limitations for low-resource languages can be overcome using artificial data generation techniques.
Dataset Structure
Data Fields
sentence: The Akan sentence text
pos_sequence: Part-of-speech tags for the… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/twi-fante-sentences-parts-of-speech-pos-10m.hebrew-wikipedia-sentences-corpus
Hebrew Wikipedia Sentences Corpus
A corpus of 10,999,257 cleaned, deduplicated Hebrew sentences extracted from 366,610 Hebrew Wikipedia articles.
Dataset Description
This dataset contains Hebrew sentences extracted from Hebrew Wikipedia (crawled 2026-02). Each sentence has been cleaned, filtered for quality, and deduplicated. The dataset is intended for Hebrew NLP tasks including language modeling, text classification, NER, sentence similarity, and more.
Source… See the full description on the dataset page: https://huggingface.co/datasets/tomron87/hebrew-wikipedia-sentences-corpus.kowiki-sentences20221001 한국어 위키를 kss(backend=mecab)을 이용해서 문장 단위로 분리한 데이터
549262 articles, 4724064 sentences
한국어 비중이 50% 이하거나 한국어 글자가 10자 이하인 경우를 제외
US-Presidents-Spoken-and-Written-SentencesUS Presidents' Spoken and Written Sentences
We obtained transcriptions of spoken language from the Miller Center of Public Affairs, University of Virginia, which covers transcriptions from George Washington to the present time. For the writing samples, we used ten complete books written by presidents, three of which we obtained from Project Gutenberg. To ensure the accuracy of calculations, all the pages that were not part of the main content were removed. Furthermore, multiple… See the full description on the dataset page: https://huggingface.co/datasets/Mina-Rajaei-Moghadam/US-Presidents-Spoken-and-Written-Sentences.agentlans__multilingual-sentences__paired_10_stsSentences from agentlans/multilingual-sentences in Spanish, and processed with Sentence Similarity Cosine Scores with model hiiamsid/sentence_similarity_spanish_es
Each sentence in original dataset was randomly assigned 10 rows (sentences) within a batch of 1000, calculate the sentence similarity, and then deleted duplicate pairs
The code for processing can be found here
Useful for data distillation, training or benchmarking.
Its recommended resampling the dataset to undersample to get a… See the full description on the dataset page: https://huggingface.co/datasets/erickfmm/agentlans__multilingual-sentences__paired_10_sts.all_continuations_sentence_doubt_dist_doubtful_sentences_embedded_doubtful_sentencesnsina-sentences-raw
Sinhala Raw Sentences - NSINA
Raw Sinhala sentences extracted and sentence-split from the sinhala-nlp/NSINA news corpus using the SinLing sentence tokenizer. This is an intermediate dataset used in the construction of Minuri/diverse_sinhala_dataset.
Dataset Structure
Column
Description
text
Raw Sinhala sentence
char_count
Character count of the sentence
word_count
Word count of the sentence
Split
Rows
train
5,415,583… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/nsina-sentences-raw.ccnews-sentences-parallel-splitbccnews-sentences-parallel-splitnamuwiki-sentences
38,015,081 rows
ccnews-sentences-corruption-splitquran-sentences-ar
Dataset Card for Quran Sentences (Hafs)
Dataset Details
Dataset Description
This dataset provides the text of the Holy Quran (Narrated by Hafs) segmented into granular semantic units (sentences/phrases). Unlike traditional datasets that offer text at the "Ayah" (Verse) level, this dataset utilizes the standard Waqf (Pause) Marks found in the Othmani script to break verses into smaller, meaningful components.
This granular approach is particularly useful for… See the full description on the dataset page: https://huggingface.co/datasets/MohammedHemed/quran-sentences-ar.smol-discharge-sentences-sftml-wiki-sentences
ml-wiki-sentences
Malayalam Wikipedia sentences extracted from article text, segmented into individual sentences.
Dataset Description
This dataset contains sentence-segmented text from Malayalam (ml) Wikipedia articles. Each row represents a single sentence extracted from Wikipedia articles, with metadata linking it back to the source article.
Data Source
Source: Malayalam Wikipedia (ml.wikipedia.org)
Dump Date: March 2025
Original Dump: Wikimedia Enterprise… See the full description on the dataset page: https://huggingface.co/datasets/smcproject/ml-wiki-sentences.BookMIA-in-sentencesDeepSeek-R1-Distill-Llama-8B-MATH-labeled-sentences
DeepSeek-R1-Distill-Llama-8B MATH Labeled Sentences
Sentence-level function-tag labels for reasoning traces from DeepSeek-R1-Distill-Llama-8B on MATH problems.
Trace model: deepseek-ai/DeepSeek-R1-Distill-Llama-8B (served via vLLM)
Label model: gpt-4o-mini
Source traces: jrosseruk/DeepSeek-R1-Distill-Llama-8B-MATH-traces-balanced
Total sentences: 435,525
Backtrack sentences: 39,998 (9.2%)
Traces: 4,413 (2,496 correct, 1,917 incorrect)
Each sentence in a chain-of-thought trace is… See the full description on the dataset page: https://huggingface.co/datasets/jrosseruk/DeepSeek-R1-Distill-Llama-8B-MATH-labeled-sentences.SKIML-ICL_incontext_webq_v2-with-predicted-answer-sentences_with_nli_distributionneuronovo-utc-hate-speech18-sentencesall_annotated_sentences_25000
Dataset Summary
For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/all_annotated_sentences_25000
Additional Information
This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 25,000 sentences taken from the meeting minutes of the 25 central banks referenced in our paper.
Label… See the full description on the dataset page: https://huggingface.co/datasets/gh22994/all_annotated_sentences_25000.beer-com-sentencesall_annotated_sentences_25000
Dataset Summary
For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/all_annotated_sentences_25000
Additional Information
This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 25,000 sentences taken from the meeting minutes of the 25 central banks referenced in our paper.
Label… See the full description on the dataset page: https://huggingface.co/datasets/YMETHO/all_annotated_sentences_25000.ears_dataset_sentences_tags
