datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
finepdfs_lang_classificationjigsaw-toxic-comment-classification-challenge
Dataset Description
You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are:
toxic
severe_toxic
obscene
threat
insult
identity_hate
You must create a model which predicts a probability of each type of toxicity for each comment.
File descriptions
train.csv - the training set, contains comments with their binary labels
test.csv - the test set, you must predict the toxicity… See the full description on the dataset page: https://huggingface.co/datasets/thesofakillers/jigsaw-toxic-comment-classification-challenge.Cobot_Magic_classification_of_tableware
Cobot_Magic_classification_of_tableware
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_tableware.Cobot_Magic_classification_of_fruits_and_vegetables
Cobot_Magic_classification_of_fruits_and_vegetables
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_fruits_and_vegetables.tabular-benchmark-797-classificationCobot_Magic_classification_of_fruits_and_vegetables_a
Cobot_Magic_classification_of_fruits_and_vegetables_a
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_fruits_and_vegetables_a.mbti_classification_dataset_fullPostscis5300-text-classification
Complex Word Identification (CIS 5300)
Dataset Description
This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple.
CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.python-image-copilot-training-using-class-knowledge-graphs
Python Copilot Image Training using Class Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 312277
Size: 304.3 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.cucumber-place-classifier-eval071526-v1-trimThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "vibeboard_follower_tilt",
"total_episodes": 74,
"total_frames": 2908,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:74"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-classifier-eval071526-v1-trim.cucumber-place-classifier-filtered071126This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-classifier-filtered071126.methyl-classification
DNA Methylation Tissue Classification Dataset
Dataset Summary
Homepage: https://github.com/ylaboratory/methylation-classification
Pubmed: False
Public: True
This data resource is vast, curated reference atlas of DNA methylation (DNAm) profiles
spanning 16,959 healthy primary human tissue and cell samples profiled on Illumina 450K arrays.
Samples cover 86 unique tissues and cell types and are manually mapped to a common set of terms in the UBERON anatomical… See the full description on the dataset page: https://huggingface.co/datasets/ylab/methyl-classification.python-audio-copilot-training-using-class-knowledge-graphs
Python Copilot Audio Training using Class with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.ClassiCC-PT
📚 ClassiCC-PT: Classified Common Crawl Corpus for Portuguese
📖 Overview
ClassiCC-PT (Classified Common Crawl – Portuguese) is a large-scale web corpus containing ~120B Portuguese tokens extracted from Common Crawl snapshots. It is specifically curated for training large language models in Portuguese, with a focus on data quality, language specificity, and targeted filtering.
This corpus was created as part of a study on continued pretraining for adapting English-trained… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/ClassiCC-PT.classical-greek
Classical Greek Corpus
Ancient and classical Greek (grc) text segments drawn from the open scholarly
corpora of the Perseus Digital Library and the OpenGreekAndLatin project — the
classical/secular comparand within the NuBerea corpus estate, alongside its biblical,
Second-Temple, and patristic Greek collections. Coverage runs from the archaic canon
(Homer, Hesiod, the tragedians, the historians, Plato, Aristotle) through Hellenistic
and imperial prose (Plutarch, Lucian, Galen)… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/classical-greek.source-classifications
NuBerea Source Gold Set
Curated source-critical classifications for the Hebrew Bible, New Testament, and Septuagint — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, translation traditions in the Septuagint) expressed as structured, verse-level data, together with statistical validation summaries and characteristic-vocabulary ("hallmark") term lists.
This dataset is part of the NuBerea curated corpus… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-classifications.composition-classifications
NuBerea Composition Classifications
A curated reference set of scholarly-consensus composition history for the biblical corpus: the traditions behind the Old Testament, Deuterocanon, New Testament, and Old Testament Pseudepigrapha, and the source-critical relationships among them (e.g. Documentary Hypothesis strands, Markan priority, canonical collection, translation into the Septuagint). The dataset is a direct transcription of established scholarship — no machine learning or… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/composition-classifications.toxicity_classification_jigsaw
Dataset info
Training Dataset:
You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are:
toxic
severe_toxic
obscene
threat
insult
identity_hate
The original dataset can be found here: jigsaw_toxic_classification
Our training dataset is a sampled version from the original dataset, containing equal number of samples for both clean and toxic classes.
Dataset creation:… See the full description on the dataset page: https://huggingface.co/datasets/Arsive/toxicity_classification_jigsaw.the-stack-dedup-python-filtered-classes_importsThis is a copy of bigcode/the-stack-dedup with some filters applied.
The filters filtered in this dataset are:
remove_classes
remove_unused_imports
remove_delete_markers
aac_c4_deberta_classifiedThis dataset contains sentences from the Colossal Clean Crawled Corpus corpus.
Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication.
See our EMNLP 2025 paper for details.
Mental-Health_Text-Classification_Dataset
Mental Health Text Classification Dataset (4-Class)
Dataset Description
This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models.
The repository includes:
An… See the full description on the dataset page: https://huggingface.co/datasets/ourafla/Mental-Health_Text-Classification_Dataset.python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27
Python Copilot Audio Training using Class with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27.classifier_source
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a random sample of both https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality and https://huggingface.co/datasets/lapa-llm/pretraining-high-quality to transfer classifiers from English language to Ukrainian.It was used to transfer the following models from this collection https://huggingface.co/collections/lapa-llm/lapa-v012-pretraining:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/classifier_source.COPA-SR
COPA-SR
(The dataset uses cyrillic script. For the latin version, see this dataset.)
The COPA-SR dataset (Choice of plausible alternatives in Serbian) is a translation of the English COPA dataset by following the XCOPA dataset translation methodology .
The dataset consists of 1,000 premises (My body cast a shadow over the grass), each given a question (What is the cause? / What happened as a result?), and two choices (The sun was rising; The grass was cut), with a label encoding… See the full description on the dataset page: https://huggingface.co/datasets/classla/COPA-SR.bi-so101-fruits-classificationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_so101_follower",
"total_episodes": 2,
"total_frames": 2910,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jy13/bi-so101-fruits-classification.JDReview-classification
Dataset Card for "JDReview-classification"
More Information needed
Indonesian-ASR-11-Class-Dataset
Indonesian ASR 11-Class Dataset
Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts.
Dataset summary
Audio files: 104,500 WAV files
Real/human recordings: 104,368
Synthetic repair files: 132
Sentence classes: 11 Indonesian sentence categories
Canonical balanced sentence slots: 209 (11 categories × 19 retained slots)
Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs*
Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.class-released-products
CLASS released measurement products
The preview renders the released Q_POLARISATION field from
class_dr1_40GHz_skymap_n128 in its source RING order.
This dataset contains LAMBDA's three CLASS DR1 40 GHz maps, three 90 GHz EE
spectra, and 40 GHz circular-polarization limits. The release's
transfer functions,
beam,
bandpass,
simulations,
masks,
synchrotron-beta,
reobserved and
combined
auxiliary maps, and software are excluded. Configuration names are source
filename stems.… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/class-released-products.sft_classicpython-image-copilot-training-using-class-knowledge-graphs-2024-01-27
Python Copilot Image Training using Class Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 312836
Size: 294.1 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs-2024-01-27.
