CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceFW /finepdfs_lang_classificationtabular1M<n<10M4 likes17k downloads11mo agoHugging Face02thesofakillers /jigsaw-toxic-comment-classification-challenge Dataset Description You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are: toxic severe_toxic obscene threat insult identity_hate You must create a model which predicts a probability of each type of toxicity for each comment. File descriptions train.csv - the training set, contains comments with their binary labels test.csv - the test set, you must predict the toxicity… See the full description on the dataset page: https://huggingface.co/datasets/thesofakillers/jigsaw-toxic-comment-classification-challenge.tabular100K<n<1M13 likes14k downloads2y agoHugging Face03RoboCOIN /Cobot_Magic_classification_of_tablewaregated Cobot_Magic_classification_of_tableware 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: agilex_cobot_decoupled_magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊 Dataset Statistics Metric… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_tableware.tabularrobotics100K<n<1M0 likes1.5k downloads9mo agoHugging Face04RoboCOIN /Cobot_Magic_classification_of_fruits_and_vegetablesgated Cobot_Magic_classification_of_fruits_and_vegetables 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: agilex_cobot_decoupled_magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_fruits_and_vegetables.tabularrobotics100K<n<1M0 likes1.3k downloads9mo agoHugging Face05imodels /tabular-benchmark-797-classificationtabular1K<n<10K0 likes1k downloads3y agoHugging Face06RoboCOIN /Cobot_Magic_classification_of_fruits_and_vegetables_agated Cobot_Magic_classification_of_fruits_and_vegetables_a 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: agilex_cobot_decoupled_magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_fruits_and_vegetables_a.tabularrobotics100K<n<1M0 likes993 downloads9mo agoHugging Face07ClaudiaRichard /mbti_classification_dataset_fullPoststabular1K<n<10K1 likes924 downloads3y agoHugging Face08CCB /cis5300-text-classification Complex Word Identification (CIS 5300) Dataset Description This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple. CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.tabulartext-classification1K<n<10K0 likes816 downloads5mo agoHugging Face09matlok /python-image-copilot-training-using-class-knowledge-graphs Python Copilot Image Training using Class Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 312277 Size: 304.3 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.tabulartext-to-imagen<1K0 likes709 downloads3y agoHugging Face10VibeCuisine /cucumber-place-classifier-eval071526-v1-trimThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "vibeboard_follower_tilt", "total_episodes": 74, "total_frames": 2908, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:74" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-classifier-eval071526-v1-trim.tabularrobotics1K<n<10K0 likes597 downloads2mo agoHugging Face11VibeCuisine /cucumber-place-classifier-filtered071126This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos", "tilt.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-classifier-filtered071126.tabularrobotics1K<n<10K0 likes588 downloads2mo agoHugging Face12ylab /methyl-classification DNA Methylation Tissue Classification Dataset Dataset Summary Homepage: https://github.com/ylaboratory/methylation-classification Pubmed: False Public: True This data resource is vast, curated reference atlas of DNA methylation (DNAm) profiles spanning 16,959 healthy primary human tissue and cell samples profiled on Illumina 450K arrays. Samples cover 86 unique tissues and cell types and are manually mapped to a common set of terms in the UBERON anatomical… See the full description on the dataset page: https://huggingface.co/datasets/ylab/methyl-classification.tabulartabular-classification10K<n<100K1 likes507 downloads1y agoHugging Face13matlok /python-audio-copilot-training-using-class-knowledge-graphs Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.tabulartext-to-audion<1K0 likes476 downloads3y agoHugging Face14ClassiCC-Corpus /ClassiCC-PT 📚 ClassiCC-PT: Classified Common Crawl Corpus for Portuguese 📖 Overview ClassiCC-PT (Classified Common Crawl – Portuguese) is a large-scale web corpus containing ~120B Portuguese tokens extracted from Common Crawl snapshots. It is specifically curated for training large language models in Portuguese, with a focus on data quality, language specificity, and targeted filtering. This corpus was created as part of a study on continued pretraining for adapting English-trained… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/ClassiCC-PT.tabular10M<n<100M15 likes476 downloads8mo agoHugging Face15NuBerea /classical-greekgated Classical Greek Corpus Ancient and classical Greek (grc) text segments drawn from the open scholarly corpora of the Perseus Digital Library and the OpenGreekAndLatin project — the classical/secular comparand within the NuBerea corpus estate, alongside its biblical, Second-Temple, and patristic Greek collections. Coverage runs from the archaic canon (Homer, Hesiod, the tragedians, the historians, Plato, Aristotle) through Hellenistic and imperial prose (Plutarch, Lucian, Galen)… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/classical-greek.tabulartext-generation10M<n<100M0 likes441 downloads10d agoHugging Face16NuBerea /source-classificationsgated NuBerea Source Gold Set Curated source-critical classifications for the Hebrew Bible, New Testament, and Septuagint — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, translation traditions in the Septuagint) expressed as structured, verse-level data, together with statistical validation summaries and characteristic-vocabulary ("hallmark") term lists. This dataset is part of the NuBerea curated corpus… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-classifications.tabulartext-generation10K<n<100K0 likes430 downloads2mo agoHugging Face17NuBerea /composition-classificationsgated NuBerea Composition Classifications A curated reference set of scholarly-consensus composition history for the biblical corpus: the traditions behind the Old Testament, Deuterocanon, New Testament, and Old Testament Pseudepigrapha, and the source-critical relationships among them (e.g. Documentary Hypothesis strands, Markan priority, canonical collection, translation into the Septuagint). The dataset is a direct transcription of established scholarship — no machine learning or… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/composition-classifications.tabulartext-classificationn<1K0 likes415 downloads2mo agoHugging Face18Arsive /toxicity_classification_jigsaw Dataset info Training Dataset: You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are: toxic severe_toxic obscene threat insult identity_hate The original dataset can be found here: jigsaw_toxic_classification Our training dataset is a sampled version from the original dataset, containing equal number of samples for both clean and toxic classes. Dataset creation:… See the full description on the dataset page: https://huggingface.co/datasets/Arsive/toxicity_classification_jigsaw.tabulartext-classification100K<n<1M5 likes393 downloads3y agoHugging Face19ytzi /the-stack-dedup-python-filtered-classes_importsThis is a copy of bigcode/the-stack-dedup with some filters applied. The filters filtered in this dataset are: remove_classes remove_unused_imports remove_delete_markers tabular10M<n<100M0 likes386 downloads2y agoHugging Face20figmtu /aac_c4_deberta_classifiedThis dataset contains sentences from the Colossal Clean Crawled Corpus corpus. Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication. See our EMNLP 2025 paper for details. tabular1B<n<10B1 likes386 downloads5mo agoHugging Face21ourafla /Mental-Health_Text-Classification_Dataset Mental Health Text Classification Dataset (4-Class) Dataset Description This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models. The repository includes: An… See the full description on the dataset page: https://huggingface.co/datasets/ourafla/Mental-Health_Text-Classification_Dataset.texttext-classification10K<n<100K9 likes381 downloads9mo agoHugging Face22matlok /python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27 Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27.tabulartext-to-audion<1K0 likes367 downloads3y agoHugging Face23lapa-llm /classifier_source Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a random sample of both https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality and https://huggingface.co/datasets/lapa-llm/pretraining-high-quality to transfer classifiers from English language to Ukrainian.It was used to transfer the following models from this collection https://huggingface.co/collections/lapa-llm/lapa-v012-pretraining:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/classifier_source.tabulartext-generation1M<n<10M0 likes356 downloads10mo agoHugging Face24classla /COPA-SR COPA-SR (The dataset uses cyrillic script. For the latin version, see this dataset.) The COPA-SR dataset (Choice of plausible alternatives in Serbian) is a translation of the English COPA dataset by following the XCOPA dataset translation methodology . The dataset consists of 1,000 premises (My body cast a shadow over the grass), each given a question (What is the cause? / What happened as a result?), and two choices (The sun was rising; The grass was cut), with a label encoding… See the full description on the dataset page: https://huggingface.co/datasets/classla/COPA-SR.tabulartext-classification1K<n<10K0 likes347 downloads3y agoHugging Face25jy13 /bi-so101-fruits-classificationThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "bi_so101_follower", "total_episodes": 2, "total_frames": 2910, "total_tasks": 1, "total_videos": 6, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jy13/bi-so101-fruits-classification.tabularrobotics10K<n<100K0 likes338 downloads1y agoHugging Face26C-MTEB /JDReview-classification Dataset Card for "JDReview-classification" More Information needed tabular1K<n<10K1 likes324 downloads3y agoHugging Face27Atika88 /Indonesian-ASR-11-Class-Dataset Indonesian ASR 11-Class Dataset Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts. Dataset summary Audio files: 104,500 WAV files Real/human recordings: 104,368 Synthetic repair files: 132 Sentence classes: 11 Indonesian sentence categories Canonical balanced sentence slots: 209 (11 categories × 19 retained slots) Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs* Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.tabularautomatic-speech-recognition100K<n<1M0 likes324 downloads13d agoHugging Face28astro-legacy-archive /class-released-products CLASS released measurement products The preview renders the released Q_POLARISATION field from class_dr1_40GHz_skymap_n128 in its source RING order. This dataset contains LAMBDA's three CLASS DR1 40 GHz maps, three 90 GHz EE spectra, and 40 GHz circular-polarization limits. The release's transfer functions, beam, bandpass, simulations, masks, synchrotron-beta, reobserved and combined auxiliary maps, and software are excluded. Configuration names are source filename stems.… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/class-released-products.tabularn<1K0 likes306 downloads4h agoHugging Face29formalmathatepfl /sft_classictabular1M<n<10M0 likes287 downloads1mo agoHugging Face30matlok /python-image-copilot-training-using-class-knowledge-graphs-2024-01-27 Python Copilot Image Training using Class Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 312836 Size: 294.1 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs-2024-01-27.tabulartext-to-imagen<1K0 likes277 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.