CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceFW /finepdfs_lang_classificationtabular1M<n<10M4 likes19k downloads11mo agoHugging Face02RoboCOIN /Cobot_Magic_classification_of_tablewaregated Cobot_Magic_classification_of_tableware 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: agilex_cobot_decoupled_magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_tableware.tabularrobotics100K<n<1M0 likes1.7k downloads9mo agoHugging Face03RoboCOIN /Cobot_Magic_classification_of_fruits_and_vegetablesgated Cobot_Magic_classification_of_fruits_and_vegetables 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: agilex_cobot_decoupled_magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_fruits_and_vegetables.tabularrobotics100K<n<1M0 likes1.4k downloads9mo agoHugging Face04RoboCOIN /Cobot_Magic_classification_of_fruits_and_vegetables_agated Cobot_Magic_classification_of_fruits_and_vegetables_a 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: agilex_cobot_decoupled_magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_fruits_and_vegetables_a.tabularrobotics100K<n<1M0 likes1.1k downloads9mo agoHugging Face05ClaudiaRichard /mbti_classification_dataset_fullPoststabular1K<n<10K1 likes934 downloads3y agoHugging Face06matlok /python-image-copilot-training-using-class-knowledge-graphs Python Copilot Image Training using Class Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 312277 Size: 304.3 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.tabulartext-to-imagen<1K0 likes739 downloads3y agoHugging Face07VibeCuisine /cucumber-place-classifier-eval071526-v1-trimThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "vibeboard_follower_tilt", "total_episodes": 74, "total_frames": 2908, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:74" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-classifier-eval071526-v1-trim.tabularrobotics1K<n<10K0 likes671 downloads2mo agoHugging Face08VibeCuisine /cucumber-place-classifier-filtered071126This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos", "tilt.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-classifier-filtered071126.tabularrobotics1K<n<10K0 likes663 downloads3mo agoHugging Face09ClassiCC-Corpus /ClassiCC-PT 📚 ClassiCC-PT: Classified Common Crawl Corpus for Portuguese 📖 Overview ClassiCC-PT (Classified Common Crawl – Portuguese) is a large-scale web corpus containing ~120B Portuguese tokens extracted from Common Crawl snapshots. It is specifically curated for training large language models in Portuguese, with a focus on data quality, language specificity, and targeted filtering. This corpus was created as part of a study on continued pretraining for adapting English-trained… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/ClassiCC-PT.tabular10M<n<100M15 likes556 downloads8mo agoHugging Face10matlok /python-audio-copilot-training-using-class-knowledge-graphs Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.tabulartext-to-audion<1K0 likes478 downloads3y agoHugging Face11NuBerea /source-classificationsgated NuBerea Source Gold Set Curated source-critical classifications for the Hebrew Bible, New Testament, and Septuagint — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, translation traditions in the Septuagint) expressed as structured, verse-level data, together with statistical validation summaries and characteristic-vocabulary ("hallmark") term lists. This dataset is part of the NuBerea curated corpus… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-classifications.tabulartext-generation10K<n<100K0 likes428 downloads2mo agoHugging Face12ytzi /the-stack-dedup-python-filtered-classes_importsThis is a copy of bigcode/the-stack-dedup with some filters applied. The filters filtered in this dataset are: remove_classes remove_unused_imports remove_delete_markers tabular10M<n<100M0 likes422 downloads2y agoHugging Face13NuBerea /composition-classificationsgated NuBerea Composition Classifications A curated reference set of scholarly-consensus composition history for the biblical corpus: the traditions behind the Old Testament, Deuterocanon, New Testament, and Old Testament Pseudepigrapha, and the source-critical relationships among them (e.g. Documentary Hypothesis strands, Markan priority, canonical collection, translation into the Septuagint). The dataset is a direct transcription of established scholarship — no machine learning or… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/composition-classifications.tabulartext-classificationn<1K0 likes415 downloads2mo agoHugging Face14figmtu /aac_c4_deberta_classifiedThis dataset contains sentences from the Colossal Clean Crawled Corpus corpus. Each sentence is scored according to how similar it was to a spoken (dialogue_prob) or written (forum_prob) communication. See our EMNLP 2025 paper for details. tabular1B<n<10B1 likes388 downloads5mo agoHugging Face15matlok /python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27 Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27.tabulartext-to-audion<1K0 likes376 downloads3y agoHugging Face16lapa-llm /classifier_source Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a random sample of both https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality and https://huggingface.co/datasets/lapa-llm/pretraining-high-quality to transfer classifiers from English language to Ukrainian.It was used to transfer the following models from this collection https://huggingface.co/collections/lapa-llm/lapa-v012-pretraining:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/classifier_source.tabulartext-generation1M<n<10M0 likes365 downloads11mo agoHugging Face17formalmathatepfl /sft_classictabular1M<n<10M0 likes352 downloads1mo agoHugging Face18jy13 /bi-so101-fruits-classificationThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "bi_so101_follower", "total_episodes": 2, "total_frames": 2910, "total_tasks": 1, "total_videos": 6, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jy13/bi-so101-fruits-classification.tabularrobotics10K<n<100K0 likes344 downloads1y agoHugging Face19matlok /python-image-copilot-training-using-class-knowledge-graphs-2024-01-27 Python Copilot Image Training using Class Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 312836 Size: 294.1 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs-2024-01-27.tabulartext-to-imagen<1K0 likes336 downloads3y agoHugging Face20C-MTEB /JDReview-classification Dataset Card for "JDReview-classification" More Information needed tabular1K<n<10K1 likes334 downloads3y agoHugging Face21astro-legacy-archive /class-released-products CLASS released measurement products The preview renders the released Q_POLARISATION field from class_dr1_40GHz_skymap_n128 in its source RING order. This dataset contains LAMBDA's three CLASS DR1 40 GHz maps, three 90 GHz EE spectra, and 40 GHz circular-polarization limits. The release's transfer functions, beam, bandpass, simulations, masks, synchrotron-beta, reobserved and combined auxiliary maps, and software are excluded. Configuration names are source filename stems.… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/class-released-products.tabularn<1K0 likes324 downloads3d agoHugging Face22sunovivid /sit-latents-ode-heun-1000-class-0_1000-samples-segment-100-199tabular100K<n<1M0 likes308 downloads10mo agoHugging Face23Atika88 /Indonesian-ASR-11-Class-Dataset Indonesian ASR 11-Class Dataset Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts. Dataset summary Audio files: 104,500 WAV files Real/human recordings: 104,368 Synthetic repair files: 132 Sentence classes: 11 Indonesian sentence categories Canonical balanced sentence slots: 209 (11 categories × 19 retained slots) Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs* Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.tabularautomatic-speech-recognition100K<n<1M0 likes304 downloads16d agoHugging Face24Finnish-NLP /finepdf_fi_edu_score_topic_classifiedtabular1M<n<10M0 likes248 downloads1y agoHugging Face25Finnish-NLP /Fineweb2_fi_edu_score_topic_classifiedtabular10M<n<100M0 likes227 downloads10mo agoHugging Face26dvgodoy /CUAD_v1_Contract_Understanding_clause_classification Dataset Card for Contract Understanding Atticus Dataset (CUAD) Clause Classification This dataset contains 13,155 labeled clauses extracted from 509 commercial legal contracts from the original CUAD dataset. One of the original 510 contracts was removed due to being a scanned copy. The text was cleaned using clean-text. You can easily and quickly load it: dataset = load_dataset("dvgodoy/CUAD_v1_Contract_Understanding_clause_classification") Dataset({ features: ['file_name'… See the full description on the dataset page: https://huggingface.co/datasets/dvgodoy/CUAD_v1_Contract_Understanding_clause_classification.tabulartext-classification10K<n<100K0 likes200 downloads2y agoHugging Face27atin5551 /reddit-story-niche-classification-dataset 🧠 Reddit Niche Classification Dataset This dataset contains 13,061 Reddit posts annotated with a custom niche label (e.g. advice, drama, humor, unknown, etc). It includes structured features engineered from post metadata, not raw text — making it ideal for lightweight classification models. 🧾 Schema Column Type Description title string Post title selftext string Post body text subreddit string Subreddit the post belongs to flair string Flair… See the full description on the dataset page: https://huggingface.co/datasets/atin5551/reddit-story-niche-classification-dataset.tabulartext-classification10K<n<100K1 likes199 downloads1y agoHugging Face28VedantPadwal /clean-visual-webarena-classifiedstabular10K<n<100K0 likes193 downloads2y agoHugging Face29HyaDoo /ko-voicephishing-binary-classificationtabular1K<n<10K0 likes191 downloads2y agoHugging Face30minthanthtoo-cs /Burmese-Classics-OCR-RAW Burmese (Myanmar) Books Dataset – Burmese Classics OCR (Daily Rolling Project) Overview A daily rolling dataset of Burmese books, built for AI, OCR, and NLP research.Each entry includes OCR text with metadata: title, author, page index, and Burmese character ratio. This project fills a critical gap in Burmese-language resources: Scarcity of public-domain Burmese text. High technical and financial barriers to corpus building. Enables incremental, open access for… See the full description on the dataset page: https://huggingface.co/datasets/minthanthtoo-cs/Burmese-Classics-OCR-RAW.tabular1M<n<10M1 likes189 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.