CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01notrichardren /misconceptions_tf Dataset Card for "misconceptions_tf" More Information needed tabular1K<n<10K0 likes2k downloads3y agoHugging Face02vancevo /misconception_miningtabular10K<n<100K0 likes85 downloads5mo agoHugging Face03open-athena /synthetic-misconceptions-conversations Synthetic Misconceptions Conversations All data in this dataset is synthetic. No conversation here was had by a real person. The only human-authored source material is Wikipedia text: the corrections in List of common misconceptions about science, technology, and mathematics (260 entries), plus entries from List of conspiracy theories and Category:Health-related conspiracy theories (85 entries, filtered — see below). All of it is CC BY-SA licensed on Wikipedia. Everything… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/synthetic-misconceptions-conversations.tabulartext-generation1K<n<10K0 likes81 downloads8d agoHugging Face04Eedi /Eedi-Misconceptions-Graph Eedi Misconceptions Graph v1.0 A map of mathematical misconceptions and the curriculum constructs they appear in, released by Eedi under a CC BY 4.0 licence. Eedi defines a misconception as a flawed conceptual structure, or a gap in conceptual understanding, that manifests as a systematic and predictable error pattern across problems involving the same mathematical concept. A construct is a small, specific element of mathematics — for example, "Order fractions with the same… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Eedi-Misconceptions-Graph.tabular10K<n<100K0 likes79 downloads10d agoHugging Face05Morty0311 /misc-cfo-testing-cfo-analysis misc-cfo-testing CFO analysis This folder is the self-contained analysis output for: C:\Users\15255\Desktop\Research\CSE237D\morty_data\misc-cfo-testing The source data and the Weyl pipeline are read-only. All generated scripts, fingerprints, statistics, logs, PNGs, and SVGs remain in this analysis folder. Start with RESULTS.md. Dataset and estimator parameters Experiments: faraday (4 min), reboot (5 min), reboot-10m (10 min) Receiver: pluto11 Input sample rate:… See the full description on the dataset page: https://huggingface.co/datasets/Morty0311/misc-cfo-testing-cfo-analysis.imagen<1K0 likes75 downloads2mo agoHugging Face06miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes62 downloads1y agoHugging Face07miscovery /General_Facts_in_English_Arabic_Egyptian_Arabic 🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized) The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.tabularquestion-answering10K<n<100K12 likes47 downloads1y agoHugging Face08open-llm-leaderboard /sthenno-com__miscii-14b-1225-detailsgated Dataset Card for Evaluation run of sthenno-com/miscii-14b-1225 Dataset automatically created during the evaluation run of model sthenno-com/miscii-14b-1225 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sthenno-com__miscii-14b-1225-details.tabular10K<n<100K1 likes42 downloads2y agoHugging Face09instinct-org /miscellaneous_yt_chunked_tokenizedgated miscellaneous_yt_chunked_48k_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_tokenized.tabulartext-to-speech100K<n<1M0 likes39 downloads4mo agoHugging Face10andersonbcdefg /misc_sts_pairs_v2tabular10M<n<100M0 likes36 downloads3y agoHugging Face11vancevo /misconception_mining_asagtabular10K<n<100K0 likes34 downloads5mo agoHugging Face12Mediocreatmybest /Miscellany_of_Australian_Historical_Photographyimage1K<n<10K1 likes24 downloads4y agoHugging Face13Miscanthus /record-banana-greencube-smolvlatabular10K<n<100K0 likes20 downloads7mo agoHugging Face14guldasta /Math_misconceptiontabular10K<n<100K0 likes20 downloads4mo agoHugging Face15open-llm-leaderboard /sthenno-com__miscii-14b-1028-detailsgated Dataset Card for Evaluation run of sthenno-com/miscii-14b-1028 Dataset automatically created during the evaluation run of model sthenno-com/miscii-14b-1028 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sthenno-com__miscii-14b-1028-details.tabular10K<n<100K1 likes19 downloads2y agoHugging Face16jogarulfop /2026-04-30_miscellaneous_tasks_usb_chalk_magnet_glassThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower_dragontactile", "total_episodes": 9, "total_frames": 11079, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:9" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/jogarulfop/2026-04-30_miscellaneous_tasks_usb_chalk_magnet_glass.tabularrobotics10K<n<100K0 likes18 downloads5mo agoHugging Face17ClarusC64 /clinical-quad-central-lab-change-assay-drift-biomarker-noise-endpoint-misclassification-v0.1Clinical Quad Central Lab Change Assay Drift Biomarker Noise Endpoint Misclassification v0.1 Each row is a lab monthly snapshot. Core quad Central lab changeAssay driftBiomarker noiseEndpoint misclassification Target label_primary_fail_next_90d Files data/train.csvdata/tester.csvscorer.py Evaluation Run model on data/tester.csvReturn predictions row alignedScore with scorer.py License MIT tabulartext-classificationn<1K0 likes17 downloads7mo agoHugging Face18osanseviero /ag_misclassificationsThis dataset contains a slice of 200 samples from the AG News dataset (test split). The picked 200 samples are potential misclassifications of the original test data. Approach Fine-tune DistilBERT with 10k samples from the training data (out of 120k) Do a forward pass with the model, storing the loss Sort the samples based on the loss This is a repository for demonstration purposes tabularn<1K0 likes16 downloads3y agoHugging Face19Miscanthus /record-bananaThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 5, "total_frames": 4060, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Miscanthus/record-banana.tabularrobotics1K<n<10K0 likes16 downloads7mo agoHugging Face20electricsheepafrica /africa-cloud-misconfig-dataset Cloud Misconfiguration (Africa) | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-cloud-misconfig-dataset.tabulartabular-classification10K<n<100K0 likes16 downloads2mo agoHugging Face21miscovery /arabic_egypt_english_world_facts 🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized) The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.tabularquestion-answering10K<n<100K13 likes15 downloads1y agoHugging Face225hadytru /so101_misc_1tabular10K<n<100K0 likes15 downloads11mo agoHugging Face23Miscanthus /record-banana-greencubeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 26743, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Miscanthus/record-banana-greencube.tabularrobotics10K<n<100K0 likes14 downloads7mo agoHugging Face24misclassified /meps_speeches_with_translation.csvtabular10K<n<100K0 likes13 downloads3y agoHugging Face25open-llm-leaderboard /win10__miscii-14b-1M-0128-detailsgated Dataset Card for Evaluation run of win10/miscii-14b-1M-0128 Dataset automatically created during the evaluation run of model win10/miscii-14b-1M-0128 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/win10__miscii-14b-1M-0128-details.tabular10K<n<100K0 likes13 downloads2y agoHugging Face26Mischa01 /Teleoperationentabular1K<n<10K0 likes13 downloads3mo agoHugging Face27open-llm-leaderboard /bamec66557__MISCHIEVOUS-12B-Mix_0.3v-detailsgated Dataset Card for Evaluation run of bamec66557/MISCHIEVOUS-12B-Mix_0.3v Dataset automatically created during the evaluation run of model bamec66557/MISCHIEVOUS-12B-Mix_0.3v The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bamec66557__MISCHIEVOUS-12B-Mix_0.3v-details.tabular10K<n<100K0 likes9 downloads2y agoHugging Face28open-llm-leaderboard /bamec66557__MISCHIEVOUS-12B-Mix_0.4v-detailsgated Dataset Card for Evaluation run of bamec66557/MISCHIEVOUS-12B-Mix_0.4v Dataset automatically created during the evaluation run of model bamec66557/MISCHIEVOUS-12B-Mix_0.4v The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bamec66557__MISCHIEVOUS-12B-Mix_0.4v-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face29open-llm-leaderboard /bamec66557__MISCHIEVOUS-12B-Mix_0.5v-detailsgated Dataset Card for Evaluation run of bamec66557/MISCHIEVOUS-12B-Mix_0.5v Dataset automatically created during the evaluation run of model bamec66557/MISCHIEVOUS-12B-Mix_0.5v The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bamec66557__MISCHIEVOUS-12B-Mix_0.5v-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face30open-llm-leaderboard /bamec66557__MISCHIEVOUS-12B-Mix_Neo-detailsgated Dataset Card for Evaluation run of bamec66557/MISCHIEVOUS-12B-Mix_Neo Dataset automatically created during the evaluation run of model bamec66557/MISCHIEVOUS-12B-Mix_Neo The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bamec66557__MISCHIEVOUS-12B-Mix_Neo-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.