datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
misconceptions_tf
Dataset Card for "misconceptions_tf"
More Information needed
misconception_miningsynthetic-misconceptions-conversations
Synthetic Misconceptions Conversations
All data in this dataset is synthetic. No conversation here was had by a real
person. The only human-authored source material is Wikipedia text: the corrections
in List of common misconceptions about science, technology, and
mathematics
(260 entries), plus entries from List of conspiracy
theories and
Category:Health-related conspiracy
theories
(85 entries, filtered — see below). All of it is CC BY-SA licensed on
Wikipedia. Everything… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/synthetic-misconceptions-conversations.Eedi-Misconceptions-Graph
Eedi Misconceptions Graph v1.0
A map of mathematical misconceptions and the curriculum constructs they appear in, released by Eedi under a CC BY 4.0 licence.
Eedi defines a misconception as a flawed conceptual structure, or a gap in conceptual understanding, that manifests as a systematic and predictable error pattern across problems involving the same mathematical concept. A construct is a small, specific element of mathematics — for example, "Order fractions with the same… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Eedi-Misconceptions-Graph.misc-cfo-testing-cfo-analysis
misc-cfo-testing CFO analysis
This folder is the self-contained analysis output for:
C:\Users\15255\Desktop\Research\CSE237D\morty_data\misc-cfo-testing
The source data and the Weyl pipeline are read-only. All generated scripts,
fingerprints, statistics, logs, PNGs, and SVGs remain in this analysis folder.
Start with RESULTS.md.
Dataset and estimator parameters
Experiments: faraday (4 min), reboot (5 min), reboot-10m (10 min)
Receiver: pluto11
Input sample rate:… See the full description on the dataset page: https://huggingface.co/datasets/Morty0311/misc-cfo-testing-cfo-analysis.Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.General_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.sthenno-com__miscii-14b-1225-details
Dataset Card for Evaluation run of sthenno-com/miscii-14b-1225
Dataset automatically created during the evaluation run of model sthenno-com/miscii-14b-1225
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sthenno-com__miscii-14b-1225-details.miscellaneous_yt_chunked_tokenized
miscellaneous_yt_chunked_48k_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_tokenized.misc_sts_pairs_v2misconception_mining_asagMiscellany_of_Australian_Historical_Photographyrecord-banana-greencube-smolvlaMath_misconceptionsthenno-com__miscii-14b-1028-details
Dataset Card for Evaluation run of sthenno-com/miscii-14b-1028
Dataset automatically created during the evaluation run of model sthenno-com/miscii-14b-1028
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sthenno-com__miscii-14b-1028-details.2026-04-30_miscellaneous_tasks_usb_chalk_magnet_glassThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower_dragontactile",
"total_episodes": 9,
"total_frames": 11079,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:9"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/jogarulfop/2026-04-30_miscellaneous_tasks_usb_chalk_magnet_glass.clinical-quad-central-lab-change-assay-drift-biomarker-noise-endpoint-misclassification-v0.1Clinical Quad Central Lab Change Assay Drift Biomarker Noise Endpoint Misclassification v0.1
Each row is a lab monthly snapshot.
Core quad
Central lab changeAssay driftBiomarker noiseEndpoint misclassification
Target
label_primary_fail_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
ag_misclassificationsThis dataset contains a slice of 200 samples from the AG News dataset (test split).
The picked 200 samples are potential misclassifications of the original test data.
Approach
Fine-tune DistilBERT with 10k samples from the training data (out of 120k)
Do a forward pass with the model, storing the loss
Sort the samples based on the loss
This is a repository for demonstration purposes
record-bananaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 5,
"total_frames": 4060,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Miscanthus/record-banana.africa-cloud-misconfig-dataset
Cloud Misconfiguration (Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-cloud-misconfig-dataset.arabic_egypt_english_world_facts
🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized)
The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.so101_misc_1record-banana-greencubeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 26743,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Miscanthus/record-banana-greencube.meps_speeches_with_translation.csvwin10__miscii-14b-1M-0128-details
Dataset Card for Evaluation run of win10/miscii-14b-1M-0128
Dataset automatically created during the evaluation run of model win10/miscii-14b-1M-0128
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/win10__miscii-14b-1M-0128-details.Teleoperationenbamec66557__MISCHIEVOUS-12B-Mix_0.3v-details
Dataset Card for Evaluation run of bamec66557/MISCHIEVOUS-12B-Mix_0.3v
Dataset automatically created during the evaluation run of model bamec66557/MISCHIEVOUS-12B-Mix_0.3v
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bamec66557__MISCHIEVOUS-12B-Mix_0.3v-details.bamec66557__MISCHIEVOUS-12B-Mix_0.4v-details
Dataset Card for Evaluation run of bamec66557/MISCHIEVOUS-12B-Mix_0.4v
Dataset automatically created during the evaluation run of model bamec66557/MISCHIEVOUS-12B-Mix_0.4v
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bamec66557__MISCHIEVOUS-12B-Mix_0.4v-details.bamec66557__MISCHIEVOUS-12B-Mix_0.5v-details
Dataset Card for Evaluation run of bamec66557/MISCHIEVOUS-12B-Mix_0.5v
Dataset automatically created during the evaluation run of model bamec66557/MISCHIEVOUS-12B-Mix_0.5v
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bamec66557__MISCHIEVOUS-12B-Mix_0.5v-details.bamec66557__MISCHIEVOUS-12B-Mix_Neo-details
Dataset Card for Evaluation run of bamec66557/MISCHIEVOUS-12B-Mix_Neo
Dataset automatically created during the evaluation run of model bamec66557/MISCHIEVOUS-12B-Mix_Neo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bamec66557__MISCHIEVOUS-12B-Mix_Neo-details.
