datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
misconceptions_tf
Dataset Card for "misconceptions_tf"
More Information needed
misconception_miningsynthetic-misconceptions-conversations
Synthetic Misconceptions Conversations
All data in this dataset is synthetic. No conversation here was had by a real
person. The only human-authored source material is Wikipedia text: the corrections
in List of common misconceptions about science, technology, and
mathematics
(260 entries), plus entries from List of conspiracy
theories and
Category:Health-related conspiracy
theories
(85 entries, filtered — see below). All of it is CC BY-SA licensed on
Wikipedia. Everything… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/synthetic-misconceptions-conversations.misc_sts_pairs_v2misconception_mining_asagMiscellany_of_Australian_Historical_PhotographyMath_misconception2026-04-30_miscellaneous_tasks_usb_chalk_magnet_glassThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower_dragontactile",
"total_episodes": 9,
"total_frames": 11079,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:9"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/jogarulfop/2026-04-30_miscellaneous_tasks_usb_chalk_magnet_glass.record-bananaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 5,
"total_frames": 4060,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Miscanthus/record-banana.africa-cloud-misconfig-dataset
Cloud Misconfiguration (Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-cloud-misconfig-dataset.so101_misc_1record-banana-greencubeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 26743,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Miscanthus/record-banana-greencube.Teleoperationensthenno-com__miscii-14b-1028sthenno-com__miscii-14b-1225miscellaneous_yt_chunked_speech_restorised_nfa_aligned
miscellaneous_yt_chunked_speech_restorised_nfa_aligned
Public, manually gated NFA-aligned Uzbek speech dataset derived from instinct-org/miscellaneous_yt_chunked_speech_restorised.
Contents
Parquet shards: 130
Rows: 528,187
Approx hours: 863.88
Audio column: audio with embedded FLAC bytes
Transcript column: transcription
Alignment columns: nfa_token_alignments, nfa_word_alignments, nfa_segment_alignments, nfa_character_alignments
Access And Use… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised_nfa_aligned.agnostics-misc
Overview
eval-results: evaluation results in .parquet files. I don't think this works with datasets.load_datasets. Instead, you should directly open each Parquet file.
