datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trivia_qa
Dataset Card for "trivia_qa"
Dataset Summary
TriviaqQA is a reading comprehension dataset containing over 650K
question-answer-evidence triples. TriviaqQA includes 95K question-answer
pairs authored by trivia enthusiasts and independently gathered evidence
documents, six per question on average, that provide high quality distant
supervision for answering the questions.
Supported Tasks and Leaderboards
More Information Needed
Languages… See the full description on the dataset page: https://huggingface.co/datasets/mandarjoshi/trivia_qa.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.project_gutenberg
Dataset Card for "Project Gutenberg"
Project Gutenberg is a library of over 70,000 free eBooks, hosted at https://www.gutenberg.org/.
All examples correspond to a single book, and contain a header and a footer of a few lines (delimited by a *** Start of *** and *** End of *** tags).
Usage
from datasets import load_dataset
ds = load_dataset("manu/project_gutenberg", split="fr", streaming=True)
print(next(iter(ds)))
License
Full license is available here:… See the full description on the dataset page: https://huggingface.co/datasets/manu/project_gutenberg.Mantis-Instruct
Mantis-Instruct
Paper | Website | Github | Models | Demo
Summaries
Mantis-Instruct is a fully text-image interleaved multimodal instruction tuning dataset,
containing 721K examples from 14 subsets and covering multi-image skills including co-reference, reasoning, comparing, temporal understanding.
It's been used to train Mantis Model families
Mantis-Instruct has a total of 721K instances, consisting of 14 subsets to cover all the multi-image skills.
Among the… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/Mantis-Instruct.arvo-cybergym-2000
ARVO CyberGym-format 2000-task dataset
This dataset is shaped to be loaded by Harbor's CyberGym adapter.
It combines jm-rt/arvo-cybergym-1000 with the second 1000-task
small-target ARVO batch built outside the original CyberGym set.
genshin-voice-v3.3-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.sacrebleu_manualNiji_1_Man-metadocument-review-data
Document Review Data
Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package.
Current Title Extraction Dataset Surface
Canonical prefix:
datasets/title_extraction/
Effective datasets:
datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/
datasets/title_extraction/evaluation/real_device_280_v1/
datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/
The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.hindi_audio_dataset_testappworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.40546875
Action score: 0.475
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.4046875
Action score: 0.4703125
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.39921875
Action score: 0.44375
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.38359375
Action score: 0.4703125
Valid samples: 320/320
code-20b
Dataset Card for "code_20b2"
More Information needed
kitchen-workspace-understanding-safe-manipulation
Kitchen Workspace Understanding & Safe Manipulation
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/kitchen-workspace-understanding-safe-manipulation.mmu_manga
mmu_manga HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_manga.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be installed via… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_manga.Indian-Laws
Dataset Card for Indian Laws
This is a comprehensive collection of primary legal documents pertinent to the Indian legal system.
It is designed to serve as a foundational resource for supervised fine-tuning (SFT) to make language models, particularly those focused on legal applications tailored for Indian law.
art_manip_datacode_20b
Dataset Card for "code_20b"
More Information needed
Mana-TTS
ManaTTS-Persian-Speech-Dataset
ManaTTS is the largest publicly available single-speaker Persian corpus, comprising over 114 hours of high-quality audio (sampled at 44.1 kHz). Released under the permissive CC-0 license, this dataset is freely usable for both educational and commercial purposes.
Collected from Nasl-e-Mana magazine, the dataset covers a diverse range of topics, making it ideal for training robust text-to-speech (TTS) models. The release includes a fully transparent… See the full description on the dataset page: https://huggingface.co/datasets/MahtaFetrat/Mana-TTS.agentic-ai-options-resultsToolRet-Queries🔧 Retrieving useful tools from a large-scale toolset is an important step for Large language model (LLMs) in tool learning. This project (ToolRet) contribute to (i) the first comprehensive tool retrieval benchmark to systematically evaluate existing information retrieval (IR) models on tool retrieval tasks; and (ii) a large-scale training dataset to optimize the expertise of IR models on this tool retrieval task.
See the official Github for more details.
A concrete example for our evaluation… See the full description on the dataset page: https://huggingface.co/datasets/mangopy/ToolRet-Queries.Infinity-Instruct
Infinity Instruct
Beijing Academy of Artificial Intelligence (BAAI)
[Paper][Code][🤗] (would be released soon)
The quality and scale of instruction data are crucial for model performance. Recently, open-source models have increasingly relied on fine-tuning datasets comprising millions of instances, necessitating both high quality and large scale. However, the open-source community has long been constrained by the high costs associated with building such extensive and… See the full description on the dataset page: https://huggingface.co/datasets/manifoldlabs/Infinity-Instruct.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🌌 Omni-Frontier Distillation SFT
The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise
Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
"The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.partnet-manualpa
PartNet-ManualPA
6,871 objects with original shape-frame part meshes and
70,923 rendered assembly PNGs. One object per row; all assets
are embedded in Parquet. No source-specific dataloader is needed.
from datasets import load_dataset
objects = load_dataset("AssemblyWorld/partnet-manualpa", split="all")
sample = objects[0]
Requires datasets >= 5.0.1 and Pillow. Pin revision to a release commit for
reproducibility. For local loading use the prepared directory instead of the Hub… See the full description on the dataset page: https://huggingface.co/datasets/AssemblyWorld/partnet-manualpa.rlt-maniskill-PegInsertionSide-v1-400-succ
RLT ManiSkill Joint
Dataset Summary
rlt_maniskill_joint is a LeRobot-style dataset for joint-control Robot Learning Token (RLT) training on the ManiSkill peg insertion task.
It is designed for the RLinf + OpenPI pi05_rlt_joint pipeline and is used in three stages:
OpenPI supervised fine-tuning (SFT) base policy training
RLT Stage 1 RL-token training
RLT Stage 2 online RL initialization and normalization
The dataset corresponds to the ManiSkill task:… See the full description on the dataset page: https://huggingface.co/datasets/RLinf/rlt-maniskill-PegInsertionSide-v1-400-succ.manariafriends
Bangumi Image Base of Manaria Friends
This is the image base of bangumi Manaria Friends, we detected 45 characters, 2199 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/manariafriends.SpatialLM-Testset
SpatialLM Testset
Project page | Paper | Code
We provide a test set of 107 preprocessed point clouds and their corresponding GT layouts, point clouds are reconstructed from RGB videos using MASt3R-SLAM. SpatialLM-Testset is quite challenging compared to prior clean RGBD scan datasets due to the noises and occlusions in the point clouds reconstructed from monocular RGB videos.
Folder Structure
Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Testset.ManipDreamer3D_data
Dataset of paper ManipDreamer3D
This repository contains the dataset for the paper ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory.
Data Structure
The dataset is organized as follows:
manipdreamer3d_data/
├── 000000/ # data processed from bridge-v1
├── 000000_v2/ # data processed from bridge-v2
│ ├── data.json # contains gripper state, camera params, etc.
│ ├── depth_0000.png… See the full description on the dataset page: https://huggingface.co/datasets/myendless/ManipDreamer3D_data.
