Man
Datasets
All datasets matching “Man”many-peptides-md
[!IMPORTANT]
Critical Update
The original 8AA TICA models within subsampled_trajectories/*/8AA/*.npz employed a CA-only atom selection. These models are not valid for comparison to results in our paper.
Updated files (uploaded 15/12/2025) now contain corrected models. If you previously downloaded this dataset, please re-download to ensure accurate results.
Note: Codebase references to tica_features_ca must now be replaced with tica_features. This was resolved in our codebase by PR #26.
Note:… See the full description on the dataset page: https://huggingface.co/datasets/transferable-samplers/many-peptides-md.trivia_qa
Dataset Card for "trivia_qa"
Dataset Summary
TriviaqQA is a reading comprehension dataset containing over 650K
question-answer-evidence triples. TriviaqQA includes 95K question-answer
pairs authored by trivia enthusiasts and independently gathered evidence
documents, six per question on average, that provide high quality distant
supervision for answering the questions.
Supported Tasks and Leaderboards
More Information Needed
Languages… See the full description on the dataset page: https://huggingface.co/datasets/mandarjoshi/trivia_qa.PhysicalAI-Robotics-Manipulation-Kitchen-Demos
PhysicalAI-Robotics-Manipulation-Kitchen-Demos
We provide a 600 hours of human-teleoperated demonstrations across 316 different tasks, totalling 55k trajectories.
The datasets are collected using Franka Panda robot with an Omron mobile base.
The datasets follow the LeRobot format. Here is an overview of important elements of each dataset:
Click to expand dataset structure
lerobot/
├── meta/ # Metadata files describing the dataset
│ ├── info.json… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Kitchen-Demos.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.IROS-2025-Challenge-Manip
IROS-2025-Challenge-Manip
Dataset Summary 📖
This dataset contains the IROS Challenge - Manipulation Track benchmark, organized into pretrain, train, and validation splits.
Pretrain split: ~20,000 single pick-and-place trajectories, packaged into tar files (each containing ~1,000 trajectories).
Train split: task-specific demonstrations, with ~100 trajectories provided per task.
Validation split: includes the test-time scenes and object assets in USD format.
Each… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/IROS-2025-Challenge-Manip.maniparena-dataset
ManipArena Dataset
Training dataset for ManipArena, a real-robot benchmark and competition for bimanual manipulation at the CVPR 2026 Embodied AI Workshop.
This dataset provides rich multi-modal demonstrations in LeRobot format, covering 20 real-robot tasks and 3 simulation tasks. Beyond standard end-effector trajectories, we provide joint positions, velocities, currents, camera views, and mobile-base states — giving participants the freedom to explore diverse input representations.… See the full description on the dataset page: https://huggingface.co/datasets/ManipArena/maniparena-dataset.

