datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v1.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v1.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private.moltbook-corpus
Moltbook Corpus: Agent Social Behavior Dataset
This dataset provides for the research paper:
"FORM WITHOUT FUNCTION: AGENT SOCIAL BEHAVIOR IN THE MOLTBOOK NETWORK"📄 https://arxiv.org/abs/2604.13052
📊 Dataset Statistics
Category
Total Count
Collection Period
Jan 27, 2026 – Mar 8, 2026
Total Posts
1,312,238
Total Comments
6,691,460
Total Profiles
120,811
Total Submolts Metadata
108,490
Annotated Posts
394,221
Annotated Comments
2,100,589… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/moltbook-corpus.lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v3.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v3.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private.pad-auto-solver-reviewed
PAD Reviewed Dataset
Canonical reviewed PAD board/orb artifacts for dw-indie/pad-auto-solver-reviewed. This repository
contains immutable reviewed package revisions and does not contain raw captures,
training runs, checkpoints, or model binaries.
Packages exported: 28
Active catalog datasets: 14
Catalog schema: 3
Layout
packages/<dataset_id>.tar: deterministic self-contained reviewed package
catalog.json: active revision heads and coverage summary… See the full description on the dataset page: https://huggingface.co/datasets/dw-indie/pad-auto-solver-reviewed.put_beaker_on_pad_25_07_22_lerobotv21pi0_conversion_no_pad_videoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 1417,
"total_frames": 166855,
"total_tasks": 33,
"total_videos": 2834,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:1417"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlfu7/pi0_conversion_no_pad_video.put_bottle_on_pad_25_07_23_lerobotv21nfqa-multilingual-dataset
NFQA Multilingual Dataset
A large-scale multilingual dataset for Non-Factoid Question Answering (NFQA) classification, covering 49 languages and 8 question categories.
Dataset Statistics
Split
Examples
Train
28,653
Validation
3,539
Test
3,671
Total (Balanced)
35,863
Full Dataset (High Quality)
63,647
Dataset Composition
Languages (49 total)
Arabic (ar), Azerbaijani (az), Bulgarian (bg), Bengali (bn), Catalan (ca)… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/nfqa-multilingual-dataset.ultrafeedback_llama3_pad_rm
Dataset Card for "ultrafeedback_llama3_pad_rm"
More Information needed
piper_3cube_padfixThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "piper",
"total_episodes": 150,
"total_frames": 62990,
"total_tasks": 3,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:150"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/StyTJU/piper_3cube_padfix.agilex_put_mouse_on_padThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 19,
"total_frames": 6867,
"total_tasks": 1,
"total_videos": 57,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:19"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_put_mouse_on_pad.agent-launch-pad-trajectories
agent-launch-pad trajectories
Multi-bench trajectory dataset collected by agent-launch-pad.
Each row is one (agent × model × task) cell with the full sharegpt-format conversation
and a grade_pass signal from the bench's own verifier (pytest, reward.txt, etc).
Coverage
Total trajectories: 1380
grade_pass=True: 145 (10.5%)
Per benchmark
terminal-bench-2: 1204 cells, 138 grade_pass (11.5%)
scienceagentbench: 176 cells, 7 grade_pass (4.0%)
Per model… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/agent-launch-pad-trajectories.plug_socket_single_paddingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 22752,
"total_tasks": 1,
"total_videos": 150,
"total_audio": 150,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/plug_socket_single_padding.cube-pad-placementThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 90,
"total_frames": 61850,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vanniew/cube-pad-placement.PADBen
PADBen: Paraphrase and AI-Generated Text Detection Benchmark
📊 Dataset Overview
PADBen is a comprehensive benchmark for evaluating AI-generated text detection methods, specifically designed to test detection capabilities across various paraphrasing scenarios and attack vectors. For detailed implementation of how this dataset is generated/curated, please see https://github.com/JonathanZha47/PadBen-Paraphrase-Attack-Benchmark.
Total Dataset Size: 486,990 samples across 46… See the full description on the dataset page: https://huggingface.co/datasets/JonathanZha/PADBen.026-place-sponge-on-hot-pad-cleanThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/RevanthGundala/026-place-sponge-on-hot-pad-clean.swemera-10tasks-pyconfHere’s your Markdown text, organized for clear readability:
Task Description
Instances
Instance ID
Short Title
reframe-0
Performance threshold goes to -inf when it should be zero.
pyflakes-1
Walrus operator + annotation can cause F821
sqlglot-2
MySQL dialect fails to parse PRIMARY KEY USING BTREE syntax
matchms-3
matchms fails when reading spectra where abundance is in scientific notation #809
guarddog-4
Add Mach-O magic bytes to bundled binary detector… See the full description on the dataset page: https://huggingface.co/datasets/padamenko/swemera-10tasks-pyconf.super_chatton_annotated_pad26This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/super_chatton_annotated_pad26.chicago-crime-datatest_ds2
SWE-MERA
Continuously updated SWE-MERA dataset
SWE-MERA splits:
dev: for testing (10 samples)
lite: presented at the leaderboard here (750 samples)
full: continuously updated to collect more data (2738 samples)
Load dataset
from datasets import load_dataset
ds = load_dataset("MERA-evaluation/SWE-MERA", split='dev')
Evaluation
Description
The main tool to validate tasks is repotest (available at PyPI or GitHub)
data.jsonl -… See the full description on the dataset page: https://huggingface.co/datasets/padamenko/test_ds2.019-place-sponge-on-hot-pad-dagger15-trimmedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/RevanthGundala/019-place-sponge-on-hot-pad-dagger15-trimmed.Padebornomx_f_PAD_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omx_f",
"total_episodes": 2,
"total_frames": 822,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sangha1221/omx_f_PAD_2.vui-english-nonverbaltts-fluac-q9-22khz-v1
VUI English NonverbalTTS FLUAC Q9 22kHz v1
English NonverbalTTS subset prepared for VUI decoder distillation.
Processing
Source dataset: deepvk/NonverbalTTS
Text source: validated Result / fallback text
Audio codec: FLUAC Q9 22kHz
fluac_codes shape: [9, T]
No EOS/PAD frames included
Training collator should add EOS if needed
Final [pause] is appended to match VUI-style prompts
Unsupported emoji/nonverbal events are dropped
Supported mappings:
🌬 -> [breath]
🤣 -> [laugh]… See the full description on the dataset page: https://huggingface.co/datasets/padmanabansambath/vui-english-nonverbaltts-fluac-q9-22khz-v1.IS_cube_grasping_06_500_paddedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 500,
"total_frames": 14681,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/WillMandil001/IS_cube_grasping_06_500_padded.STARVLA_RoboTwin-Clean_place_mouse_pad_lerobotv30omx_f_PAD_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omx_f",
"total_episodes": 1,
"total_frames": 42,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sangha1221/omx_f_PAD_1.STARVLA_RoboTwin-Clean_move_stapler_pad_lerobotv30omx_f_PAD_2_curated_v1This dataset was created using Physical AI Tools and LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omx_f",
"total_episodes": 49,
"total_frames": 17865,
"total_tasks": 1,
"total_videos": 49,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:48"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sangha1221/omx_f_PAD_2_curated_v1.posterior_right_padding
