datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-papers-by-subject
arXiv Papers by Subject
A reorganised version of the nick007x/arxiv-papers dataset, partitioned by subject code, year, and month for efficient selective access.
Dataset Description
This dataset contains metadata for over 2.5 million arXiv papers, organised into a hierarchical directory structure that allows users to download only the specific subjects and time periods they need, rather than the entire dataset.
Motivation
The original… See the full description on the dataset page: https://huggingface.co/datasets/permutans/arxiv-papers-by-subject.wdc-common-crawl-embedded-jsonldPerMo
PerMo
Standalone PerMo release normalized to the repository SMPL-H convention.
Summary
PerMo is released here as a standalone motion dataset with SMPL-H motions,
single-motion captions, and style editing / style transfer instruction pairs.
Motions: 6,610 clips, 924,726 frames, 8.56 hours at 30 FPS.
Train split: 6,543 clips, 915,162 frames, 8.47 hours.
Test split: 67 clips, 9,564 frames, 0.09 hours.
Editing train split: 6,176 Neutral-to-style pairs, 869,194 target… See the full description on the dataset page: https://huggingface.co/datasets/ZeyuLing/PerMo.github-code-permissive-sampleSampling from codeparrot/github-code under more permissive license ['mit', 'apache-2.0', 'bsd-3-clause', 'bsd-2-clause', 'cc0-1.0'].It is intended to be used for training code language classifier.
AIC_final_safe_red_permuted_500AgiBotWorld-Beta_G1_task_532_Filler_permanent_magnet_ingot
agibot_task_532
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 填料永磁锭
total_episodes: 4367
total_tasks: 1
size: 81G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├── observation.images.back_left_fisheye… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_532_Filler_permanent_magnet_ingot.finephrase_permissivehive-partitioned-data-demomoral_education_permissive
Moral Education Permissive
Permissive-source subset of locuslab/moral_education, filtered using row-level metadata.url with the local permissiveness rules in this workspace. The output preserves the original row fields and adds url, idx, dump, language, source_config, is_permissive, is_oss, and permissive_reason where applicable.
Source configs included: score_4_morals, score_5_morals. Source split: train.
Rows scanned: 2,806,450. Rows kept: 72,943. Keep rate: 2.5991%.
See… See the full description on the dataset page: https://huggingface.co/datasets/ontocord/moral_education_permissive.powertron-global-permafrost-corpus
Dataset Card: Powertron Global PermaFrost Corpus
Important Disambiguation: This corpus documents PermaFrost® NMR, a trademarked HVAC efficiency treatment product. It contains HVAC/refrigeration efficiency data (chillers, RTUs, DX systems, refrigeration). This corpus has NO connection to geological permafrost (frozen ground), climate science, or Arctic research. The name "PermaFrost" is a product trademark reflecting thermal transfer properties, not a geological term.… See the full description on the dataset page: https://huggingface.co/datasets/powertronglobal/powertron-global-permafrost-corpus.c4-bbc-news
Dataset Card for BBC News from C4
This dataset provides a filtered subset of BBC News articles from the realnewslike subset of the C4 dataset, containing approximately 77k articles from BBC News domains.
Dataset Details
Dataset Sources
Repository: https://huggingface.co/datasets/permutans/c4-bbc-news
Source Dataset: allenai/c4 (realnewslike subset)
Paper: https://arxiv.org/abs/1910.10683 (C4 paper)
Uses
Direct Use
Suitable for text… See the full description on the dataset page: https://huggingface.co/datasets/permutans/c4-bbc-news.Word_Permuations_EnglishPERMA
PERMA: Benchmarking Personalized Memory Agents
TL;DR
PERMA is a benchmark for evaluating personalized memory agents in long-horizon conversations where user preferences evolve over time.Instead of static retrieval, models must track event-driven preference evolution and maintain persona consistency under realistic interaction noise.
This dataset supports two complementary evaluation protocols:
Multiple-choice evaluation for granular capability probing (task completion… See the full description on the dataset page: https://huggingface.co/datasets/ustclsc/PERMA.ExpansionRx_OpenADMET_Caco-2_Permeability_Papp_AB
ExpansionRx-OpenADMET Caco-2 Permeability Papp A>B
Caco-2 Permeability Papp A>B dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict Caco-2 Permeability Papp A>B of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
3773
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1:… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_Caco-2_Permeability_Papp_AB.object-permanence
Object Permanence
The training corpus of WROP (World Reasoning with Object Permanence): 1.5M Blender-rendered
video-continuation samples across 150 hand-designed cognitive tasks, one tar per task.
Abstract
Object permanence is the hallmark of human cognitive priors. Recent studies show that video models… See the full description on the dataset page: https://huggingface.co/datasets/Hokin/object-permanence.atlas-16-verifier-permission-prompt-ablation
ATLAS report 16: does the orchestrator's "cannot solve" clause suppress candidate verification?
Complete raw products of the ATLAS rl-training report 16 experiment
(GitHub issue #36). Two system-prompt arms of the same model over the
same 78 fixed states, greedy decoding, one shared vLLM server.
What the experiment did
The ATLAS orchestrator's frozen system prompt contains the clause
You cannot solve the problem yourself; you decide when to explore
further and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-16-verifier-permission-prompt-ablation.freesound-commercially-permissive-subset-with-captionsAIC_final_safe_red_permuted_500_labeledpermuformer_dataExpansionRx_OpenADMET_Caco-2_Permeability_Efflux
ExpansionRx-OpenADMET Caco-2 Permeability Efflux
Caco-2 Permeability Efflux dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict Caco-2 Permeability Efflux of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
3777
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1:… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_Caco-2_Permeability_Efflux.emoji-liif
Dataset Card for emoji-liif
This dataset contains 3,165 high-resolution emoji images that have been upscaled using the LIIF (Learning Implicit Image Function) method. The images were generated from original Apple emoji assets and are provided for research and academic purposes under fair use.
Dataset Details
Dataset Description
The emoji-liif dataset consists of upscaled emoji images generated from Apple's emoji assets. Each image has been enlarged to 2000x2000… See the full description on the dataset page: https://huggingface.co/datasets/permutans/emoji-liif.object-permanence-benchmark
Object Permanence Benchmark
The 300-question exam of WROP (World Reasoning with Object Permanence) and 14 video models' answers to it:
our own model PWM-WROP, five open-source baselines and eight hosted commercial models. Scores and
the human-preference leaderboard live on the project page; this repository holds the… See the full description on the dataset page: https://huggingface.co/datasets/Hokin/object-permanence-benchmark.Ultra_Permissive_TestAgiBotWorld-Beta_G1_task_584_Packaging_permanent_magnet_ingots
agibot_task_584
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 包装永磁锭
total_episodes: 101
total_tasks: 1
size: 5.5G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├── observation.images.back_left_fisheye… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_584_Packaging_permanent_magnet_ingots.PerMedCQA
PerMedCQA: Persian Medical Consumer QA Benchmark
PerMedCQA: Benchmarking Large Language Models on Medical Consumer Question Answering in Persian
PerMedCQA is the first large-scale, real-world benchmark for Persian-language medical consumer question answering. It contains anonymized medical inquiries from Persian-speaking users paired with professional responses, enabling rigorous evaluation of large language models in low-resource, health-related domains.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/NaghmehAI/PerMedCQA.dummy-lang-subset-dataset-1m-chunksDummy dataset uploaded to test the process, benefits of and constraints on/drawbacks to uploading subsets by a partition of a dataset.
This dataset is made up of fake data to illustrate a long tail of rarer languages similar to Wikipedia/Wikidata's distribution.
The dataset metadata was written automatically by the datasets library.
The config name is the language, and we iterate over all languages to do this.
Since the data is synthetic, we have the list of languages as a variable, without… See the full description on the dataset page: https://huggingface.co/datasets/permutans/dummy-lang-subset-dataset-1m-chunks.us-building-permits
PermitBase — U.S. Residential Building Permits 1980–2024
The most comprehensive historical residential building permit dataset available at the place level.
Place-level | Annual | 1980–2024 | SF/MF differentiated | 51 jurisdictions | 683,986 records
Dataset Description
This dataset contains annual residential building permit data for permit-issuing places
(cities, towns, and unincorporated county areas) across the United States, covering
1980 through 2024. It is derived… See the full description on the dataset page: https://huggingface.co/datasets/thanna94/us-building-permits.permanitai-framework
⚠️ LIVING WORK DOCUMENT — DRAFT STATE ⚠️
This dataset is part of the AUGMANITAI Compendium, a living research work document, continuously updated. Each entry is a priority anchor for terminological provenance — not a final reference. Errors, omissions and improvements are expected and explicitly part of the evolving methodology.
LEBENDES ARBEITSDOKUMENT — ENTWURFSSTADIUM. Laufend aktualisiert. Prioritäts-Anker, nicht finale Referenz.
Author: Andreas Ehstand · ORCID: 0009-0006-3773-7796 ·… See the full description on the dataset page: https://huggingface.co/datasets/AndreasEhstand/permanitai-framework.permutation_invariant_rewardPermaBotLaboThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 180,
"total_frames": 111945,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:180"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/PermaBotLabo.
