datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
petrodb
PetroDB
Public petroleum datasets served as a partitioned parquet tree, designed for
remote columnar queries with DuckDB httpfs (HTTP Range pushdown).
This Hugging Face dataset repo hosts the parquet bytes only. It mirrors the
parquet/ root and is queried directly via resolve URLs, which honour HTTP
Range requests so consumers fetch only the row groups a predicate needs.
INSTALL httpfs; LOAD httpfs;
SELECT *
FROM… See the full description on the dataset page: https://huggingface.co/datasets/sumpalabs/petrodb.longhorizon-orchestrator-benchmark
LongHorizon Orchestrator Benchmark
An offline benchmark for the judgements a manipulation orchestrator delegates to a
vision-language model. An orchestrator wraps a frozen low-level policy and replaces a compound
instruction ("put everything in the bin") with a stream of single-object subtasks; to do so it must
plan (decompose the instruction into subtasks), verify (judge from pixels whether the
current subtask is finished), track state (know which goals are already done), and —… See the full description on the dataset page: https://huggingface.co/datasets/petkopetkov/longhorizon-orchestrator-benchmark.typescript-codeIndustryCorpus2_petrochemical
IndustryCorpus2: Petrochemicals
This repository contains the IndustryCorpus2: Petrochemicals domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year =… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_petrochemical.merge_battery_30fpsyoutube-commons-small
📺 YouTube-Commons-Small 📺
This is a smaller subset of the YouTube-Commons dataset, which is a collection of audio transcripts from videos shared on YouTube under a CC-By license.
Dataset Description
This smaller version contains a subset of the original dataset, maintaining the same structure and features. It's designed for easier experimentation and testing purposes.
Features
The dataset includes the following information for each video:
Video ID and link… See the full description on the dataset page: https://huggingface.co/datasets/dm-petrov/youtube-commons-small.ICNDelay
ICNDelay: Multimodal Language–Time-Series Flight Delay Regression Dataset for Incheon International Airport
This dataset provides a structured, monthly collection of air traffic management scenarios for flight delay analysis and prediction.
Each scenario aggregates aircraft trajectory information, operational context, and structured flight attributes into a unified representation suitable for machine learning research.
The data are organized by month to support temporal… See the full description on the dataset page: https://huggingface.co/datasets/petchthwr/ICNDelay.danbooru2026-index
Danbooru 2026 Mirror Index
This dataset contains only the derived byte-range lookup artifact used by the
danbooru-mirror Space. It contains no image payloads.
tar-index.parquet has one row per image with id, shard, offset_bytes,
and length_bytes. The Space loads this Parquet file into PostgreSQL with
pg_parquet and uses the primary-key range index for lookups.
The repository may also contain metadata-smoke.parquet and
tar-index-smoke.parquet. These are three-row fixtures used… See the full description on the dataset page: https://huggingface.co/datasets/petersunde/danbooru2026-index.record-pet001This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 60,
"total_frames": 17370,
"total_tasks": 1,
"total_videos": 120,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mocacoffee/record-pet001.lm-eval-results-PetroGPT-WestSeverus-7B-DPO-v2-private
Dataset Card for Evaluation run of PetroGPT/WestSeverus-7B-DPO-v2
Dataset automatically created during the evaluation run of model PetroGPT/WestSeverus-7B-DPO-v2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-PetroGPT-WestSeverus-7B-DPO-v2-private.pettahdambulla_vegforecastpan2020_dict_author_fandom_doc
PAN2020 Fanfiction Author-Fandom-Disjoint Train/Validation Split
PAN 2020 / PAN 2021 fanfiction authorship verification data with Train/Validation split. The training data has been pre-split into Train and Validation under Author-Fandom-Disjoint constraints as is appropriate for PAN21 test data.
The training data is one row per document to allow easy recombination. The PAN21 validation and test splits consist of fixed document pairs for consistent scoring. The string fields are the… See the full description on the dataset page: https://huggingface.co/datasets/peterkirby/pan2020_dict_author_fandom_doc.standup_petbottleThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 38,
"total_frames": 37497,
"total_tasks": 1,
"total_videos": 38,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:38"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/standup_petbottle.catalytic-orfs-90pidworld-petroleum-assessment-units
World Petroleum Assessment Units
A modernized, AI/API-ready GeoParquet dataset of USGS World Petroleum Assessment 2000 Assessment Units.
Overview
The World Petroleum Assessment Units dataset is a modernization of the U.S. Geological Survey's World Petroleum Assessment 2000 (WPA 2000) Assessment Units — a globally distributed collection of petroleum assessment units defined by USGS geoscientists based on geological knowledge, exploration and production… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/world-petroleum-assessment-units.record-petbottle_revenge02This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 15,
"total_frames": 6346,
"total_tasks": 1,
"total_videos": 30,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mocacoffee/record-petbottle_revenge02.environment-multi-labels-even
Dataset Card for "environment-multi-labels-even"
More Information needed
Massive-STEPS-Petaling-Jaya
Massive-STEPS-Petaling Jaya
Dataset Summary
Massive-STEPSis a large-scale dataset of semantic trajectories intended for understanding POI check-ins. The dataset is derived from the Semantic Trails Dataset and Foursquare Open Source Places, and includes check-in data from 15 cities across 10 countries. The dataset is designed to facilitate research in various domains, including trajectory prediction, POI recommendation, and urban modeling. Massive-STEPS emphasizes… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/Massive-STEPS-Petaling-Jaya.industrial-sensor-anomaly-data
Industrial Equipment Sensor Anomaly Data
Overview
Synthetic multivariate sensor data from a simulated manufacturing plant with 5 equipment units (EQ-001 through EQ-005). Each unit generates 10,000 one-minute-interval readings across 11 sensor channels, 2 metadata fields, 3 derived features, and equipment operating mode labels.
The dataset is designed for anomaly detection benchmarking. It embeds 4 distinct anomaly types at approximately 4.5% prevalence:
Thermal runaway —… See the full description on the dataset page: https://huggingface.co/datasets/Petsteb/industrial-sensor-anomaly-data.combined-wildchat-qwen3-8b-petri-judged-responsesthermal_wrist_transparent_PET_pnpThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5e_follower",
"total_episodes": 62,
"total_frames": 16610,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:62"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Arururu12/thermal_wrist_transparent_PET_pnp.petrosafe-rag-corpus-fa
PetroSafe RAG Corpus (FA/EN)
Bilingual (Persian/English) knowledge corpus for process safety and HSE in oil, gas, and
petrochemical operations. Built for alirezaaminzadeh/petrosafe-rag-fa, the retrieval architecture
is inherited unchanged from hse-multimodal-rag-corpus
(hybrid BM25 + word/char TF-IDF, mandatory citations, abstention) — this repo supplies new
domain content, not a new retrieval method.
Data honesty (please read before citing any number from this… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/petrosafe-rag-corpus-fa.lm-eval-results-PetroGPT-WestSeverus-7B-DPO-private
Dataset Card for Evaluation run of PetroGPT/WestSeverus-7B-DPO
Dataset automatically created during the evaluation run of model PetroGPT/WestSeverus-7B-DPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-PetroGPT-WestSeverus-7B-DPO-private.typescript-jestjobcannon-psychometric-responses
JobCannon Psychometric Response Dataset
v3 — 54,431 item-level responses across nine instruments and 25 languages.
Anonymized, item-level responses to nine open-domain psychometric instruments,
collected from real test-takers on JobCannon. Each row is
one completed assessment: the raw per-item answers, the computed dimensional
scores, and the dominant result type.
This is a first-party dataset — our own users' responses, not a
re-publication of someone else's data.… See the full description on the dataset page: https://huggingface.co/datasets/PeterKol/jobcannon-psychometric-responses.s0101_grab_the_screwThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 110,
"total_frames": 54587,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:110"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/peterrolfes/s0101_grab_the_screw.cube_in_cup_subtask_handoffThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_roll.pos",
"left_gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Peter-zieg-uga/cube_in_cup_subtask_handoff.omx-smolvla-dataset_20260813_142158This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/peter1111aaaa/omx-smolvla-dataset_20260813_142158.flight-mh370-revisited-data
Flight MH370 Revisited — oversized source files
Companion data for the research repository
gmkf7vfyfb-web/flight-mh370-revisited.
That repository holds the complete project tree — model code, source data,
reports, figures, posterior outputs, the 148-file source-library audit and the
handoff dossier — from the consolidated snapshot of 14 August 2026. Seven files
were too large to keep in Git (one is 336 MB, above GitHub's hard 100 MB
per-file limit), so they live here instead.… See the full description on the dataset page: https://huggingface.co/datasets/peteabiome/flight-mh370-revisited-data.petal_pickup_combinedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/kbenz/petal_pickup_combined.
