datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gaia2
Gaia2
Paper | Code | Project Page
Dataset Summary
Gaia2 is a benchmark dataset for evaluating AI agent capabilities in simulated environments. The dataset contains 800 scenarios that test agent performance in environments where time flows continuously and events occur dynamically.
The dataset evaluates seven core capabilities: Execution (multi-step planning and state changes), Search (information gathering and synthesis), Adaptability (dynamic response to environmental… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2.metaworld_mt50This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "metaworld",
"total_episodes": 2500,
"total_frames": 204806,
"total_tasks": 50,
"chunks_size": 1000,
"fps": 80,
"splits": {
"train": "0:2500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/metaworld_mt50.stackoverflow-with-meta-data
Dataset Card for "stackoverflow-with-meta-data"
More Information needed
the-stack-metadata
Dataset Card for The Stack Metadata
Changelog
Release
Description
v1.1
This is the first release of the metadata. It is for The Stack v1.1
v1.2
Metadata dataset matching The Stack v1.2
Dataset Summary
This is a set of additional information for repositories used for The Stack. It contains file paths, detected licenes as well as some other information for the repositories.
Supported Tasks and Leaderboards
The main… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-metadata.arxiv-metadata-snapshot
Dataset Card for "arxiv-metadata-oai-snapshot"
More Information needed
This is a mirror of the metadata portion of the arXiv dataset.
The sync will take place weekly so may fall behind the original datasets slightly if there are more regular updates to the source dataset.
Metadata
This dataset is a mirror of the original ArXiv data. This dataset contains an entry for each paper, containing:
id: ArXiv ID (can be used to access the paper, see below)
submitter:… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/arxiv-metadata-snapshot.Niji_1_Man-metadetails_meta-llama__Llama-3.1-8B-Instruct_private
Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct.
The dataset is composed of 78 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 20 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_meta-llama__Llama-3.1-8B-Instruct_private.Evo1_MetaWorld_DatasetColorectal-Liver_Metastases
Colorectal-Liver-Metastases (CRLM)
Preoperative contrast-enhanced CT scans + manual radiologist segmentations
from 197 patients who underwent hepatic resection for colorectal liver
metastases at Memorial Sloan Kettering Cancer Center (MSKCC). Released
under TCIA in 2023 alongside the Scientific Data descriptor by Simpson et
al. (2024).
Dataset Details
Field
Value
Modality
CT (preoperative, portal-venous phase, contrast-enhanced MDCT)
Body part
Liver… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/Colorectal-Liver_Metastases.visual_masked_distracting_metaworld
Visual Masked Distracting Meta-World (ground-truth masks)
Author: Georgios Tsakoumakis
Thesis: Interaction-Masked Latent Action Models for Object-Aware Manipulation under Visual Distractors (MSc, Imperial College London)
Expert Meta-World manipulation trajectories rendered with dynamic video-background
distractors, augmented with ground-truth segmentation masks and pose for the
manipulated object: the agent mask plus two per-frame fields, object_mask and
object_state.
All… See the full description on the dataset page: https://huggingface.co/datasets/tsakman23/visual_masked_distracting_metaworld.danbooru2025-metadata
🎨 Danbooru 2025 Metadata
Latest Post ID: 9,158,800
(as of Apr 16, 2025)
📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork.
Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes:
More consistent tag history tracking
Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.Taur_CoT_Analysis_Project___meta-llama__Meta-Llama-3.1-8B-Instructfineweb-edu-full-metadata[WIP]
FineWeb-Edu with Metadata
This repo contains 3 versions of the FineWeb-Edu v1 dataset:
fwedu1-metaonly/
fwedu1-text-content-zstd/
fineweb-edu-1.0.0-meta-and-text/
These are all joinable via the hash column, which is xxhash64 in pyspark, calculated on the text column. This hash is unique for all instances in the dataset. For convenience, this join is done for you in the third table
fwedu1-metaonly is just the metadata of the data exactly as it comes from the FineWeb-Edu v1… See the full description on the dataset page: https://huggingface.co/datasets/mmarone/fineweb-edu-full-metadata.Amazon_Sample_Metadata_2023
Dataset Card for Dataset Name
Original datasets can be found on: https://amazon-reviews-2023.github.io/
Dataset Details
This dataset was made as sample of several datasets from the link above.
Dataset Description
This dataset is a curated sample derived from seven filtered Amazon product category datasets(Amazon All Beauty, Amazon Fashion, Sports and Outdoors,
Health and Personal Care, Amazon Clothing Shoes and Jewlery,
Baby Products and Beauty and Personal… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Sample_Metadata_2023.5000-podcast-conversations-with-metadata-and-embedding-dataset
🗂️ ReadyAI - 5,000 Podcast Conversations with Metadata and Embedding Dataset
ReadyAI, operating subnet 33 on the Bittensor Network is an open-source initiative focused on low-cost, resource-minimal pipelines for structuring raw data for AI applications.
This dataset is part of the ReadyAI Conversational Genome Project, leveraging the Bittensor decentralized network.
AI runs on structured data — and this dataset bridges the gap between raw conversation transcripts and structured… See the full description on the dataset page: https://huggingface.co/datasets/ReadyAi/5000-podcast-conversations-with-metadata-and-embedding-dataset.RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode!
Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available.
RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.visual_distracting_metaworld
Visual Distracting Meta-World with Agent Masks
Visual Distracting Meta-World with Agent Masks contains successful expert
trajectories for all 50 Meta-World
MT50 v3 manipulation tasks. Every visual observation includes agent masks: a
ground-truth robot-arm mask obtained from the simulator, a SAM 2.1
mask predicted from the clean observation, and a separate SAM 2.1 mask predicted
from the distracted observation. Each step pairs these masks with the same
underlying simulator state… See the full description on the dataset page: https://huggingface.co/datasets/EpicPinkPenguin/visual_distracting_metaworld.stackoverflow-python-with-meta-data
Dataset Card for "stackoverflow-python-with-meta-data"
More Information needed
seamless-align-enA-jaA.speaker-embedding.metavoicemetaworld_ml45-v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "metaworld",
"total_episodes": 4391,
"total_frames": 358997,
"total_tasks": 44,
"total_videos": 0,
"total_chunks": 5,
"chunks_size": 1000,
"fps": 80,
"splits": {
"train": "0:4391"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/metaworld_ml45-v2.metaworld_mt50This dataset was created using LeRobot.
Dataset Description
This dataset contains 50 demonstrations per task from the Meta-world simulation benchmarks. Demonstrations are generated using expert policies.
Meta-world: https://arxiv.org/abs/1910.10897
We reposition the camera and flip the rendered images as follow:
Homepage: [More Information Needed]
Paper: [More Information Needed]
License: apache-2.0
Dataset Structure
meta/info.json:
{
"codebase_version":… See the full description on the dataset page: https://huggingface.co/datasets/ML-GOD/metaworld_mt50.Olympiads
Numina-Olympiads
Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers.
Dataset Information
Split: train
Original size: 32926
Filtered size: 32926
Source: olympiads
All examples contain valid boxed answers
Dataset Description
This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes:
A mathematical word problem
A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Olympiads.visual_masked_distracting_metaworld_sam
Visual Masked Distracting Meta-World (ground-truth + SAM masks)
Author: Georgios Tsakoumakis
Thesis: Interaction-Masked Latent Action Models for Object-Aware Manipulation under Visual Distractors (MSc, Imperial College London)
Expert Meta-World manipulation trajectories rendered with dynamic video-background
distractors, carrying both ground-truth and SAM-predicted segmentation masks for the
agent and the manipulated object, plus the manipulated object's pose.
This is the SAM… See the full description on the dataset page: https://huggingface.co/datasets/tsakman23/visual_masked_distracting_metaworld_sam.arxiv_metadata_by_year
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/arxiv_metadata_by_year.dataset_cards_with_metadatacrossref_metadata_2025
Dataset Overview
This dataset contains bibliographic metadata from the public Crossref snapshot released in 2025. It provides core fields for scholarly documents, including DOI, title, abstract, authorship, publication month and year, and URLs. The entire public dump (~196.94 GB) was filtered and extracted into a parquet format for efficient loading and querying.
Total size: 196.94 GB (parquet files)
Number of records: 34,308,730
Use this dataset for large-scale text mining… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_2025.model_cards_with_metadataNLU-Metaphor
SEA Metaphor
SEA Metaphor evaluates a model's ability to interpret paired figurative phrases with divergent meanings. It is sampled from Multilingual-Fig-QA for Indonesian, Javanese, and Sundanese.
Supported Tasks and Leaderboards
SEA Metaphor is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Javanese (jv)
Sundanese (su)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Metaphor.qwen3.5-metamathqaGigaMIDI
Dataset Card for GigaMIDI
The Extended GigaMIDI Dataset Summary
We present the extended GigaMIDI dataset [https://huggingface.co/datasets/Metacreation/GigaMIDI/viewer/v2.0.0], a large-scale symbolic music collection comprising over 2.1 million unique MIDI files with detailed annotations for music loop detection. Expanding on its predecessor, this release introduces a novel expressive loop detection method that captures performance nuances such as microtiming and dynamic… See the full description on the dataset page: https://huggingface.co/datasets/Metacreation/GigaMIDI.
