datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RealCam-Vid
RealCam-Vid Dataset
News
25/04/08: We provide torch dataset demo code for example usage of our RealCam-Vid.
25/03/26: Release our dataset RealCam-Vid v1 for metric-scale camera-controlled video generation, containing ~100K video clips with dedicated short/long captions and metric-scale camera annotations.
25/02/18: Initial commit of the project, we plan to release the full dataset and data processing code in several… See the full description on the dataset page: https://huggingface.co/datasets/MuteApo/RealCam-Vid.ai-model-popularity
Datamata AI Model Popularity Index
Weekly popularity of the most-downloaded and trending Hugging Face models: trailing downloads, likes, the model's task and its trending rank. One row per model from the most recent weekly snapshot.
Latest snapshot: 2026-09-20
Models in this release: 50
Updated: weekly
Licence: CC BY 4.0 — free to use and adapt, including commercially, with attribution.
Source & methodology: https://www.datamatastudios.com/datasets
Quickstart… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/ai-model-popularity.GroMo25
GroMo25: Multiview Time-Series Plant Image Dataset for Age Estimation and Leaf Counting
Dataset Summary
GroMo25 is a multiview, time-series plant image dataset designed for plant age estimation (in days) and leaf counting tasks in precision agriculture. It contains high-quality images of four crop species — Wheat, Okra, Radish, and Mustard — captured over multiple days under controlled conditions. Each plant is photographed from 24 angles across 5 vertical levels per day… See the full description on the dataset page: https://huggingface.co/datasets/MrigLabIITRopar/GroMo25.CT_DeepLesion-MedSAM2
CT_DeepLesion-MedSAM2 Dataset
Authors
Jun Ma* 1,2,
Zongxin Yang* 3,
Sumin Kim2,4,5,
Bihui Chen2,4,5,
Mohammed Baharoon2,3,5,
Adibvafa Fallahpour2,4,5,
Reza Asakereh4,7,
Hongwei Lyu4,
Bo Wang† 1,2,4,5,6
* Equal contribution † Corresponding author
1AI Collaborative Centre, University Health Network, Toronto, Canada
2Vector… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/CT_DeepLesion-MedSAM2.multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.multiblimp
MultiBLiMP
MultiBLiMP is a massively Multilingual Benchmark for Linguistic Minimal Pairs. The dataset is composed of synthetic pairs generated using Universal Dependencies and UniMorph.
The paper can be found here.
We split the data set by language: each language consists of a single .tsv file. The rows contain many attributes for a particular pair, most important are the sen and wrong_sen fields, which we use for evaluating the language models.
Using MultiBLiMP
To… See the full description on the dataset page: https://huggingface.co/datasets/jumelet/multiblimp.global-piqa-nonparallel
Global PIQA Non-Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The non-parallel split covers 136 language varieties, covering five continents, 18 language families, and 24 writing systems.
In this non-parallel split, over 50% of examples reference local foods, customs, traditions, or other culturally-specific elements.
Details are in our preprint:… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-nonparallel.panlex-meanings
Dataset Card for panlex-meanings
This is a dataset of words in several thousand languages, extracted from https://panlex.org.
Dataset Details
Dataset Description
This dataset has been extracted from https://panlex.org (the 20240301 database dump) and rearranged on the per-language basis.
Each language subset consists of expressions (words and phrases).
Each expression is associated with some meanings (if there is more than one meaning, they are in separate… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/panlex-meanings.NOAH-mini
MOAH mini
The dataset prest here is a very samll sample of NOAH dataset.
In the original dataset each satellite image is ~650MB with 234,089 images present in 11 bands.
It is not feasible to upload the complete dataset.
A sample of the dataset across diffrent modalities can be seen in the figure below:
The diffrence between NOAH and NOAH mini is hilighted in the figure below.
Each subplot is a band of Landsat 8 in NOAH.
The region hilighted in red is the region available in NOAH… See the full description on the dataset page: https://huggingface.co/datasets/mutakabbirCarleton/NOAH-mini.spotify-tracks-dataset
Content
This is a dataset of Spotify tracks over a range of 125 different genres. Each track has some audio features associated with it. The data is in CSV format which is tabular and can be loaded quickly.
Usage
The dataset can be used for:
Building a Recommendation System based on some user input or preference
Classification purposes based on audio features and available genres
Any other application that you can think of. Feel free to discuss!
Column… See the full description on the dataset page: https://huggingface.co/datasets/maharshipandya/spotify-tracks-dataset.migration-bench-java-full
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.ML-ArXiv-PapersThis dataset contains the subset of ArXiv papers with the "cs.LG" tag to indicate the paper is about Machine Learning.
The core dataset is filtered from the full ArXiv dataset hosted on Kaggle: https://www.kaggle.com/datasets/Cornell-University/arxiv. The original dataset contains roughly 2 million papers. This dataset contains roughly 100,000 papers following the category filtering.
The dataset is maintained by with requests to the ArXiv API.
The current iteration of the dataset only contains… See the full description on the dataset page: https://huggingface.co/datasets/CShorten/ML-ArXiv-Papers.global-piqa-parallel
Global PIQA Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems.
In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.ArabicMMLU
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin
MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne
Introduction
We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.Amelia42-Mini
Dataset Overview
The Amelia42-Mini dataset provides air traffic position reports for 42 major U.S. airports, including the following airports:
KATL (Hartsfield-Jackson Atlanta International Airport)
KBDL (Bradley International Airport)
KBOS (Boston Logan International Airport)
KBWI (Baltimore/Washington International Thurgood Marshall Airport)
KCLE (Cleveland Hopkins International Airport)
KCLT (Charlotte Douglas International Airport)
KDCA (Washington National Airport)
KDEN… See the full description on the dataset page: https://huggingface.co/datasets/AmeliaCMU/Amelia42-Mini.MassSpecGym
MassSpecGym provides a dataset and benchmark for the discovery and identification of new molecules from tandem mass spectrometry (MS/MS) spectra. The provided challenges abstract the process of scientific discovery of new molecules from biological and environmental samples into well-defined machine learning problems.
Papers
MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery (2025): Paper Link
MassSpecGym: A benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/roman-bushuiev/MassSpecGym.jepa-qwen3-32b-pure-baselines-2026-05-25
JEPA-Align: Qwen3-32B Safety Defense Matrix
The complete 11-condition Qwen3-32B experiment for Predictive Representation
Alignment (PRA), the paired-view objective introduced in Predictive
Representation Alignment Improves Generalization in LLM Safety.
PRA aligns adversarially rewritten prompts with clean prompts expressing the
same intent. This release contains trained adapters, attack traces, benign
capability evaluations, machine-readable results, and paper-ready tables for… See the full description on the dataset page: https://huggingface.co/datasets/memo-ozdincer/jepa-qwen3-32b-pure-baselines-2026-05-25.M4Raw_brain
M4Raw Brain v1.6
M4Raw Brain is a multi-contrast, multi-repetition, four-channel k-space
dataset acquired on a 0.3-T whole-body MRI system. The release contains T1w,
T2w, FLAIR, and GRE brain acquisitions from healthy volunteers, including
explicit motion subsets and an expanded repeated-acquisition test cohort.
Companion dataset: M4Raw-Abdomen
is a separate low-field abdominal MRI k-space and segmentation dataset and
will be made public soon. Until then, the linked private… See the full description on the dataset page: https://huggingface.co/datasets/mylyu/M4Raw_brain.sim-datasets
SIM-Datasets: A Unified Symbolic Regression Benchmark
A standardized benchmark collection designed for the Scientific Intelligent Modelling (SIM) toolkit, providing comprehensive datasets for symbolic regression research and applications.
Overview
SIM-Datasets serves as a unified benchmark for symbolic regression tasks, offering standardized datasets with consistent formatting and evaluation protocols. This collection is specifically curated to support the Scientific… See the full description on the dataset page: https://huggingface.co/datasets/scientific-intelligent-modelling/sim-datasets.suno-ai-music-dataset
Suno AI Music Dataset (Multi-Genre Curated)
A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research.
This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.panlex-meanings
Dataset Card for panlex-meanings
This is a dataset of words in several thousand languages, extracted from https://panlex.org.
Dataset Details
Dataset Description
This dataset has been extracted from https://panlex.org (the 20240301 database dump) and rearranged on the per-language basis.
Each language subset consists of expressions (words and phrases).
Each expression is associated with some meanings (if there is more than one meaning, they are in separate… See the full description on the dataset page: https://huggingface.co/datasets/gtak1/panlex-meanings.simplified_grooveThis is a copy of the Magenta Groove dataset
The script ´simplify_midi_pretty.py` reads the midi data and simplifies it, by removing any midi values that aren't kicks or snares, and quantizing the notes.
MA_Query_Expansion_MLT26usgs-global-earthquake-catalog
USGS Global Earthquake Catalog
Provides historical data on global seismic events, sourced directly from the U.S. Geological Survey (USGS) Earthquake Hazards Program via its FDSN Event Web Service.
Each record represents a single seismic event (primarily earthquakes) and contains detailed information, including:
Event Time & Location: Precise timestamp, geographic coordinates (latitude, longitude), and depth of the event.
Magnitude: The magnitude of the event (mag) and the method… See the full description on the dataset page: https://huggingface.co/datasets/mnemoraorg/usgs-global-earthquake-catalog.Atmos_MODIS_AODstock-market-data-warehouseMALTA_LIBRAS
malta_libras_minds_subset:
Dataset tensors corresponding to all 20 LIBRAS signs from MINDS dataset.
malta_libras_complete:
Complete dataset tensors of all MALTA-LIBRAS collection.
montreal_firestock_factorsspambase
Spambase
The Spambase dataset from the UCI ML repository.
Is the given mail spam?
Configurations and tasks
Configuration
Task
Description
spambase
Binary classification
Is the mail spam?
Usage
from datasets import load_dataset
dataset = load_dataset("mstz/spambase")["train"]
