datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-model-popularity
Datamata AI Model Popularity Index
Weekly popularity of the most-downloaded and trending Hugging Face models: trailing downloads, likes, the model's task and its trending rank. One row per model from the most recent weekly snapshot.
Latest snapshot: 2026-09-20
Models in this release: 50
Updated: weekly
Licence: CC BY 4.0 — free to use and adapt, including commercially, with attribution.
Source & methodology: https://www.datamatastudios.com/datasets
Quickstart… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/ai-model-popularity.aimo-validation-amc
Dataset Card for AIMO Validation AMC
All 83 come from AMC12 2022, AMC12 2023, and have been extracted from the AOPS wiki page https://artofproblemsolving.com/wiki/index.php/AMC_12_Problems_and_Solutions
This dataset serves as an internal validation set during our participation in the AIMO progress prize competition. Using data after 2021 is to avoid potential overlap with the MATH training set.
Here are the different columns in the dataset:
problem: the modified problem statement… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/aimo-validation-amc.amc_filteredaime_filteredprompt-swap-mixed12-5xlr-e1-mxfp4-mergedprompt-swap-mixed12-5xlr-e2-mxfp4-mergedprompt-swap-medium12-e2-mxfp4-mergedPolyUniMath
PolyUniMath
Dataset Summary
PolyUniMath is a large multilingual mathematics dataset of question-solution-answer pairs extracted from mathematical PDFs.
The dataset is designed for training and studying natural-language mathematical reasoning, with a strong emphasis on university-level content.
Sample count: approximately 4 million Q&A pairs
Main focus: university-level mathematics
Format: problem, optional choices, solution, final answer, and auxiliary… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/PolyUniMath.val-sampleaugmented-sample-math-aggaimo-interp-challenge-sample-fullaugmented-sample-math-agg-filteredaugmented-sample-mathaugmented-sample-math-full-filteredprompt-swap-medium12-e1-mxfp4-mergedaimo-math-problemsMathematical QA data collections combining GSM8k, MATH and historical national mathematics competitions data extracted from AoPS, such as AMC and AIME.
The dataset splits into two difficulty levels, given problems' affinity to AIMO competition.
Hard: AMC12, AIME
Not hard: GSM8K, MATH, AHSME, USAMO, USOMO USAJMO, AJHSME, AMC8, AMC10
These data are further deduplicated, and filtered to keep those with text-based description (instead of replying on images) and integer answers, to match AIMO… See the full description on the dataset page: https://huggingface.co/datasets/billxbf/aimo-math-problems.aimo-complex_decontaminatedapex_gpt-oss20b-low_gemini3-flash-lowAI-MO__NuminaMath-7B-TIR-details
Dataset Card for Evaluation run of AI-MO/NuminaMath-7B-TIR
Dataset automatically created during the evaluation run of model AI-MO/NuminaMath-7B-TIR
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AI-MO__NuminaMath-7B-TIR-details.omy_f3m_PnP_Yellow_To_RedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m",
"total_episodes": 2,
"total_frames": 230,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AImond/omy_f3m_PnP_Yellow_To_Red.AI-MO__NuminaMath-7B-CoT-details
Dataset Card for Evaluation run of AI-MO/NuminaMath-7B-CoT
Dataset automatically created during the evaluation run of model AI-MO/NuminaMath-7B-CoT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AI-MO__NuminaMath-7B-CoT-details.aimo3-ref-valaimo2-irpo-Qwen2.5-Math-7B-Instructaimo-complexaimo-interp-challenge-sample-v2ai-models-database
Convly AI Models Database
A continuously updated, hand-verified dataset of 30+ AI language models — specs, licenses, API pricing (USD per 1M tokens), and local-hardware (VRAM) requirements.
Maintained by Convly.ai · Live interactive version: https://convly.ai/models/
Fields
name, slug, convly_url, developer, model_type, modality, parameters, context_window, max_output, license, open_weights, release_date, input_price (USD/1M tokens), output_price (USD/1M tokens)… See the full description on the dataset page: https://huggingface.co/datasets/sakd99/ai-models-database.validation-datatest-aimotraining-data-oss120baimo-interp-challenge-sample-v2-full
