datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.AmpScape
AmpScape
v1.0 (2026-09-23). Generated 2026-09-16 → 09-22 on Georgia Tech PACE-ICE with streaming, checksum-verified
upload; every tier passed a full Hub-vs-plan audit; the post-run precision pass (09-22/23) re-solved 129 722 rows so
that every T1/T1W/T1R/T3 row carries its true Kirchhoff residual. Pipeline tag v1.0-pipeline (freeze) and release tag
v1.0 (GitHub and Hub revision). Cost: 16 552 core-hours. Full account: docs/generation_postmortem.md.
AmpScape is a benchmark of… See the full description on the dataset page: https://huggingface.co/datasets/Xirro/AmpScape.euler-math-logs
euler-math
1. Evaluation
-
bash scripts/run_eval.sh
2. Scoring
root_path(str): Default value is results.
datasets(list): Dataset list to grade the response of models. All datasets in directory will be evaluated if None was given.
languages(list): Language list to grade the response of models. All languages in directory will be evaluated if None was given.
bash scripts/run_score.sh
3. Language Consistency Score(LCS)
Arguments
root_path(str):… See the full description on the dataset page: https://huggingface.co/datasets/amphora/euler-math-logs.hephaestus-ccx-runs-megarepohle-verified-shortformsh-prompt-analsisResearchMath-14k
ResearchMath-14k
ResearchMath-14k is a collection of 14,056 research-level mathematical problem records extracted from papers, open-problem lists, workshop sheets, and related academic sources. Each record contains the original extracted question, a rewritten self-contained problem statement, taxonomy labels, and open-status metadata.
Paper: ResearchMath-14K: Scaling Research-Level Mathematics via Agents
Load
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/amphora/ResearchMath-14k.Tiger-AMP
Tiger-AMP
Data and ablation metadata/results for TIGER / Tiger-AMP.
Layout
data/wetlab — wet-lab species training CSVs and manifest
data/trainval_dbassp — DBAASP train/val labels/provenance (PDB zip omitted here; binaries need Xet/LFS)
data/test_external — external test sets
models/runs_ablation — ablation configs/results/logs (checkpoint .pt binaries omitted pending Xet upload)
models/runs_ablation_toxin — toxin ablation metrics/logs (.joblib checkpoints omitted… See the full description on the dataset page: https://huggingface.co/datasets/haifan-gong/Tiger-AMP.math-utility-dbAMPBench-MT
AMPBench-MT
AMPBench-MT is a homology-controlled benchmark for antimicrobial peptide endpoint prediction. The release is dated 2026-07-08.
Repository: https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT
The benchmark is organized around endpoint-aware prediction rather than binary AMP recognition alone. It contains processed task tables for AMP/non-AMP classification, species-conditioned MIC regression, activity spectrum positive-evidence audits, low-toxicity classification… See the full description on the dataset page: https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT.bon-resultsbtcusdt-amplitude-top100
BTCUSDT 全周期单根振幅 Top100 数据集
赞助商列表
交易猫(TradeCat):本项目的核心赞助方与长期支持方
本项目由 交易猫(TradeCat) 赞助与支持。交易猫为项目持续提供资金、社区与分发支持,帮助本项目保持开源迭代与长期维护。
交易猫 CA:0x8a99b8d53eff6bc331af529af74ad267f3167777
数据集简介
这是一份面向研究用途的快照型数据集,内容来自 TradeCat / apps/research 中的 btcusdt-um-amplitude-top100 工作区。
核心目标:
提供 BTCUSDT 在 1m / 5m / 15m / 1h / 4h / 1d / 1w 七个周期上的单根 K 线振幅 Top100 榜单
提供对应的描述统计、集中度统计、覆盖范围统计、牛熊阶段统计与字段字典
提供事件窗口上下文与 forward return 观察样本
提供一份可本地打开的 HTML… See the full description on the dataset page: https://huggingface.co/datasets/tradecatlabs/btcusdt-amplitude-top100.stanford-dogs-amplified
Stanford Dogs Amplified (Parquet Edition)
This dataset is an amplified and modernized version of the classic Stanford Dogs Dataset. It builds upon the original 120 dog breeds by automatically fetching, cleaning, and injecting thousands of new high-quality images scraped from Bing, filtered dynamically via YOLO object detection.
This specific repository hosts the dataset natively in Hugging Face's optimized Parquet format.
This means:
It is a strictly "Image Classification" dataset… See the full description on the dataset page: https://huggingface.co/datasets/fedehorl/stanford-dogs-amplified.ResearchMath-Reasoning-194K
ResearchMath-Reasoning-194K
ResearchMath-Reasoning-194K is a collection of 193,938 long-form reasoning traces and solutions for research-level mathematical problems, released alongside ResearchMath-14k as part of the same paper. While ResearchMath-14k provides the curated problem statements, this dataset provides model-generated solution attempts: each record contains a self-contained problem statement, a long chain-of-thought reasoning trace, and a final response.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/amphora/ResearchMath-Reasoning-194K.AMPLIFAI
Dataset access is gated — registration required.
Clicking Request access on Hugging Face is not enough.
You must complete the official
registration form,
agree to the Terms & Conditions, and wait for approval before you can download data.
Follow the
starter kit
step by step to get set up correctly.
Registration closes September 1, 2026.
@
Annotated Multi-Phase Liver Imaging for AI
Official training data repository for the AMPLIFAI Challenge — a MICCAI 2026… See the full description on the dataset page: https://huggingface.co/datasets/UM-IHC-CA2i/AMPLIFAI.finqa_suiteMCLM
Multilingual Competition Level Math (MCLM)
Link to Paper: https://arxiv.org/abs/2502.17407
Overview:MCLM is a benchmark designed to evaluate advanced mathematical reasoning in a multilingual context. It features competition-level math problems across 55 languages, moving beyond standard word problems to challenge even state-of-the-art large language models.
Dataset Composition
MCLM is constructed from two main types of reasoning problems:
Machine-translated… See the full description on the dataset page: https://huggingface.co/datasets/amphora/MCLM.stt-unified-bench
am-pranav/stt-unified-bench
Private, language/locale-partitioned mini-benchmark for STT models.
Each subset is a dataset config (e.g., en, de, en_indian_accent, hi_in) with a single split val.
Audio is staged at 16 kHz and stored in-repo for reproducibility.
Schema
audio : Audio(sampling_rate=16000, decode=False)
text : reference transcription
lang : implied by dataset config name
source : upstream dataset tag
id : source-stable id
⚠️ For internal evaluation only.… See the full description on the dataset page: https://huggingface.co/datasets/am-pranav/stt-unified-bench.Emilia-NV
NVSpeech Dataset
Overview
The NVSpeech dataset provides extensive annotations of paralinguistic vocalizations for Mandarin Chinese speech, aimed at enhancing the capabilities of automatic speech recognition (ASR) and text-to-speech (TTS) systems. The dataset features explicit word-level annotations for 18 categories of paralinguistic vocalizations, including non-verbal sounds like laughter and breathing, as well as lexicalized interjections like "uhm" and "oh."… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-NV.QwQ-LongCoT-130KAlso have a look on the second version here => QwQ-LongCoT-2
Figure 1: Just a cute picture generate with [Flux](https://huggingface.co/Shakker-Labs/FLUX.1-dev-LoRA-Logo-Design)
Today, I’m excited to release QwQ-LongCoT-130K, a SFT dataset designed for training O1-like large language models (LLMs). This dataset includes about 130k instances, each with responses generated using QwQ-32B-Preview. The dataset is available under the Apache 2.0 license, so feel free to use it as you like.… See the full description on the dataset page: https://huggingface.co/datasets/amphora/QwQ-LongCoT-130K.owm-transFC-Text-to-JSON-150kHistorical_Nifty_50_Constituent_Weights_20YSUMMARY & CONTEXT:
This dataset aims to provide a comprehensive, rolling 20-year history of the constituent stocks and their corresponding weights in India's Nifty 50 index. The data begins on January 31, 2008, and is actively maintained with monthly updates. After hitting the 20-year mark, as new monthly data is added, the oldest month's data will be removed to maintain a consistent 20-year window. This dataset was developed as a foundational feature for a graph-based model analyzing the… See the full description on the dataset page: https://huggingface.co/datasets/AMP4010/Historical_Nifty_50_Constituent_Weights_20Y.SingVERSE
SingVERSE: A Diverse, Real-World Benchmark for Singing Voice Enhancement
SingVERSE is the first real-world benchmark for singing voice enhancement, created to address the critical lack of realistic evaluation data. It provides a foundational benchmark for developing and evaluating singing voice enhancement models.
The dataset consists of 3,929 audio pairs, totaling 18.14 hours. It spans 19 distinct and diverse real-world acoustic scenarios, from reverberant concert halls to noisy… See the full description on the dataset page: https://huggingface.co/datasets/amphion/SingVERSE.dasd-stage1-50k
DASD stage1 - 50k length-filtered subset
A 50,000-example subset of the stage1 (low-temperature) config of
Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b.
Columns are input / output; output is the verbatim gpt-oss-120b <think> reasoning trace.
How it was built
Started from stage1 (104,829 rows).
Applied the Qwen3-4B-Instruct-2507 chat template and tokenized the full formatted
conversation, then dropped every example over 65,536 tokens (the 64K training… See the full description on the dataset page: https://huggingface.co/datasets/amphora/dasd-stage1-50k.q32-yisang-rmample-math
AMPLE-Math
5,319 mathematics problems, each with a verified final answer and six references to that same
answer. The references differ only in how much of the reasoning they show, which makes them useful
for studying what a teacher's reference content contributes during distillation.
Problems and original reasoning come from the metadata configuration of
OpenThoughts-114k, and keep its
Apache-2.0 attribution. A question was kept only if all six references exist, every generated… See the full description on the dataset page: https://huggingface.co/datasets/xiuyuz/ample-math.AMP2026
Citation
If you use this dataset, please cite:
@misc{meriaux2026amp2026multiplatformmarinerobotics,
title={AMP2026: A Multi-Platform Marine Robotics Dataset for Tracking and Mapping},
author={Edwin Meriaux and Shuo Wen and David Widhalm and Zhizun Wang and Junming Shi and Mariana Sosa Guzmán and Kalvik Jakkala and Bennett Carley and Elias Sokolova and Yogesh Girdhar and Monika Roznere and Jason O'Kane and Junaed Sattar and Gregory Dudek},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/edwinmeriaux/AMP2026.owm-rm-3.2mAMPS
