amp
Datasets
All datasets matching “amp”Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.AmpScape
AmpScape
v1.0 (2026-09-23). Generated 2026-09-16 → 09-22 on Georgia Tech PACE-ICE with streaming, checksum-verified
upload; every tier passed a full Hub-vs-plan audit; the post-run precision pass (09-22/23) re-solved 129 722 rows so
that every T1/T1W/T1R/T3 row carries its true Kirchhoff residual. Pipeline tag v1.0-pipeline (freeze) and release tag
v1.0 (GitHub and Hub revision). Cost: 16 552 core-hours. Full account: docs/generation_postmortem.md.
AmpScape is a benchmark of… See the full description on the dataset page: https://huggingface.co/datasets/Xirro/AmpScape.euler-math-logs
euler-math
1. Evaluation
-
bash scripts/run_eval.sh
2. Scoring
root_path(str): Default value is results.
datasets(list): Dataset list to grade the response of models. All datasets in directory will be evaluated if None was given.
languages(list): Language list to grade the response of models. All languages in directory will be evaluated if None was given.
bash scripts/run_score.sh
3. Language Consistency Score(LCS)
Arguments
root_path(str):… See the full description on the dataset page: https://huggingface.co/datasets/amphora/euler-math-logs.hephaestus-ccx-runs-megarepohle-verified-shortformsh-prompt-analsis
