CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amphion /Emilia-Datasetgated Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.audiotext-to-speech10M<n<100M489 likes46k downloads2y agoHugging Face02Xirro /AmpScape AmpScape v1.0 (2026-09-23). Generated 2026-09-16 → 09-22 on Georgia Tech PACE-ICE with streaming, checksum-verified upload; every tier passed a full Hub-vs-plan audit; the post-run precision pass (09-22/23) re-solved 129 722 rows so that every T1/T1W/T1R/T3 row carries its true Kirchhoff residual. Pipeline tag v1.0-pipeline (freeze) and release tag v1.0 (GitHub and Hub revision). Cost: 16 552 core-hours. Full account: docs/generation_postmortem.md. AmpScape is a benchmark of… See the full description on the dataset page: https://huggingface.co/datasets/Xirro/AmpScape.image-to-image100K<n<1M0 likes15k downloads2h agoHugging Face03amphora /euler-math-logs euler-math 1. Evaluation - bash scripts/run_eval.sh 2. Scoring root_path(str): Default value is results. datasets(list): Dataset list to grade the response of models. All datasets in directory will be evaluated if None was given. languages(list): Language list to grade the response of models. All languages in directory will be evaluated if None was given. bash scripts/run_score.sh 3. Language Consistency Score(LCS) Arguments root_path(str):… See the full description on the dataset page: https://huggingface.co/datasets/amphora/euler-math-logs.0 likes4.9k downloads2y agoHugging Face04amphora /hephaestus-ccx-runs-megarepo1 likes1.7k downloads2mo agoHugging Face05amphora /hle-verified-shortformtextn<1K0 likes1.6k downloads9d agoHugging Face06amphora /sh-prompt-analsis0 likes827 downloads8mo agoHugging Face07amphora /ResearchMath-14k ResearchMath-14k ResearchMath-14k is a collection of 14,056 research-level mathematical problem records extracted from papers, open-problem lists, workshop sheets, and related academic sources. Each record contains the original extracted question, a rewritten self-contained problem statement, taxonomy labels, and open-status metadata. Paper: ResearchMath-14K: Scaling Research-Level Mathematics via Agents Load from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/amphora/ResearchMath-14k.texttext-generation10K<n<100K57 likes633 downloads4mo agoHugging Face08haifan-gong /Tiger-AMP Tiger-AMP Data and ablation metadata/results for TIGER / Tiger-AMP. Layout data/wetlab — wet-lab species training CSVs and manifest data/trainval_dbassp — DBAASP train/val labels/provenance (PDB zip omitted here; binaries need Xet/LFS) data/test_external — external test sets models/runs_ablation — ablation configs/results/logs (checkpoint .pt binaries omitted pending Xet upload) models/runs_ablation_toxin — toxin ablation metrics/logs (.joblib checkpoints omitted… See the full description on the dataset page: https://huggingface.co/datasets/haifan-gong/Tiger-AMP.0 likes466 downloads2mo agoHugging Face09amphora /math-utility-db0 likes434 downloads9mo agoHugging Face10ZihengZhou06 /AMPBench-MT AMPBench-MT AMPBench-MT is a homology-controlled benchmark for antimicrobial peptide endpoint prediction. The release is dated 2026-07-08. Repository: https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT The benchmark is organized around endpoint-aware prediction rather than binary AMP recognition alone. It contains processed task tables for AMP/non-AMP classification, species-conditioned MIC regression, activity spectrum positive-evidence audits, low-toxicity classification… See the full description on the dataset page: https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT.tabulartabular-classification100K<n<1M0 likes430 downloads2mo agoHugging Face11amphora /bon-resultstext10K<n<100K0 likes396 downloads2y agoHugging Face12tradecatlabs /btcusdt-amplitude-top100 BTCUSDT 全周期单根振幅 Top100 数据集 赞助商列表 交易猫(TradeCat):本项目的核心赞助方与长期支持方 本项目由 交易猫(TradeCat) 赞助与支持。交易猫为项目持续提供资金、社区与分发支持,帮助本项目保持开源迭代与长期维护。 交易猫 CA:0x8a99b8d53eff6bc331af529af74ad267f3167777 数据集简介 这是一份面向研究用途的快照型数据集,内容来自 TradeCat / apps/research 中的 btcusdt-um-amplitude-top100 工作区。 核心目标: 提供 BTCUSDT 在 1m / 5m / 15m / 1h / 4h / 1d / 1w 七个周期上的单根 K 线振幅 Top100 榜单 提供对应的描述统计、集中度统计、覆盖范围统计、牛熊阶段统计与字段字典 提供事件窗口上下文与 forward return 观察样本 提供一份可本地打开的 HTML… See the full description on the dataset page: https://huggingface.co/datasets/tradecatlabs/btcusdt-amplitude-top100.1K<n<10K0 likes393 downloads5mo agoHugging Face13fedehorl /stanford-dogs-amplified Stanford Dogs Amplified (Parquet Edition) This dataset is an amplified and modernized version of the classic Stanford Dogs Dataset. It builds upon the original 120 dog breeds by automatically fetching, cleaning, and injecting thousands of new high-quality images scraped from Bing, filtered dynamically via YOLO object detection. This specific repository hosts the dataset natively in Hugging Face's optimized Parquet format. This means: It is a strictly "Image Classification" dataset… See the full description on the dataset page: https://huggingface.co/datasets/fedehorl/stanford-dogs-amplified.imageimage-classification10K<n<100K0 likes339 downloads6mo agoHugging Face14amphora /ResearchMath-Reasoning-194K ResearchMath-Reasoning-194K ResearchMath-Reasoning-194K is a collection of 193,938 long-form reasoning traces and solutions for research-level mathematical problems, released alongside ResearchMath-14k as part of the same paper. While ResearchMath-14k provides the curated problem statements, this dataset provides model-generated solution attempts: each record contains a self-contained problem statement, a long chain-of-thought reasoning trace, and a final response. Paper:… See the full description on the dataset page: https://huggingface.co/datasets/amphora/ResearchMath-Reasoning-194K.texttext-generation100K<n<1M7 likes336 downloads4mo agoHugging Face15UM-IHC-CA2i /AMPLIFAIgated Dataset access is gated — registration required. Clicking Request access on Hugging Face is not enough. You must complete the official registration form, agree to the Terms & Conditions, and wait for approval before you can download data. Follow the starter kit step by step to get set up correctly. Registration closes September 1, 2026. @ Annotated Multi-Phase Liver Imaging for AI Official training data repository for the AMPLIFAI Challenge — a MICCAI 2026… See the full description on the dataset page: https://huggingface.co/datasets/UM-IHC-CA2i/AMPLIFAI.image-classification1K<n<10K8 likes334 downloads1mo agoHugging Face16amphora /finqa_suitetext1K<n<10K1 likes316 downloads3y agoHugging Face17amphora /MCLM Multilingual Competition Level Math (MCLM) Link to Paper: https://arxiv.org/abs/2502.17407 Overview:MCLM is a benchmark designed to evaluate advanced mathematical reasoning in a multilingual context. It features competition-level math problems across 55 languages, moving beyond standard word problems to challenge even state-of-the-art large language models. Dataset Composition MCLM is constructed from two main types of reasoning problems: Machine-translated… See the full description on the dataset page: https://huggingface.co/datasets/amphora/MCLM.textn<1K2 likes295 downloads2y agoHugging Face18am-pranav /stt-unified-bench am-pranav/stt-unified-bench Private, language/locale-partitioned mini-benchmark for STT models. Each subset is a dataset config (e.g., en, de, en_indian_accent, hi_in) with a single split val. Audio is staged at 16 kHz and stored in-repo for reproducibility. Schema audio : Audio(sampling_rate=16000, decode=False) text : reference transcription lang : implied by dataset config name source : upstream dataset tag id : source-stable id ⚠️ For internal evaluation only.… See the full description on the dataset page: https://huggingface.co/datasets/am-pranav/stt-unified-bench.audioautomatic-speech-recognition10K<n<100K0 likes276 downloads1y agoHugging Face19amphion /Emilia-NVgated NVSpeech Dataset Overview The NVSpeech dataset provides extensive annotations of paralinguistic vocalizations for Mandarin Chinese speech, aimed at enhancing the capabilities of automatic speech recognition (ASR) and text-to-speech (TTS) systems. The dataset features explicit word-level annotations for 18 categories of paralinguistic vocalizations, including non-verbal sounds like laughter and breathing, as well as lexicalized interjections like "uhm" and "oh."… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-NV.audiotext-to-speech100K<n<1M52 likes269 downloads1y agoHugging Face20amphora /QwQ-LongCoT-130KAlso have a look on the second version here => QwQ-LongCoT-2 Figure 1: Just a cute picture generate with [Flux](https://huggingface.co/Shakker-Labs/FLUX.1-dev-LoRA-Logo-Design) Today, I’m excited to release QwQ-LongCoT-130K, a SFT dataset designed for training O1-like large language models (LLMs). This dataset includes about 130k instances, each with responses generated using QwQ-32B-Preview. The dataset is available under the Apache 2.0 license, so feel free to use it as you like.… See the full description on the dataset page: https://huggingface.co/datasets/amphora/QwQ-LongCoT-130K.texttext-generation100K<n<1M153 likes258 downloads2y agoHugging Face21amphora /owm-transtext1M<n<10M0 likes245 downloads2y agoHugging Face22amphora /FC-Text-to-JSON-150ktext100K<n<1M1 likes237 downloads8mo agoHugging Face23AMP4010 /Historical_Nifty_50_Constituent_Weights_20YSUMMARY & CONTEXT: This dataset aims to provide a comprehensive, rolling 20-year history of the constituent stocks and their corresponding weights in India's Nifty 50 index. The data begins on January 31, 2008, and is actively maintained with monthly updates. After hitting the 20-year mark, as new monthly data is added, the oldest month's data will be removed to maintain a consistent 20-year window. This dataset was developed as a foundational feature for a graph-based model analyzing the… See the full description on the dataset page: https://huggingface.co/datasets/AMP4010/Historical_Nifty_50_Constituent_Weights_20Y.texttime-series-forecastingn<1K0 likes215 downloads1y agoHugging Face24amphion /SingVERSE SingVERSE: A Diverse, Real-World Benchmark for Singing Voice Enhancement SingVERSE is the first real-world benchmark for singing voice enhancement, created to address the critical lack of realistic evaluation data. It provides a foundational benchmark for developing and evaluating singing voice enhancement models. The dataset consists of 3,929 audio pairs, totaling 18.14 hours. It spans 19 distinct and diverse real-world acoustic scenarios, from reverberant concert halls to noisy… See the full description on the dataset page: https://huggingface.co/datasets/amphion/SingVERSE.audio-to-audio1K<n<10K5 likes201 downloads1y agoHugging Face25amphora /dasd-stage1-50k DASD stage1 - 50k length-filtered subset A 50,000-example subset of the stage1 (low-temperature) config of Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b. Columns are input / output; output is the verbatim gpt-oss-120b <think> reasoning trace. How it was built Started from stage1 (104,829 rows). Applied the Qwen3-4B-Instruct-2507 chat template and tokenized the full formatted conversation, then dropped every example over 65,536 tokens (the 64K training… See the full description on the dataset page: https://huggingface.co/datasets/amphora/dasd-stage1-50k.texttext-generation10K<n<100K0 likes189 downloads2mo agoHugging Face26amphora /q32-yisang-rmtext100K<n<1M0 likes183 downloads1y agoHugging Face27xiuyuz /ample-math AMPLE-Math 5,319 mathematics problems, each with a verified final answer and six references to that same answer. The references differ only in how much of the reasoning they show, which makes them useful for studying what a teacher's reference content contributes during distillation. Problems and original reasoning come from the metadata configuration of OpenThoughts-114k, and keep its Apache-2.0 attribution. A question was kept only if all six references exist, every generated… See the full description on the dataset page: https://huggingface.co/datasets/xiuyuz/ample-math.tabulartext-generation1K<n<10K1 likes182 downloads4d agoHugging Face28edwinmeriaux /AMP2026 Citation If you use this dataset, please cite: @misc{meriaux2026amp2026multiplatformmarinerobotics, title={AMP2026: A Multi-Platform Marine Robotics Dataset for Tracking and Mapping}, author={Edwin Meriaux and Shuo Wen and David Widhalm and Zhizun Wang and Junming Shi and Mariana Sosa Guzmán and Kalvik Jakkala and Bennett Carley and Elias Sokolova and Yogesh Girdhar and Monika Roznere and Jason O'Kane and Junaed Sattar and Gregory Dudek}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/edwinmeriaux/AMP2026.image10K<n<100K1 likes177 downloads6mo agoHugging Face29amphora /owm-rm-3.2mtabular1M<n<10M1 likes168 downloads2y agoHugging Face30sarahpann /AMPStext100K<n<1M7 likes159 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.