CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DAComp /dacomp-da-zh-eval DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle ✍️ Citation If you find our work helpful, please cite as @misc{lei2025dacompbenchmarkingdataagents, title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle}, author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da-zh-eval.imagen<1K2 likes3.6k downloads10mo agoHugging Face02dacorvo /funes-handoff-recall-benchmark handover-vs-recall A long investigation bloats an agent session until each new turn costs more to carry the context than to do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task, on tasks that genuinely require the prior investigation: arm channel A branch-only switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.tabularn<1K0 likes3.1k downloads20d agoHugging Face03DAComp /dacomp-da-eval DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle ✍️ Citation If you find our work helpful, please cite as @misc{lei2025dacompbenchmarkingdataagents, title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle}, author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da-eval.imagen<1K0 likes2.9k downloads10mo agoHugging Face04dacorvo /funes-xiaowu0162-longmemeval-cleaned-s Funes recall store — LongMemEval_s cleaned corpus A funes recall store built by indexing the longmemeval_s_cleaned.json haystack of xiaowu0162/longmemeval-cleaned (LongMemEval, arXiv:2410.10813) — every unique chat session across all 500 questions' haystacks, in one corpus-wide store. What this is This is not a raw trace dataset — it is a pre-built funes index: the source sessions chunked into content blocks and embedded, stored as a Lance table (chunks.lance).… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-xiaowu0162-longmemeval-cleaned-s.tabular100K<n<1M0 likes1.4k downloads2mo agoHugging Face05TTS-AGI /commonvoice22-sidon-dacvae CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents Source sarulab-speech/commonvoice22_sidon Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.audioautomatic-speech-recognition1M<n<10M1 likes977 downloads6mo agoHugging Face06anthony-wss /librispeech_asr-audiodec_dac_16k Dataset Card for "librispeech_asr-audiodec_dac_16k" More Information needed text100K<n<1M0 likes668 downloads3y agoHugging Face07DAComp /dacomp-da-zh DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle Paper | Project Page | Code This repository contains DAComp, a benchmark of 210 tasks that mirrors complex real-world enterprise data intelligence workflows. It includes: Data Engineering (DE) tasks: Require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements. Data Analysis (DA)… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da-zh.texttext-generationn<1K0 likes664 downloads10mo agoHugging Face08MichaelMedek /dach_bike_graph DACH Bike + Rail Routing Graph A prebuilt, ready-to-route cycling + railway graph covering Germany, Austria, and Switzerland (DACH), stored as lat/lon-tiled GeoParquet. Built from OpenStreetMap by the Bike Route Optimizer for flat-preferring, surface-aware bike routing that can also hop on a train uphill. Everything routing needs is baked in — node elevations and full 3D edge geometry — so an application downloads this once and routes offline, with no Overpass and no elevation… See the full description on the dataset page: https://huggingface.co/datasets/MichaelMedek/dach_bike_graph.tabularother10M<n<100M1 likes661 downloads1mo agoHugging Face09dacthai2807 /ViMed-PET-part1 Dataset description for three years: 2017, 2018, 2019 This dataset contains data from three years (2017, 2018, 2019). Each year has several month folders, which are named as THANG {month}. Each year folder is compressed into zip files (chunks), each with an average size of approximately 2.5 GB. Please unzip the .zip files to fully extract all data folders. Folder structure after extraction Each folder named THANG {month} of a year is divided into 3 subfolders… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2807/ViMed-PET-part1.text1K<n<10K0 likes651 downloads1y agoHugging Face10dacorvo /transformers-coding-session-pi-traces dacorvo/transformers-coding-session-pi-traces pi coding-agent session traces produced by agentcap runs. Each run contributes one folder under data/<run_id>/; inside, one file per session in pi's native export format. The on-the-wire HTTP captures for these same runs live in dacorvo/transformers-coding-session-captures. Both belong to the transformers-coding-session Collection — join on run_id to align captures with traces. tabularn<1K0 likes605 downloads4mo agoHugging Face11dacorvo /hf-hub-session-pi-traces dacorvo/hf-hub-session-pi-traces pi coding-agent session traces produced by agentcap runs. Each run contributes one folder under data/<run_id>/; inside, one file per session in pi's native export format. The on-the-wire HTTP captures for these same runs live in dacorvo/hf-hub-session-captures. Both belong to the hf-hub-session Collection — join on run_id to align captures with traces. tabularn<1K0 likes572 downloads4mo agoHugging Face12dacthai2k /ViMed-PET-part3 Dataset description for year 2023 This dataset contains data from three months: October, November, and December, stored in the following folders respectively: THANG 10 THANG 11 THANG 12 The data is compressed into zip files (chunks), each with an average size of approximately 2.5 GB. Please unzip the .zip files to fully extract the data folders. Folder structure after extraction Each folder named THANG {month} is divided into 3 subfolders, corresponding to 2… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2k/ViMed-PET-part3.text1K<n<10K0 likes405 downloads1y agoHugging Face13anthony-wss /librispeech_asr-audiodec_dac_24ktext100K<n<1M0 likes395 downloads3y agoHugging Face14dac-research /longbench_synthetic_v4_1 LongBench Synthetic V4.1 Dataset statistics v4.1 = lbs_v4 train+val (verbatim) + a test split absorbing every config from dac-research/extra_evals_v1 not already in v4. Five ZeroScrolls configs whose upstream corpora collide with lbs_v3.1 train/test (gov_report, qmsum, qasper, narrative_qa, musique) were dropped outright. All test rows have rubrics backfilled via the longbench_generate_rubric_for_imported.jinja template (same path as lbs_v3.1 / lbs_v4 Stage 1b).… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v4_1.tabular10K<n<100K1 likes365 downloads4mo agoHugging Face15DAComp /dacomp-da DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle ✍️ Citation If you find our work helpful, please cite as @misc{lei2025dacompbenchmarkingdataagents, title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle}, author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da.textn<1K6 likes355 downloads10mo agoHugging Face16treadon /speech-dac-tokens-3cb Speech DAC Tokens (3 Codebooks) Pre-tokenized speech dataset using the Descript Audio Codec (DAC). Each audio clip has been encoded into discrete codebook tokens from DAC's first 3 residual vector quantization codebooks, paired with its text transcription. Dataset Summary Stat Value Total samples 241,451 Total audio ~780 hours Language English Codebooks 3 (of DAC's 9) Codebook size 1,024 entries each DAC model 44kHz Tokens per second ~258 (86 frames… See the full description on the dataset page: https://huggingface.co/datasets/treadon/speech-dac-tokens-3cb.tabulartext-to-speech100K<n<1M0 likes298 downloads6mo agoHugging Face17jjjsadhfgj /dacomp-da-zh-eval DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle ✍️ Citation If you find our work helpful, please cite as @misc{lei2025dacompbenchmarkingdataagents, title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle}, author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/jjjsadhfgj/dacomp-da-zh-eval.imagen<1K0 likes283 downloads9mo agoHugging Face18dacorvo /transformers-coding-session-captures dacorvo/transformers-coding-session-captures HTTP captures of agent ↔ model interactions — one parquet row per /v1/chat/completions call. Produced by agentcap. Native session traces for the same runs live in companion datasets named transformers-coding-session-<agent>-traces. They're all grouped under the transformers-coding-session Collection alongside this dataset. Join on run_id. Loading from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/transformers-coding-session-captures.tabular10K<n<100K0 likes247 downloads4mo agoHugging Face19TTS-AGI /mls-enhanced-dacvae Multilingual LibriSpeech converted to DAC VAE latents Source facebook/multilingual_librispeech Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.audioautomatic-speech-recognition100K<n<1M0 likes214 downloads6mo agoHugging Face20ShantanuT01 /DACTYL DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large language models Dataset The DACTYL dataset is an AI-generated text detection dataset focusing primarily on one-shot or few-shot examples. We also include texts from continued pre-trained small language models. For more information, refer to our paper. Models Used We used the following LLMs to generate texts. OpenAI’s GPT-4o-mini and GPT-4o Anthropic’s Claude Haiku and Sonnet 3.5 Mistral Small (24B)and… See the full description on the dataset page: https://huggingface.co/datasets/ShantanuT01/DACTYL.tabulartext-classification100K<n<1M0 likes207 downloads1y agoHugging Face21TTS-AGI /maestrino-data-DACVAEtext1M<n<10M0 likes202 downloads6mo agoHugging Face22TTS-AGI /balanced-audio-snippets-40x3k-DACVAEtext100K<n<1M0 likes183 downloads6mo agoHugging Face23dac-research /longbench_synthetic_v3_1 LongBench Synthetic V3.1 Dataset statistics Per-subset stats over uploaded samples. Unique ctx counts distinct context strings, while the context-length buckets count samples/rows. Token counts use Qwen/Qwen3-14B. Pool Subset Unique ctx Sample rows <8K 8-16K 16-32K >32K Median tok p90 tok Max tok eval hotpotqa 200 200 27 110 63 0 14,982 16,947 17,578 eval hotpotqa_e 286 286 116 139 31 0 9,575 16,434 17,322 eval musique 200 200 3 46 151 0 16,733 17… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v3_1.tabular10K<n<100K1 likes161 downloads4mo agoHugging Face24alex-miller /oecd-dac-crs OECD DAC CRS Project titles and descriptions All unique project titles and descriptions from the OECD DAC Creditor Reporting System (CRS). https://stats.oecd.org/Index.aspx?DataSetCode=crs1 text column is the concatenation of Project Title, Short Description, and Long Description, and is also the column on which duplicate projects were removed. Other columns are included for metadata purposes, or if you want to create a new text column as a concatenation of additional data. tabularmask-generation1M<n<10M0 likes156 downloads2y agoHugging Face25SpX-DAC /training_datatextn<1K0 likes141 downloads9mo agoHugging Face26Charitarth /dac-sdc-2023 DAC System Design Contest 2023 Dataset dataset is in coco format and with images ending in _*.jpg removed. I did not make any splits, in the effort to keep it close to the original. imageobject-detection10K<n<100K1 likes126 downloads3y agoHugging Face27chcaa /dacy-data Combined CDT, DDT and DaNE dataset This dataset merges the Danish UD treebank (DDT), Danish Dependency Treebank (DaNE) and Copenhagen Dependency Treebank (CDT). The DDT contains part-of-speech, dependency and morphology tags and has been further annotated for entities by Alexandra Institute in DaNE. DDT is based on CDT to assign tags consistent with the universal dependencies project (UD). However, this process split the data in DDT into singular sentences, therefore models… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dacy-data.texttoken-classification1K<n<10K3 likes123 downloads7d agoHugging Face28TTS-AGI /enhanced-audiosnippets-DACVAEtext1M<n<10M1 likes118 downloads6mo agoHugging Face29anthony-wss /librispeech_asr-audiodec_dac_44ktext100K<n<1M0 likes104 downloads3y agoHugging Face30TTS-AGI /vocal-bursts-taxonomy-DACVAE Vocal Bursts Taxonomy — DACVAE + MaestroClap Embeddings & Scores Processed version of with DACVAE latents, MaestroClap embeddings, derived attribute/quality/speaker scores, and Gemini-verified labels. Overview Metric Value Total samples 16,175 Categories 82 Genders male, female Female samples 8,097 Male samples 8,078 Gemini Label Verification Every sample was sent to Gemini 3.1 Flash Lite for two independent tasks: Match scoring:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/vocal-bursts-taxonomy-DACVAE.audioaudio-classificationn<1K0 likes100 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.