CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01analytics-agents-uncertainty /da-code-evaluation-results0 likes10k downloads8mo agoHugging Face02dacorvo /funes-nvidia-Open-SWE-Traces Funes recall store — NVIDIA Open-SWE-Traces (resolved) A funes recall store built by indexing the resolved==1 trajectories of nvidia/Open-SWE-Traces (65244 sessions, across both harnesses — SWE-agent and OpenHands — and both models, Minimax-M2.5 and Qwen3.5-122B). What this is This is not a raw trace dataset — it is a pre-built funes index: the source trajectories chunked into content blocks and embedded, stored as a Lance table (chunks.lance). Source… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-nvidia-Open-SWE-Traces.10M<n<100M0 likes4.5k downloads2mo agoHugging Face03DAComp /dacomp-da-zh-eval DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle ✍️ Citation If you find our work helpful, please cite as @misc{lei2025dacompbenchmarkingdataagents, title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle}, author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da-zh-eval.imagen<1K2 likes3.6k downloads10mo agoHugging Face04dacorvo /funes-handoff-recall-benchmark handover-vs-recall A long investigation bloats an agent session until each new turn costs more to carry the context than to do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task, on tasks that genuinely require the prior investigation: arm channel A branch-only switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.tabularn<1K0 likes3.1k downloads19d agoHugging Face05dachengzisks /modified_libero_rlds Modified LIBERO RLDS Datasets This repository contains the four modified LIBERO datasets used in the OpenVLA fine-tuning experiments, stored in RLDS data format. See Appendix E in the OpenVLA paper for details about the fine-tuning experiments and specific dataset modifications, and see the OpenVLA GitHub README for instructions on how to run OpenVLA in LIBERO environments. Citation BibTeX: @article{kim24openvla, title={OpenVLA: An Open-Source… See the full description on the dataset page: https://huggingface.co/datasets/dachengzisks/modified_libero_rlds.0 likes3k downloads8mo agoHugging Face06DAComp /dacomp-da-eval DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle ✍️ Citation If you find our work helpful, please cite as @misc{lei2025dacompbenchmarkingdataagents, title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle}, author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da-eval.imagen<1K0 likes3k downloads10mo agoHugging Face07Voxel51 /dacl10k Dataset Card for dacl10k dacl10k stands for damage classification 10k images and is a multi-label semantic segmentation dataset for 19 classes (13 damages and 6 objects) present on bridges. The dacl10k dataset includes images collected during concrete bridge inspections acquired from databases at authorities and engineering offices, thus, it represents real-world scenarios. Concrete bridges represent the most common building type, besides steel, steel composite, and wooden bridges.… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/dacl10k.imageimage-classification1K<n<10K5 likes2.3k downloads2y agoHugging Face08dacorvo /funes-xiaowu0162-longmemeval-cleaned-s Funes recall store — LongMemEval_s cleaned corpus A funes recall store built by indexing the longmemeval_s_cleaned.json haystack of xiaowu0162/longmemeval-cleaned (LongMemEval, arXiv:2410.10813) — every unique chat session across all 500 questions' haystacks, in one corpus-wide store. What this is This is not a raw trace dataset — it is a pre-built funes index: the source sessions chunked into content blocks and embedded, stored as a Lance table (chunks.lance).… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-xiaowu0162-longmemeval-cleaned-s.tabular100K<n<1M0 likes1.4k downloads2mo agoHugging Face09TTS-AGI /commonvoice22-sidon-dacvae CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents Source sarulab-speech/commonvoice22_sidon Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.audioautomatic-speech-recognition1M<n<10M1 likes924 downloads6mo agoHugging Face10anthony-wss /librispeech_asr-audiodec_dac_16k Dataset Card for "librispeech_asr-audiodec_dac_16k" More Information needed text100K<n<1M0 likes674 downloads3y agoHugging Face11DAComp /dacomp-da-zh DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle Paper | Project Page | Code This repository contains DAComp, a benchmark of 210 tasks that mirrors complex real-world enterprise data intelligence workflows. It includes: Data Engineering (DE) tasks: Require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements. Data Analysis (DA)… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da-zh.texttext-generationn<1K0 likes664 downloads10mo agoHugging Face12dacthai2807 /ViMed-PET-part1 Dataset description for three years: 2017, 2018, 2019 This dataset contains data from three years (2017, 2018, 2019). Each year has several month folders, which are named as THANG {month}. Each year folder is compressed into zip files (chunks), each with an average size of approximately 2.5 GB. Please unzip the .zip files to fully extract all data folders. Folder structure after extraction Each folder named THANG {month} of a year is divided into 3 subfolders… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2807/ViMed-PET-part1.text1K<n<10K0 likes649 downloads1y agoHugging Face13dacoolkid44 /Data2 likes643 downloads2y agoHugging Face14MichaelMedek /dach_bike_graph DACH Bike + Rail Routing Graph A prebuilt, ready-to-route cycling + railway graph covering Germany, Austria, and Switzerland (DACH), stored as lat/lon-tiled GeoParquet. Built from OpenStreetMap by the Bike Route Optimizer for flat-preferring, surface-aware bike routing that can also hop on a train uphill. Everything routing needs is baked in — node elevations and full 3D edge geometry — so an application downloads this once and routes offline, with no Overpass and no elevation… See the full description on the dataset page: https://huggingface.co/datasets/MichaelMedek/dach_bike_graph.tabularother10M<n<100M1 likes616 downloads1mo agoHugging Face15dacorvo /transformers-coding-session-pi-traces dacorvo/transformers-coding-session-pi-traces pi coding-agent session traces produced by agentcap runs. Each run contributes one folder under data/<run_id>/; inside, one file per session in pi's native export format. The on-the-wire HTTP captures for these same runs live in dacorvo/transformers-coding-session-captures. Both belong to the transformers-coding-session Collection — join on run_id to align captures with traces. tabularn<1K0 likes606 downloads4mo agoHugging Face16dacorvo /hf-hub-session-pi-traces dacorvo/hf-hub-session-pi-traces pi coding-agent session traces produced by agentcap runs. Each run contributes one folder under data/<run_id>/; inside, one file per session in pi's native export format. The on-the-wire HTTP captures for these same runs live in dacorvo/hf-hub-session-captures. Both belong to the hf-hub-session Collection — join on run_id to align captures with traces. tabularn<1K0 likes570 downloads4mo agoHugging Face17dacorvo /transformers-gh-memory huggingface/transformers issues and pull requests, as a funes memory Every issue and pull request of huggingface/transformers with activity since 2024-01-01 — opening bodies, comments, reviews, inline review comments and PR diffs — chunked, embedded and written to a Lance table by funes, so the tracker can be searched by meaning and read back thread by thread. Kept fresh every few minutes by the funes-github Space. Use it Set funes up for your agent the usual way… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/transformers-gh-memory.100K<n<1M0 likes496 downloads3m agoHugging Face18dac-research /longbench_synthetic_v4_1 LongBench Synthetic V4.1 Dataset statistics v4.1 = lbs_v4 train+val (verbatim) + a test split absorbing every config from dac-research/extra_evals_v1 not already in v4. Five ZeroScrolls configs whose upstream corpora collide with lbs_v3.1 train/test (gov_report, qmsum, qasper, narrative_qa, musique) were dropped outright. All test rows have rubrics backfilled via the longbench_generate_rubric_for_imported.jinja template (same path as lbs_v3.1 / lbs_v4 Stage 1b).… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v4_1.tabular10K<n<100K1 likes469 downloads4mo agoHugging Face19jhoncrad /dacl10k Dataset Card for dacl10k dacl10k stands for damage classification 10k images and is a multi-label semantic segmentation dataset for 19 classes (13 damages and 6 objects) present on bridges. The dacl10k dataset includes images collected during concrete bridge inspections acquired from databases at authorities and engineering offices, thus, it represents real-world scenarios. Concrete bridges represent the most common building type, besides steel, steel composite, and wooden… See the full description on the dataset page: https://huggingface.co/datasets/jhoncrad/dacl10k.imageimage-classification1K<n<10K0 likes428 downloads4mo agoHugging Face20dacthai2k /ViMed-PET-part3 Dataset description for year 2023 This dataset contains data from three months: October, November, and December, stored in the following folders respectively: THANG 10 THANG 11 THANG 12 The data is compressed into zip files (chunks), each with an average size of approximately 2.5 GB. Please unzip the .zip files to fully extract the data folders. Folder structure after extraction Each folder named THANG {month} is divided into 3 subfolders, corresponding to 2… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2k/ViMed-PET-part3.text1K<n<10K0 likes400 downloads1y agoHugging Face21anthony-wss /librispeech_asr-audiodec_dac_24ktext100K<n<1M0 likes397 downloads3y agoHugging Face22DAComp /dacomp-de-gold DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle ✍️ Citation If you find our work helpful, please cite as @misc{lei2025dacompbenchmarkingdataagents, title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle}, author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-de-gold.0 likes333 downloads10mo agoHugging Face23DAComp /dacomp-da DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle ✍️ Citation If you find our work helpful, please cite as @misc{lei2025dacompbenchmarkingdataagents, title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle}, author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da.textn<1K6 likes330 downloads10mo agoHugging Face24treadon /speech-dac-tokens-3cb Speech DAC Tokens (3 Codebooks) Pre-tokenized speech dataset using the Descript Audio Codec (DAC). Each audio clip has been encoded into discrete codebook tokens from DAC's first 3 residual vector quantization codebooks, paired with its text transcription. Dataset Summary Stat Value Total samples 241,451 Total audio ~780 hours Language English Codebooks 3 (of DAC's 9) Codebook size 1,024 entries each DAC model 44kHz Tokens per second ~258 (86 frames… See the full description on the dataset page: https://huggingface.co/datasets/treadon/speech-dac-tokens-3cb.tabulartext-to-speech100K<n<1M0 likes297 downloads6mo agoHugging Face25jjjsadhfgj /dacomp-da-zh-eval DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle ✍️ Citation If you find our work helpful, please cite as @misc{lei2025dacompbenchmarkingdataagents, title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle}, author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/jjjsadhfgj/dacomp-da-zh-eval.imagen<1K0 likes282 downloads9mo agoHugging Face26dacorvo /transformers-coding-session-captures dacorvo/transformers-coding-session-captures HTTP captures of agent ↔ model interactions — one parquet row per /v1/chat/completions call. Produced by agentcap. Native session traces for the same runs live in companion datasets named transformers-coding-session-<agent>-traces. They're all grouped under the transformers-coding-session Collection alongside this dataset. Join on run_id. Loading from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/transformers-coding-session-captures.tabular10K<n<100K0 likes252 downloads3mo agoHugging Face27smallopen1145141919810 /dacl10k Dataset Card for dacl10k dacl10k stands for damage classification 10k images and is a multi-label semantic segmentation dataset for 19 classes (13 damages and 6 objects) present on bridges. The dacl10k dataset includes images collected during concrete bridge inspections acquired from databases at authorities and engineering offices, thus, it represents real-world scenarios. Concrete bridges represent the most common building type, besides steel, steel composite, and wooden… See the full description on the dataset page: https://huggingface.co/datasets/smallopen1145141919810/dacl10k.imageimage-classification1K<n<10K0 likes228 downloads3mo agoHugging Face28TTS-AGI /mls-enhanced-dacvae Multilingual LibriSpeech converted to DAC VAE latents Source facebook/multilingual_librispeech Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.audioautomatic-speech-recognition100K<n<1M0 likes217 downloads6mo agoHugging Face29ShantanuT01 /DACTYL DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large language models Dataset The DACTYL dataset is an AI-generated text detection dataset focusing primarily on one-shot or few-shot examples. We also include texts from continued pre-trained small language models. For more information, refer to our paper. Models Used We used the following LLMs to generate texts. OpenAI’s GPT-4o-mini and GPT-4o Anthropic’s Claude Haiku and Sonnet 3.5 Mistral Small (24B)and… See the full description on the dataset page: https://huggingface.co/datasets/ShantanuT01/DACTYL.tabulartext-classification100K<n<1M0 likes202 downloads1y agoHugging Face30TTS-AGI /maestrino-data-DACVAEtext1M<n<10M0 likes200 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.