datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dacomp-da-zh-eval
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
✍️ Citation
If you find our work helpful, please cite as
@misc{lei2025dacompbenchmarkingdataagents,
title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle},
author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da-zh-eval.funes-handoff-recall-benchmark
handover-vs-recall
A long investigation bloats an agent session until each new turn costs more to carry the context than to
do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the
ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task,
on tasks that genuinely require the prior investigation:
arm
channel
A branch-only
switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.dacomp-da-eval
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
✍️ Citation
If you find our work helpful, please cite as
@misc{lei2025dacompbenchmarkingdataagents,
title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle},
author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da-eval.funes-xiaowu0162-longmemeval-cleaned-s
Funes recall store — LongMemEval_s cleaned corpus
A funes recall store built by indexing the
longmemeval_s_cleaned.json haystack of
xiaowu0162/longmemeval-cleaned
(LongMemEval, arXiv:2410.10813) — every unique
chat session across all 500 questions' haystacks, in one corpus-wide store.
What this is
This is not a raw trace dataset — it is a pre-built funes index: the source
sessions chunked into content blocks and embedded, stored as a
Lance table (chunks.lance).… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-xiaowu0162-longmemeval-cleaned-s.commonvoice22-sidon-dacvae
CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents
Source
sarulab-speech/commonvoice22_sidon
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.librispeech_asr-audiodec_dac_16k
Dataset Card for "librispeech_asr-audiodec_dac_16k"
More Information needed
dacomp-da-zh
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
Paper | Project Page | Code
This repository contains DAComp, a benchmark of 210 tasks that mirrors complex real-world enterprise data intelligence workflows. It includes:
Data Engineering (DE) tasks: Require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements.
Data Analysis (DA)… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da-zh.dach_bike_graph
DACH Bike + Rail Routing Graph
A prebuilt, ready-to-route cycling + railway graph covering Germany, Austria, and Switzerland (DACH), stored as lat/lon-tiled GeoParquet. Built from OpenStreetMap by the Bike Route Optimizer for flat-preferring, surface-aware bike routing that can also hop on a train uphill.
Everything routing needs is baked in — node elevations and full 3D edge geometry — so an application downloads this once and routes offline, with no Overpass and no elevation… See the full description on the dataset page: https://huggingface.co/datasets/MichaelMedek/dach_bike_graph.ViMed-PET-part1
Dataset description for three years: 2017, 2018, 2019
This dataset contains data from three years (2017, 2018, 2019). Each year has several month folders, which are named as THANG {month}.
Each year folder is compressed into zip files (chunks), each with an average size of approximately 2.5 GB.
Please unzip the .zip files to fully extract all data folders.
Folder structure after extraction
Each folder named THANG {month} of a year is divided into 3 subfolders… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2807/ViMed-PET-part1.transformers-coding-session-pi-traces
dacorvo/transformers-coding-session-pi-traces
pi coding-agent session traces produced by
agentcap runs. Each run
contributes one folder under data/<run_id>/; inside, one file per
session in pi's native export format.
The on-the-wire HTTP captures for these same runs live in
dacorvo/transformers-coding-session-captures.
Both belong to the
transformers-coding-session Collection
— join on run_id to align captures with traces.
hf-hub-session-pi-traces
dacorvo/hf-hub-session-pi-traces
pi coding-agent session traces produced by
agentcap runs. Each run
contributes one folder under data/<run_id>/; inside, one file per
session in pi's native export format.
The on-the-wire HTTP captures for these same runs live in
dacorvo/hf-hub-session-captures.
Both belong to the
hf-hub-session Collection
— join on run_id to align captures with traces.
ViMed-PET-part3
Dataset description for year 2023
This dataset contains data from three months: October, November, and December, stored in the following folders respectively:
THANG 10
THANG 11
THANG 12
The data is compressed into zip files (chunks), each with an average size of approximately 2.5 GB.
Please unzip the .zip files to fully extract the data folders.
Folder structure after extraction
Each folder named THANG {month} is divided into 3 subfolders, corresponding to 2… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2k/ViMed-PET-part3.librispeech_asr-audiodec_dac_24klongbench_synthetic_v4_1
LongBench Synthetic V4.1
Dataset statistics
v4.1 = lbs_v4 train+val (verbatim) + a test split absorbing every config from dac-research/extra_evals_v1 not already in v4. Five ZeroScrolls configs whose upstream corpora collide with lbs_v3.1 train/test (gov_report, qmsum, qasper, narrative_qa, musique) were dropped outright. All test rows have rubrics backfilled via the longbench_generate_rubric_for_imported.jinja template (same path as lbs_v3.1 / lbs_v4 Stage 1b).… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v4_1.dacomp-da
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
✍️ Citation
If you find our work helpful, please cite as
@misc{lei2025dacompbenchmarkingdataagents,
title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle},
author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da.speech-dac-tokens-3cb
Speech DAC Tokens (3 Codebooks)
Pre-tokenized speech dataset using the Descript Audio Codec (DAC). Each audio clip has been encoded into discrete codebook tokens from DAC's first 3 residual vector quantization codebooks, paired with its text transcription.
Dataset Summary
Stat
Value
Total samples
241,451
Total audio
~780 hours
Language
English
Codebooks
3 (of DAC's 9)
Codebook size
1,024 entries each
DAC model
44kHz
Tokens per second
~258 (86 frames… See the full description on the dataset page: https://huggingface.co/datasets/treadon/speech-dac-tokens-3cb.dacomp-da-zh-eval
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
✍️ Citation
If you find our work helpful, please cite as
@misc{lei2025dacompbenchmarkingdataagents,
title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle},
author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/jjjsadhfgj/dacomp-da-zh-eval.transformers-coding-session-captures
dacorvo/transformers-coding-session-captures
HTTP captures of agent ↔ model interactions — one parquet row per
/v1/chat/completions call. Produced by
agentcap.
Native session traces for the same runs live in companion datasets
named transformers-coding-session-<agent>-traces. They're all grouped under the
transformers-coding-session Collection
alongside this dataset. Join on run_id.
Loading
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/transformers-coding-session-captures.mls-enhanced-dacvae
Multilingual LibriSpeech converted to DAC VAE latents
Source
facebook/multilingual_librispeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.DACTYL
DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large language models Dataset
The DACTYL dataset is an AI-generated text detection dataset focusing primarily on one-shot or few-shot examples. We also include texts from continued pre-trained small language models.
For more information, refer to our paper.
Models Used
We used the following LLMs to generate texts.
OpenAI’s GPT-4o-mini and GPT-4o
Anthropic’s Claude Haiku and Sonnet 3.5
Mistral Small (24B)and… See the full description on the dataset page: https://huggingface.co/datasets/ShantanuT01/DACTYL.maestrino-data-DACVAEbalanced-audio-snippets-40x3k-DACVAElongbench_synthetic_v3_1
LongBench Synthetic V3.1
Dataset statistics
Per-subset stats over uploaded samples. Unique ctx counts distinct context strings, while the context-length buckets count samples/rows. Token counts use Qwen/Qwen3-14B.
Pool
Subset
Unique ctx
Sample rows
<8K
8-16K
16-32K
>32K
Median tok
p90 tok
Max tok
eval
hotpotqa
200
200
27
110
63
0
14,982
16,947
17,578
eval
hotpotqa_e
286
286
116
139
31
0
9,575
16,434
17,322
eval
musique
200
200
3
46
151
0
16,733
17… See the full description on the dataset page: https://huggingface.co/datasets/dac-research/longbench_synthetic_v3_1.oecd-dac-crs
OECD DAC CRS Project titles and descriptions
All unique project titles and descriptions from the OECD DAC Creditor Reporting System (CRS). https://stats.oecd.org/Index.aspx?DataSetCode=crs1
text column is the concatenation of Project Title, Short Description, and Long Description, and is also the column on which duplicate projects were removed. Other columns are included for metadata purposes, or if you want to create a new text column as a concatenation of additional data.
training_datadac-sdc-2023
DAC System Design Contest 2023 Dataset
dataset is in coco format and with images ending in _*.jpg removed. I did not make any splits, in the effort to keep it close to the original.
dacy-data
Combined CDT, DDT and DaNE dataset
This dataset merges the Danish UD treebank (DDT), Danish Dependency Treebank (DaNE) and Copenhagen Dependency Treebank (CDT). The DDT contains part-of-speech, dependency and morphology tags and has been further annotated for entities by Alexandra Institute in DaNE. DDT is based on CDT to assign tags consistent with the universal dependencies project (UD). However, this process split the data in DDT into singular sentences, therefore models… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dacy-data.enhanced-audiosnippets-DACVAElibrispeech_asr-audiodec_dac_44kvocal-bursts-taxonomy-DACVAE
Vocal Bursts Taxonomy — DACVAE + MaestroClap Embeddings & Scores
Processed version of with DACVAE latents, MaestroClap embeddings, derived attribute/quality/speaker scores, and Gemini-verified labels.
Overview
Metric
Value
Total samples
16,175
Categories
82
Genders
male, female
Female samples
8,097
Male samples
8,078
Gemini Label Verification
Every sample was sent to Gemini 3.1 Flash Lite for two independent tasks:
Match scoring:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/vocal-bursts-taxonomy-DACVAE.
