datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
funes-handoff-recall-benchmark
handover-vs-recall
A long investigation bloats an agent session until each new turn costs more to carry the context than to
do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the
ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task,
on tasks that genuinely require the prior investigation:
arm
channel
A branch-only
switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.dacomp-da-zh
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
Paper | Project Page | Code
This repository contains DAComp, a benchmark of 210 tasks that mirrors complex real-world enterprise data intelligence workflows. It includes:
Data Engineering (DE) tasks: Require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements.
Data Analysis (DA)… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da-zh.ViMed-PET-part1
Dataset description for three years: 2017, 2018, 2019
This dataset contains data from three years (2017, 2018, 2019). Each year has several month folders, which are named as THANG {month}.
Each year folder is compressed into zip files (chunks), each with an average size of approximately 2.5 GB.
Please unzip the .zip files to fully extract all data folders.
Folder structure after extraction
Each folder named THANG {month} of a year is divided into 3 subfolders… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2807/ViMed-PET-part1.transformers-coding-session-pi-traces
dacorvo/transformers-coding-session-pi-traces
pi coding-agent session traces produced by
agentcap runs. Each run
contributes one folder under data/<run_id>/; inside, one file per
session in pi's native export format.
The on-the-wire HTTP captures for these same runs live in
dacorvo/transformers-coding-session-captures.
Both belong to the
transformers-coding-session Collection
— join on run_id to align captures with traces.
hf-hub-session-pi-traces
dacorvo/hf-hub-session-pi-traces
pi coding-agent session traces produced by
agentcap runs. Each run
contributes one folder under data/<run_id>/; inside, one file per
session in pi's native export format.
The on-the-wire HTTP captures for these same runs live in
dacorvo/hf-hub-session-captures.
Both belong to the
hf-hub-session Collection
— join on run_id to align captures with traces.
ViMed-PET-part3
Dataset description for year 2023
This dataset contains data from three months: October, November, and December, stored in the following folders respectively:
THANG 10
THANG 11
THANG 12
The data is compressed into zip files (chunks), each with an average size of approximately 2.5 GB.
Please unzip the .zip files to fully extract the data folders.
Folder structure after extraction
Each folder named THANG {month} is divided into 3 subfolders, corresponding to 2… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2k/ViMed-PET-part3.dacomp-da
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
✍️ Citation
If you find our work helpful, please cite as
@misc{lei2025dacompbenchmarkingdataagents,
title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle},
author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da.space-law-corpus
Space Law Corpus
A neutral, provenance-first, machine-readable record of international and national space law. Every record carries its official source, retrieval date, citation, language, an authoritative-status flag, and a SHA-256 content hash; texts are verified against official sources.
Source of truth / build history: https://github.com/dacheah/space-law-corpus
Archived & citable: concept DOI 10.5281/zenodo.21185483 (resolves to the latest Zenodo-archived GitHub release)… See the full description on the dataset page: https://huggingface.co/datasets/dacheah/space-law-corpus.funes-recall-session-pi-traces
dacorvo/funes-recall-session-pi-traces
pi coding-agent session traces produced by
agentcap runs. Each run
contributes one folder under data/<run_id>/; inside, one file per
session in pi's native export format.
The on-the-wire HTTP captures for these same runs live in
dacorvo/funes-recall-session-captures.
Both belong to the
funes-recall-session Collection
— join on run_id to align captures with traces.
bbnj-high-seas-treaty-corpus
BBNJ / High Seas Treaty Corpus
A neutral, provenance-first, machine-readable record of the 2023 BBNJ Agreement (the "High Seas Treaty", in force 17 January 2026) in all six authentic UN languages, and its implementing framework. Every record carries its official source, retrieval date, citation, authentic language, an authoritative-status flag, a SHA-256 content hash, and an honest per-language fidelity flag (extracted_verified → extracted_unverified → ocr_unverified).
Source… See the full description on the dataset page: https://huggingface.co/datasets/dacheah/bbnj-high-seas-treaty-corpus.deep-seabed-mining-law-corpus
Deep Seabed Mining Law Corpus
A neutral, provenance-first, machine-readable record of the law governing mineral resources of "the Area" (the seabed beyond national jurisdiction): the international ISA/UNCLOS regime and the US non-UNCLOS parallel track. Every record carries its official source, retrieval date, citation, language, an authoritative-status flag, and a SHA-256 content hash; texts are verified against official sources.
Source of truth / build history:… See the full description on the dataset page: https://huggingface.co/datasets/dacheah/deep-seabed-mining-law-corpus.dacomp-da-zh
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
Paper | Project Page | Code
This repository contains DAComp, a benchmark of 210 tasks that mirrors complex real-world enterprise data intelligence workflows. It includes:
Data Engineering (DE) tasks: Require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements.
Data Analysis (DA)… See the full description on the dataset page: https://huggingface.co/datasets/jjjsadhfgj/dacomp-da-zh.funes-recall-session-hermes-traces
dacorvo/funes-recall-session-hermes-traces
hermes coding-agent session traces produced by
agentcap runs. Each run
contributes one folder under data/<run_id>/; inside, one file per
session in hermes's native export format.
The on-the-wire HTTP captures for these same runs live in
dacorvo/funes-recall-session-captures.
Both belong to the
funes-recall-session Collection
— join on run_id to align captures with traces.
dac6-instruct
DAC6 instruct (11-12-2023)
“DAC 6” refers to European Council Directive (EU) 2018/822 of May 25, 2018 relating to the automatic and mandatory exchange of information on cross-border arrangements requiring declaration. It aims to strengthen cooperation between tax administrations in EU countries on potentially aggressive tax planning arrangements.
This project focuses on fine-tuning pre-trained language models to create efficient and accurate models for tax practice.
Fine-tuning is… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/dac6-instruct.DACMini_Refined
Dataset di ricerca
DACMini_Refined è un dataset creato a scopo di ricerca e sviluppo per migliorare le capacità del modello compatto DACMini-IT, un modello linguistico italiano da 109 milioni di parametri.
L’obiettivo del dataset è incrementare la qualità delle risposte del modello di base attraverso un processo supervisionato multi-stadio, sfruttando modelli di dimensioni maggiori come generatore e validatore.
Metodologia di generazione
Generazione automatica di… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DACMini_Refined.SD4mindw96__DeepSeek-llama3.3-Bllossom-8B-DACON-LLM3-details
Dataset Card for Evaluation run of mindw96/DeepSeek-llama3.3-Bllossom-8B-DACON-LLM3
Dataset automatically created during the evaluation run of model mindw96/DeepSeek-llama3.3-Bllossom-8B-DACON-LLM3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mindw96__DeepSeek-llama3.3-Bllossom-8B-DACON-LLM3-details.yingjiyuanDAC-Think
DAC-Think
Dataset Name: DAC-ThinkCreator: MattimaxOrganization: MINCLicense: MITLanguage: ItalianoNumber of rows: 24,505
Overview
DAC-Think è un dataset di ragionamento esclusivamente in lingua italiana, progettato per task di generazione di testo e conversational AI. Ogni esempio contiene un prompt e una risposta strutturata, con tag <think> che evidenziano la parte di ragionamento del modello, seguita dalla risposta finale.
Il dataset è organizzato in questo… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DAC-Think.DAC-Reasoning-ITA
Descrizione del dataset
Questo dataset è stato generato sinteticamente da Mattia (“Mattimax”) per l’azienda M.INC.Serve per lo sviluppo e la valutazione di modelli in grado di ragionare in italiano e fornire risposte strutturate con tracciamento del ragionamento.I dati non sono garantiti accurati e sono destinati esclusivamente a scopi di ricerca e sperimentazione.
Fonte
Profilo autore: https://huggingface.co/Mattimax
Organizzazione: https://huggingface.co/MINC01… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DAC-Reasoning-ITA.dac-judge-v4-datasetxiaofang1
