datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dacomp-da-zh
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
Paper | Project Page | Code
This repository contains DAComp, a benchmark of 210 tasks that mirrors complex real-world enterprise data intelligence workflows. It includes:
Data Engineering (DE) tasks: Require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements.
Data Analysis (DA)… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da-zh.dacomp-de
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
This repository contains the DAComp benchmark, presented in the paper DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle.
Project page: https://da-comp.github.io/
Code: https://github.com/anonymous/DAComp
✍️ Citation
If you find our work helpful, please cite as
@misc{lei2025dacompbenchmarkingdataagents,
title={DAComp: Benchmarking Data Agents across the Full… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-de.dacr-lt-training
DACR Recurrent-Depth Training Data
A large enriched reasoning corpus derived from the BOTCOIN/DACR data pipeline and adjusted for preliminary recurrent-depth natural-language experiments.
This dataset is not intended to be treated as a single fixed training split. It is better understood as a reusable source corpus containing several export categories that can be pruned, reshaped, and filtered depending on the training objective.
What Is Included… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/dacr-lt-training.DA-Code
[EMNLP2024] DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models
DA-Code is a comprehensive evaluation dataset designed to assess the data analysis and code generation capabilities of LLM in agent-based data science tasks. Our papers and experiment reports have been published on Arxiv.
Dataset Overview
500 complex real-world data analysis tasks across Data Wrangling (DW), Machine Learning (ML), and Exploratory Data Analysis (EDA).
Tasks cover… See the full description on the dataset page: https://huggingface.co/datasets/Luo2003/DA-Code.dacomp-da-zh
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
Paper | Project Page | Code
This repository contains DAComp, a benchmark of 210 tasks that mirrors complex real-world enterprise data intelligence workflows. It includes:
Data Engineering (DE) tasks: Require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements.
Data Analysis (DA)… See the full description on the dataset page: https://huggingface.co/datasets/jjjsadhfgj/dacomp-da-zh.DACTYL-Pretraining
DACTYL Pretraining Corpus
This corpus contains human texts that have passed a quality check from the textdescriptives library.
These texts have been used to further train various Llama 3.2 1B Instruct models by domain and split. For a domain/split combination, you can use the corresponding model:
ShantanuT01/fine-tuned-Llama-3.2-1B-Instruct-apollo-mini-{domain}-{split}
Citation
@misc{thorat2025dactyldiverseadversarialcorpus,
title={DACTYL: Diverse Adversarial… See the full description on the dataset page: https://huggingface.co/datasets/ShantanuT01/DACTYL-Pretraining.dactylic-hexameter-latin-poetry-corpus
Dactylic Hexameter Latin Poetry Corpus
This repository contains a curated and processed corpus of Classical Latin poetry written in dactylic hexameter. It serves as the raw training data ("Dataset V3") for the Master's Thesis titled "A Hybrid Post Hoc Feedback Framework for Latin Dactylic Hexameter" submitted to KU Leuven (2025).
Dataset Description
This corpus was constructed to fine-tune Large Language Models (LLMs) for the generation of metrically valid Latin poetry.… See the full description on the dataset page: https://huggingface.co/datasets/KaanGoker/dactylic-hexameter-latin-poetry-corpus.dac6-instruct
DAC6 instruct (11-12-2023)
“DAC 6” refers to European Council Directive (EU) 2018/822 of May 25, 2018 relating to the automatic and mandatory exchange of information on cross-border arrangements requiring declaration. It aims to strengthen cooperation between tax administrations in EU countries on potentially aggressive tax planning arrangements.
This project focuses on fine-tuning pre-trained language models to create efficient and accurate models for tax practice.
Fine-tuning is… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/dac6-instruct.DACMini_Refined
Dataset di ricerca
DACMini_Refined è un dataset creato a scopo di ricerca e sviluppo per migliorare le capacità del modello compatto DACMini-IT, un modello linguistico italiano da 109 milioni di parametri.
L’obiettivo del dataset è incrementare la qualità delle risposte del modello di base attraverso un processo supervisionato multi-stadio, sfruttando modelli di dimensioni maggiori come generatore e validatore.
Metodologia di generazione
Generazione automatica di… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DACMini_Refined.DAC-Think
DAC-Think
Dataset Name: DAC-ThinkCreator: MattimaxOrganization: MINCLicense: MITLanguage: ItalianoNumber of rows: 24,505
Overview
DAC-Think è un dataset di ragionamento esclusivamente in lingua italiana, progettato per task di generazione di testo e conversational AI. Ogni esempio contiene un prompt e una risposta strutturata, con tag <think> che evidenziano la parte di ragionamento del modello, seguita dalla risposta finale.
Il dataset è organizzato in questo… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DAC-Think.DAC-Reasoning-ITA
Descrizione del dataset
Questo dataset è stato generato sinteticamente da Mattia (“Mattimax”) per l’azienda M.INC.Serve per lo sviluppo e la valutazione di modelli in grado di ragionare in italiano e fornire risposte strutturate con tracciamento del ragionamento.I dati non sono garantiti accurati e sono destinati esclusivamente a scopi di ricerca e sperimentazione.
Fonte
Profilo autore: https://huggingface.co/Mattimax
Organizzazione: https://huggingface.co/MINC01… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DAC-Reasoning-ITA.
