CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DAComp /dacomp-da-zh DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle Paper | Project Page | Code This repository contains DAComp, a benchmark of 210 tasks that mirrors complex real-world enterprise data intelligence workflows. It includes: Data Engineering (DE) tasks: Require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements. Data Analysis (DA)… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da-zh.texttext-generationn<1K0 likes666 downloads10mo agoHugging Face02DAComp /dacomp-de DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle This repository contains the DAComp benchmark, presented in the paper DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle. Project page: https://da-comp.github.io/ Code: https://github.com/anonymous/DAComp ✍️ Citation If you find our work helpful, please cite as @misc{lei2025dacompbenchmarkingdataagents, title={DAComp: Benchmarking Data Agents across the Full… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-de.text-generation4 likes101 downloads10mo agoHugging Face03botcoinmoney /dacr-lt-training DACR Recurrent-Depth Training Data A large enriched reasoning corpus derived from the BOTCOIN/DACR data pipeline and adjusted for preliminary recurrent-depth natural-language experiments. This dataset is not intended to be treated as a single fixed training split. It is better understood as a reusable source corpus containing several export categories that can be pruned, reshaped, and filtered depending on the training objective. What Is Included… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/dacr-lt-training.text-generation100K<n<1M0 likes70 downloads5mo agoHugging Face04Luo2003 /DA-Codegated [EMNLP2024] DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models DA-Code is a comprehensive evaluation dataset designed to assess the data analysis and code generation capabilities of LLM in agent-based data science tasks. Our papers and experiment reports have been published on Arxiv. Dataset Overview 500 complex real-world data analysis tasks across Data Wrangling (DW), Machine Learning (ML), and Exploratory Data Analysis (EDA). Tasks cover… See the full description on the dataset page: https://huggingface.co/datasets/Luo2003/DA-Code.texttable-question-answeringn<1K4 likes38 downloads2y agoHugging Face05jjjsadhfgj /dacomp-da-zh DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle Paper | Project Page | Code This repository contains DAComp, a benchmark of 210 tasks that mirrors complex real-world enterprise data intelligence workflows. It includes: Data Engineering (DE) tasks: Require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements. Data Analysis (DA)… See the full description on the dataset page: https://huggingface.co/datasets/jjjsadhfgj/dacomp-da-zh.texttext-generationn<1K0 likes37 downloads9mo agoHugging Face06ShantanuT01 /DACTYL-Pretraining DACTYL Pretraining Corpus This corpus contains human texts that have passed a quality check from the textdescriptives library. These texts have been used to further train various Llama 3.2 1B Instruct models by domain and split. For a domain/split combination, you can use the corresponding model: ShantanuT01/fine-tuned-Llama-3.2-1B-Instruct-apollo-mini-{domain}-{split} Citation @misc{thorat2025dactyldiverseadversarialcorpus, title={DACTYL: Diverse Adversarial… See the full description on the dataset page: https://huggingface.co/datasets/ShantanuT01/DACTYL-Pretraining.texttext-generation100K<n<1M0 likes34 downloads1y agoHugging Face07KaanGoker /dactylic-hexameter-latin-poetry-corpus Dactylic Hexameter Latin Poetry Corpus This repository contains a curated and processed corpus of Classical Latin poetry written in dactylic hexameter. It serves as the raw training data ("Dataset V3") for the Master's Thesis titled "A Hybrid Post Hoc Feedback Framework for Latin Dactylic Hexameter" submitted to KU Leuven (2025). Dataset Description This corpus was constructed to fine-tune Large Language Models (LLMs) for the generation of metrically valid Latin poetry.… See the full description on the dataset page: https://huggingface.co/datasets/KaanGoker/dactylic-hexameter-latin-poetry-corpus.texttext-generation10K<n<100K0 likes27 downloads9mo agoHugging Face08louisbrulenaudet /dac6-instruct DAC6 instruct (11-12-2023) “DAC 6” refers to European Council Directive (EU) 2018/822 of May 25, 2018 relating to the automatic and mandatory exchange of information on cross-border arrangements requiring declaration. It aims to strengthen cooperation between tax administrations in EU countries on potentially aggressive tax planning arrangements. This project focuses on fine-tuning pre-trained language models to create efficient and accurate models for tax practice. Fine-tuning is… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/dac6-instruct.texttext-generationn<1K0 likes23 downloads2y agoHugging Face09Mattimax /DACMini_Refined Dataset di ricerca DACMini_Refined è un dataset creato a scopo di ricerca e sviluppo per migliorare le capacità del modello compatto DACMini-IT, un modello linguistico italiano da 109 milioni di parametri. L’obiettivo del dataset è incrementare la qualità delle risposte del modello di base attraverso un processo supervisionato multi-stadio, sfruttando modelli di dimensioni maggiori come generatore e validatore. Metodologia di generazione Generazione automatica di… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DACMini_Refined.texttext-generation10K<n<100K0 likes20 downloads11mo agoHugging Face10Mattimax /DAC-Thinkgated DAC-Think Dataset Name: DAC-ThinkCreator: MattimaxOrganization: MINCLicense: MITLanguage: ItalianoNumber of rows: 24,505 Overview DAC-Think è un dataset di ragionamento esclusivamente in lingua italiana, progettato per task di generazione di testo e conversational AI. Ogni esempio contiene un prompt e una risposta strutturata, con tag <think> che evidenziano la parte di ragionamento del modello, seguita dalla risposta finale. Il dataset è organizzato in questo… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DAC-Think.texttext-generation10K<n<100K0 likes8 downloads9mo agoHugging Face11Mattimax /DAC-Reasoning-ITA Descrizione del dataset Questo dataset è stato generato sinteticamente da Mattia (“Mattimax”) per l’azienda M.INC.Serve per lo sviluppo e la valutazione di modelli in grado di ragionare in italiano e fornire risposte strutturate con tracciamento del ragionamento.I dati non sono garantiti accurati e sono destinati esclusivamente a scopi di ricerca e sperimentazione. Fonte Profilo autore: https://huggingface.co/Mattimax Organizzazione: https://huggingface.co/MINC01… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DAC-Reasoning-ITA.texttext-generation10K<n<100K0 likes5 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.