CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dacorvo /funes-handoff-recall-benchmark handover-vs-recall A long investigation bloats an agent session until each new turn costs more to carry the context than to do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task, on tasks that genuinely require the prior investigation: arm channel A branch-only switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.tabularn<1K0 likes3.1k downloads20d agoHugging Face02DAComp /dacomp-da-zh DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle Paper | Project Page | Code This repository contains DAComp, a benchmark of 210 tasks that mirrors complex real-world enterprise data intelligence workflows. It includes: Data Engineering (DE) tasks: Require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements. Data Analysis (DA)… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da-zh.texttext-generationn<1K0 likes664 downloads10mo agoHugging Face03dacthai2807 /ViMed-PET-part1 Dataset description for three years: 2017, 2018, 2019 This dataset contains data from three years (2017, 2018, 2019). Each year has several month folders, which are named as THANG {month}. Each year folder is compressed into zip files (chunks), each with an average size of approximately 2.5 GB. Please unzip the .zip files to fully extract all data folders. Folder structure after extraction Each folder named THANG {month} of a year is divided into 3 subfolders… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2807/ViMed-PET-part1.text1K<n<10K0 likes651 downloads1y agoHugging Face04dacorvo /transformers-coding-session-pi-traces dacorvo/transformers-coding-session-pi-traces pi coding-agent session traces produced by agentcap runs. Each run contributes one folder under data/<run_id>/; inside, one file per session in pi's native export format. The on-the-wire HTTP captures for these same runs live in dacorvo/transformers-coding-session-captures. Both belong to the transformers-coding-session Collection — join on run_id to align captures with traces. tabularn<1K0 likes605 downloads4mo agoHugging Face05dacorvo /hf-hub-session-pi-traces dacorvo/hf-hub-session-pi-traces pi coding-agent session traces produced by agentcap runs. Each run contributes one folder under data/<run_id>/; inside, one file per session in pi's native export format. The on-the-wire HTTP captures for these same runs live in dacorvo/hf-hub-session-captures. Both belong to the hf-hub-session Collection — join on run_id to align captures with traces. tabularn<1K0 likes572 downloads4mo agoHugging Face06dacthai2k /ViMed-PET-part3 Dataset description for year 2023 This dataset contains data from three months: October, November, and December, stored in the following folders respectively: THANG 10 THANG 11 THANG 12 The data is compressed into zip files (chunks), each with an average size of approximately 2.5 GB. Please unzip the .zip files to fully extract the data folders. Folder structure after extraction Each folder named THANG {month} is divided into 3 subfolders, corresponding to 2… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2k/ViMed-PET-part3.text1K<n<10K0 likes405 downloads1y agoHugging Face07DAComp /dacomp-da DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle ✍️ Citation If you find our work helpful, please cite as @misc{lei2025dacompbenchmarkingdataagents, title={DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle}, author={Fangyu Lei and Jinxiang Meng and Yiming Huang and Junjie Zhao and Yitong Zhang and Jianwen Luo and Xin Zou and Ruiyi Yang and Wenbo Shi and Yan Gao and Shizhu He and Zuo Wang and Qian Liu and… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da.textn<1K6 likes355 downloads10mo agoHugging Face08dacheah /space-law-corpus Space Law Corpus A neutral, provenance-first, machine-readable record of international and national space law. Every record carries its official source, retrieval date, citation, language, an authoritative-status flag, and a SHA-256 content hash; texts are verified against official sources. Source of truth / build history: https://github.com/dacheah/space-law-corpus Archived & citable: concept DOI 10.5281/zenodo.21185483 (resolves to the latest Zenodo-archived GitHub release)… See the full description on the dataset page: https://huggingface.co/datasets/dacheah/space-law-corpus.texttext-retrievaln<1K0 likes76 downloads2mo agoHugging Face09dacorvo /funes-recall-session-pi-traces dacorvo/funes-recall-session-pi-traces pi coding-agent session traces produced by agentcap runs. Each run contributes one folder under data/<run_id>/; inside, one file per session in pi's native export format. The on-the-wire HTTP captures for these same runs live in dacorvo/funes-recall-session-captures. Both belong to the funes-recall-session Collection — join on run_id to align captures with traces. tabularn<1K0 likes59 downloads3mo agoHugging Face10dacheah /bbnj-high-seas-treaty-corpus BBNJ / High Seas Treaty Corpus A neutral, provenance-first, machine-readable record of the 2023 BBNJ Agreement (the "High Seas Treaty", in force 17 January 2026) in all six authentic UN languages, and its implementing framework. Every record carries its official source, retrieval date, citation, authentic language, an authoritative-status flag, a SHA-256 content hash, and an honest per-language fidelity flag (extracted_verified → extracted_unverified → ocr_unverified). Source… See the full description on the dataset page: https://huggingface.co/datasets/dacheah/bbnj-high-seas-treaty-corpus.texttext-retrievaln<1K0 likes48 downloads2mo agoHugging Face11dacheah /deep-seabed-mining-law-corpus Deep Seabed Mining Law Corpus A neutral, provenance-first, machine-readable record of the law governing mineral resources of "the Area" (the seabed beyond national jurisdiction): the international ISA/UNCLOS regime and the US non-UNCLOS parallel track. Every record carries its official source, retrieval date, citation, language, an authoritative-status flag, and a SHA-256 content hash; texts are verified against official sources. Source of truth / build history:… See the full description on the dataset page: https://huggingface.co/datasets/dacheah/deep-seabed-mining-law-corpus.texttext-retrieval1K<n<10K0 likes45 downloads2mo agoHugging Face12jjjsadhfgj /dacomp-da-zh DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle Paper | Project Page | Code This repository contains DAComp, a benchmark of 210 tasks that mirrors complex real-world enterprise data intelligence workflows. It includes: Data Engineering (DE) tasks: Require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements. Data Analysis (DA)… See the full description on the dataset page: https://huggingface.co/datasets/jjjsadhfgj/dacomp-da-zh.texttext-generationn<1K0 likes38 downloads9mo agoHugging Face13dacorvo /funes-recall-session-hermes-traces dacorvo/funes-recall-session-hermes-traces hermes coding-agent session traces produced by agentcap runs. Each run contributes one folder under data/<run_id>/; inside, one file per session in hermes's native export format. The on-the-wire HTTP captures for these same runs live in dacorvo/funes-recall-session-captures. Both belong to the funes-recall-session Collection — join on run_id to align captures with traces. tabularn<1K0 likes31 downloads3mo agoHugging Face14louisbrulenaudet /dac6-instruct DAC6 instruct (11-12-2023) “DAC 6” refers to European Council Directive (EU) 2018/822 of May 25, 2018 relating to the automatic and mandatory exchange of information on cross-border arrangements requiring declaration. It aims to strengthen cooperation between tax administrations in EU countries on potentially aggressive tax planning arrangements. This project focuses on fine-tuning pre-trained language models to create efficient and accurate models for tax practice. Fine-tuning is… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/dac6-instruct.texttext-generationn<1K0 likes25 downloads2y agoHugging Face15Mattimax /DACMini_Refined Dataset di ricerca DACMini_Refined è un dataset creato a scopo di ricerca e sviluppo per migliorare le capacità del modello compatto DACMini-IT, un modello linguistico italiano da 109 milioni di parametri. L’obiettivo del dataset è incrementare la qualità delle risposte del modello di base attraverso un processo supervisionato multi-stadio, sfruttando modelli di dimensioni maggiori come generatore e validatore. Metodologia di generazione Generazione automatica di… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DACMini_Refined.texttext-generation10K<n<100K0 likes19 downloads11mo agoHugging Face16daceteate /SD4tabularn<1K0 likes9 downloads11mo agoHugging Face17open-llm-leaderboard /mindw96__DeepSeek-llama3.3-Bllossom-8B-DACON-LLM3-detailsgated Dataset Card for Evaluation run of mindw96/DeepSeek-llama3.3-Bllossom-8B-DACON-LLM3 Dataset automatically created during the evaluation run of model mindw96/DeepSeek-llama3.3-Bllossom-8B-DACON-LLM3 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mindw96__DeepSeek-llama3.3-Bllossom-8B-DACON-LLM3-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face18dachengzisks /yingjiyuantextn<1K0 likes7 downloads2y agoHugging Face19Mattimax /DAC-Thinkgated DAC-Think Dataset Name: DAC-ThinkCreator: MattimaxOrganization: MINCLicense: MITLanguage: ItalianoNumber of rows: 24,505 Overview DAC-Think è un dataset di ragionamento esclusivamente in lingua italiana, progettato per task di generazione di testo e conversational AI. Ogni esempio contiene un prompt e una risposta strutturata, con tag <think> che evidenziano la parte di ragionamento del modello, seguita dalla risposta finale. Il dataset è organizzato in questo… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DAC-Think.texttext-generation10K<n<100K0 likes7 downloads9mo agoHugging Face20Mattimax /DAC-Reasoning-ITA Descrizione del dataset Questo dataset è stato generato sinteticamente da Mattia (“Mattimax”) per l’azienda M.INC.Serve per lo sviluppo e la valutazione di modelli in grado di ragionare in italiano e fornire risposte strutturate con tracciamento del ragionamento.I dati non sono garantiti accurati e sono destinati esclusivamente a scopi di ricerca e sperimentazione. Fonte Profilo autore: https://huggingface.co/Mattimax Organizzazione: https://huggingface.co/MINC01… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DAC-Reasoning-ITA.texttext-generation10K<n<100K0 likes5 downloads11mo agoHugging Face21bazobehram /dac-judge-v4-datasettextn<1K0 likes1 downloads7mo agoHugging Face22dachengzisks /xiaofang1textn<1K0 likes2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.