datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dacomp-da-zh
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
Paper | Project Page | Code
This repository contains DAComp, a benchmark of 210 tasks that mirrors complex real-world enterprise data intelligence workflows. It includes:
Data Engineering (DE) tasks: Require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements.
Data Analysis (DA)… See the full description on the dataset page: https://huggingface.co/datasets/DAComp/dacomp-da-zh.dacomp-da-zh
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
Paper | Project Page | Code
This repository contains DAComp, a benchmark of 210 tasks that mirrors complex real-world enterprise data intelligence workflows. It includes:
Data Engineering (DE) tasks: Require repository-level engineering on industrial schemas, including designing and building multi-stage SQL pipelines from scratch and evolving existing systems under evolving requirements.
Data Analysis (DA)… See the full description on the dataset page: https://huggingface.co/datasets/jjjsadhfgj/dacomp-da-zh.dac6-instruct
DAC6 instruct (11-12-2023)
“DAC 6” refers to European Council Directive (EU) 2018/822 of May 25, 2018 relating to the automatic and mandatory exchange of information on cross-border arrangements requiring declaration. It aims to strengthen cooperation between tax administrations in EU countries on potentially aggressive tax planning arrangements.
This project focuses on fine-tuning pre-trained language models to create efficient and accurate models for tax practice.
Fine-tuning is… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/dac6-instruct.DACMini_Refined
Dataset di ricerca
DACMini_Refined è un dataset creato a scopo di ricerca e sviluppo per migliorare le capacità del modello compatto DACMini-IT, un modello linguistico italiano da 109 milioni di parametri.
L’obiettivo del dataset è incrementare la qualità delle risposte del modello di base attraverso un processo supervisionato multi-stadio, sfruttando modelli di dimensioni maggiori come generatore e validatore.
Metodologia di generazione
Generazione automatica di… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DACMini_Refined.DAC-Think
DAC-Think
Dataset Name: DAC-ThinkCreator: MattimaxOrganization: MINCLicense: MITLanguage: ItalianoNumber of rows: 24,505
Overview
DAC-Think è un dataset di ragionamento esclusivamente in lingua italiana, progettato per task di generazione di testo e conversational AI. Ogni esempio contiene un prompt e una risposta strutturata, con tag <think> che evidenziano la parte di ragionamento del modello, seguita dalla risposta finale.
Il dataset è organizzato in questo… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DAC-Think.DAC-Reasoning-ITA
Descrizione del dataset
Questo dataset è stato generato sinteticamente da Mattia (“Mattimax”) per l’azienda M.INC.Serve per lo sviluppo e la valutazione di modelli in grado di ragionare in italiano e fornire risposte strutturate con tracciamento del ragionamento.I dati non sono garantiti accurati e sono destinati esclusivamente a scopi di ricerca e sperimentazione.
Fonte
Profilo autore: https://huggingface.co/Mattimax
Organizzazione: https://huggingface.co/MINC01… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DAC-Reasoning-ITA.
