CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01artefactory /ledger-long-context-KPI-QA LEDGER — Long-Context KPI Question Answering & Page Retrieval This dataset is part of the LEDGER (Long-context Evaluation of Documents for Grounded Extraction and Retrieval) benchmark. It supports two of the three LEDGER tasks: Page-level KPI retrieval — given a natural-language question about a financial KPI and the corresponding annual report, retrieve the relevant page(s). Each row includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.tabularquestion-answering100K<n<1M14 likes3k downloads1mo agoHugging Face02csoai /gspc-art5 GSPC — art5 safeguard bank (Art5Bench) Council of AI measurement bank. Measurement, not certification. Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026. Live measurement. This bank stands behind the art5-safeguard row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=art5-safeguard (family, kind, status and n are on that row… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-art5.tabularquestion-answeringn<1K0 likes938 downloads2d agoHugging Face03abhilash88 /aim-technical-articles Analytics India Magazine Technical Articles Dataset 🚀 Dataset Description This comprehensive dataset contains 25,685 high-quality technical articles from Analytics India Magazine, one of India's leading publications covering artificial intelligence, machine learning, data science, and emerging technologies. ✨ Dataset Highlights 📚 Comprehensive Coverage: Latest AI models, frameworks, and tools 🔬 Technical Depth: Extracted keywords and complexity scoring 🏭… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/aim-technical-articles.tabulartext-classification10K<n<100K2 likes226 downloads1y agoHugging Face04BAAI /IndustryInstruction_Artificial-Intelligence IndustryInstruction: Artificial Intelligence This repository contains the IndustryInstruction: Artificial Intelligence domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Artificial-Intelligence.tabularquestion-answering100K<n<1M2 likes212 downloads1mo agoHugging Face05artefactory /Argimi-Legal-French-Jurisprudence The ArGiMi French Jurisprudence Dataset This dataset contains a comprehensive collection of French case law, sourced from the official archives of French jurisprudence. It is divided into three distinct subdivisions: Constitutional ("constit"), Administrative ("cetat"), and Judiciary ("juri"). This dataset was created for the ArGiMi project, an open-source initiative dedicated to promoting open data and knowledge sharing. The project is a collaborative effort between Giskard… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Legal-French-Jurisprudence.tabularquestion-answering100K<n<1M10 likes187 downloads1y agoHugging Face06dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas Dataset summary It is an end-to-end evaluation dataset (using SAS metric) for Justicio. Domain: Legal, Law, Spanish Constitution Language: Spanish SAS summary The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth. Justicio summary Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas.tabularquestion-answeringn<1K1 likes101 downloads2y agoHugging Face07TonicAI /vectrix-art-e Vectrix ART-E: Synthetic Email Agent Benchmark A fully synthetic email corpus and task dataset for training and evaluating email search agents, built as a drop-in replacement for the Enron corpus used in OpenPipe's ART-E benchmark. Key Result A Qwen3.5-35B-A3B fine-tuned via GRPO on this synthetic dataset beats o3 on real Enron emails (86% vs 85%) — despite never seeing a single real email during training. Dataset Contents The dataset is available in two… See the full description on the dataset page: https://huggingface.co/datasets/TonicAI/vectrix-art-e.tabularquestion-answering1K<n<10K1 likes101 downloads6mo agoHugging Face08dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas Dataset summary It is an end-to-end evaluation dataset (using SAS metric) for Justicio. Domain: Legal, Law, Spanish Constitution Language: Spanish SAS summary The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth. Justicio summary Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas.tabularquestion-answeringn<1K0 likes91 downloads2y agoHugging Face09dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas Dataset summary It is an end-to-end evaluation dataset (using SAS metric) for Justicio. Domain: Legal, Law, Spanish Constitution Language: Spanish SAS summary The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth. Justicio summary Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas.tabularquestion-answeringn<1K0 likes76 downloads2y agoHugging Face10qurancn /islamic-articles-corpus ☪ Islamic Articles Corpus - English RAG Dataset Dataset Description Islamic Articles Corpus is a curated English-language RAG corpus containing 33 articles covering Muslim travel guides, mosque visits, halal food, prayer room directories, and Islamic community documentation. Every article preserves complete full-text content with all 608 embedded image references. Content focuses heavily on Singapore, Iran, Japan, Oman, and Qatar mosque and travel documentation.… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/islamic-articles-corpus.tabulartext-generationn<1K0 likes70 downloads3mo agoHugging Face11DariusTheGeek /mhqa-itu-artifacts MHQA · ITU · Zindi Challenge — Artifacts DariusTheGeek/mhqa-itu-artifacts · the data + precomputed features that let the code repo reproduce submission sub_v40 (public LB 0.728509) for the ITU Multilingual Health QA in Low-Resource African Languages challenge. Code (which pulls this at runtime) lives on GitHub; trained weights are in the model repo DariusTheGeek/mhqa-itu-adapters. This is a reproducibility artifact bundle, not a raw dataset. It holds derived features and the… See the full description on the dataset page: https://huggingface.co/datasets/DariusTheGeek/mhqa-itu-artifacts.tabularquestion-answering1M<n<10M0 likes65 downloads3mo agoHugging Face12ArthurSrz /comparag-tool-votes CompaRAG — Tool Votes Dataset CompaRAG is a blind comparison platform for MCP (Model Context Protocol) tools, built by The Borges Graph as a spinoff of comparIA, a French government initiative (Ministère de la Culture / Beta.gouv.fr). What is this dataset? This dataset contains human preference votes collected on the CompaRAG platform. Users submit a task and a goal, two MCP tools respond anonymously, and the user votes for the best result — without knowing which… See the full description on the dataset page: https://huggingface.co/datasets/ArthurSrz/comparag-tool-votes.tabulartext-generationn<1K1 likes64 downloads2d agoHugging Face13Astound /Art-GenEvalGPT Dataset Card Dataset Details Dataset Description The dataset includes synthetic dialogues in the art domain that can be used for training a chatbot to discuss artworks within a museum setting. Leveraging Large Language Models (LLMs), particularly ChatGPT, the dataset comprises over 13,000 dialogues generated using prompt-engineering techniques. The dialogues cover a wide range of user and chatbot behaviors, including expert guidance, tutoring, and handling… See the full description on the dataset page: https://huggingface.co/datasets/Astound/Art-GenEvalGPT.tabularquestion-answering10K<n<100K3 likes42 downloads2y agoHugging Face14louisbrulenaudet /code-artisanat Code de l'artisanat, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-artisanat.tabulartext-generationn<1K0 likes40 downloads1y agoHugging Face15crawlfeeds /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes38 downloads6mo agoHugging Face16hybridfree /phoronix-articles Phoronix Articles Dataset: The Archive of Open-Source Computing Journalism The definitive dataset of Phoronix - your gateway to years of open-source hardware/software evolution, performance analysis, and Linux ecosystem journalism. 🚀 What's Inside? This dataset contains the complete archive of Phoronix articles - from bleeding-edge hardware launches to deep-dive Linux kernel analysis. Perfect for researchers, developers, and AI enthusiasts who need high-quality technical… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/phoronix-articles.tabulartext-generation10K<n<100K0 likes35 downloads8mo agoHugging Face17UWV /wim-schema-org-wiki-articles Dutch Wikipedia Aligned Articles aligned with Schema.org Classes Dataset Version: 1.0 (2025-06-04)Point of Contact: UWV Netherlands (UWV organization on Hugging Face)License: CC BY-SA 4.0Dataset: UWV/wim_schema_org_wiki_articles Dataset Description This dataset provides alignments between Schema.org classes and relevant Dutch Wikipedia articles. Each Schema.org class from a processed subset is linked to up to 20 distinct Wikipedia articles, including their full text, a… See the full description on the dataset page: https://huggingface.co/datasets/UWV/wim-schema-org-wiki-articles.tabulartext-classification10K<n<100K1 likes33 downloads1y agoHugging Face18Adel-Elwan /Artificial-intelligence-dataset-for-IR-systems Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards information-retrieval semantic-search Languages English Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Adel-Elwan/Artificial-intelligence-dataset-for-IR-systems.tabularquestion-answering100K<n<1M0 likes25 downloads3y agoHugging Face19prithivMLmods /Content-Articles Content-Articles Dataset Overview The Content-Articles dataset is a collection of academic articles and research papers across various subjects, including Computer Science, Physics, and Mathematics. This dataset is designed to facilitate research and analysis in these fields by providing structured data on article titles, abstracts, and subject classifications. Dataset Details Modalities Tabular: The dataset is structured in a tabular format. Text:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Content-Articles.tabulartext-generation10K<n<100K3 likes24 downloads2y agoHugging Face20abhilash88 /techcrunch-articles TechCrunch News Articles Dataset 📊 Dataset Overview This dataset contains 10,265 high-quality news articles scraped from TechCrunch, one of the leading technology news websites. The dataset includes comprehensive article content, metadata, and quality assessments suitable for various NLP tasks including text classification, sentiment analysis, summarization, and content generation. 🎯 Key Features 10,265 articles with full text content High-quality filtering… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/techcrunch-articles.tabulartext-classification10K<n<100K1 likes23 downloads1y agoHugging Face21Adilbai /kz-rus-articles-comprehensive 🇰🇿🇷🇺 Kazakh-Russian Articles Comprehensive Dataset A high-quality bilingual corpus for cross-lingual NLP research 📋 Dataset Overview The Kazakh-Russian Articles Comprehensive Dataset is a meticulously curated bilingual corpus designed to advance natural language processing research for Kazakh and Russian languages. This dataset addresses the critical need for high-quality parallel and comparable text resources in Central Asian language pairs, particularly… See the full description on the dataset page: https://huggingface.co/datasets/Adilbai/kz-rus-articles-comprehensive.tabulartranslationn<1K1 likes18 downloads1y agoHugging Face22david-sprague /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes17 downloads4mo agoHugging Face23CentificAIResearch /Med-ART_Clinical_Agent_EHR_Datasetgated ART — Action-based Reasoning Tasks (Subset) 120-task stratified sample from the ART benchmark introduced in: ART: Action-based Reasoning Task Benchmarking for Medical AI Agents Ananya Mantravadi, Shivali Dalmia, Abhishek Mukherji arXiv:2601.08988 ART is a programmatically generated clinical decision benchmark built on real FHIR patient data. It targets three dominant error categories in medical AI reasoning — retrieval failures, aggregation errors, and conditional logic… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/Med-ART_Clinical_Agent_EHR_Dataset.tabulartext-generationn<1K2 likes15 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.