CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ArtificialAnalysis /big_bench_audio Artificial Analysis Big Bench Audio Dataset Summary Big Bench Audio is an audio version of a subset of Big Bench Hard questions. The dataset can be used for evaluating the reasoning capabilities of models that support audio input. The dataset includes 1000 audio recordings for all questions from the following Big Bench Hard categories. Descriptions are taken from Suzgun et al. (2022): Formal Fallacies Syllogisms Negation (Formal Fallacies) - 250 questions Given a context… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/big_bench_audio.audioaudio-to-audio1K<n<10K38 likes15k downloads2y agoHugging Face02ArtificialAnalysis /AA-LCR Artificial Analysis Long Context Reasoning (AA-LCR) Dataset AA-LCR includes 100 hard text-based questions that require reasoning across multiple real-world documents, with each document set averaging ~100k input tokens. Questions are designed such that answers cannot be directly retrieved from documents and must instead be reasoned from multiple information sources. New in Version 1.1 (September 2026) Sixteen corrected answer keys. Each one was re-verified… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-LCR.tabularn<1K38 likes8.5k downloads20d agoHugging Face03ArtificialAnalysis /AA-Omniscience-Public Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies. Leaderboard and detailed results Paper Introduction We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Omniscience-Public.documentquestion-answeringn<1K50 likes7.9k downloads1mo agoHugging Face04ArtificialAnalysis /Earnings22-Cleaned-AA Earnings22-Cleaned-AA Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.audioautomatic-speech-recognitionn<1K6 likes6.8k downloads7mo agoHugging Face05artificialguybr /veo3-video-prompts Veo 3 Video Generation Dataset English | Português do Brasil English Summary A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant. Videos: 5,811 Input images: 1,354 Configurations: 6 Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.imagetext-to-video1K<n<10K0 likes5.3k downloads1mo agoHugging Face06ArtificialAnalysis /ITBench-AA ITBench-AA Artificial Analysis' release of the public scenarios from IBM's ITBench benchmark, used for the ITBench-AA leaderboard. This repo currently contains the SRE subset (sre config). Each row is a Kubernetes incident scenario with its expected contributing-factor entities. An agent under evaluation is given access to an offline snapshot of the affected cluster (alerts, events, traces, topology) and must identify the entity (Deployment, Pod, ConfigMap, etc.) responsible for… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/ITBench-AA.textquestion-answeringn<1K47 likes2.5k downloads4mo agoHugging Face07ArtificialAnalysis /AA-Briefcase-Lite AA-Briefcase-Lite The public example scenario for AA-Briefcase, Artificial Analysis' frontier agentic evaluation of realistic, long-horizon knowledge work. Leaderboard and detailed results Launch article AA-Briefcase extends frontier model benchmarking beyond coding and short-form reasoning to the professional deliverables knowledge workers produce day to day. It consists of four private scenarios in which agents complete realistic professional workflows across data science… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Briefcase-Lite.documentothern<1K11 likes2.4k downloads3mo agoHugging Face08vidore /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes1.8k downloads1y agoHugging Face09yalesunxiatao /artificial-foveated-perception Artificial Foveated Perception (AFP) Dataset Training data for Artificial Foveated Perception (AFP), a task-conditioned mask predictor for robotic foundation models (paper, code, labeling tool). The dataset contains 786 robot-manipulation episodes from real-world and simulated manipulation data. Every frame is paired with a continuous task-relevance matte: an alpha map in [0, 1] that is 1 on the task-relevant objects and the robot end-effector, 0 on the background, and graded in… See the full description on the dataset page: https://huggingface.co/datasets/yalesunxiatao/artificial-foveated-perception.image-segmentation0 likes1.6k downloads17d agoHugging Face10ArtificialAnalysis /VoxPopuli-Cleaned-AA VoxPopuli-Cleaned-AA Quick links: AA Speech to Text Leaderboard | AA-WER v2.0 article VoxPopuli-Cleaned-AA is a cleaned subset of the English VoxPopuli test data from esb/datasets, a speech dataset derived from European Parliament recordings. This cleaned subset is the VoxPopuli portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation of Speech to Text (STT) models. This dataset is part of AA-WER… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/VoxPopuli-Cleaned-AA.audioautomatic-speech-recognitionn<1K7 likes1.2k downloads7mo agoHugging Face11mteb /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes1k downloads8mo agoHugging Face12ArtificialAnalysis /hf-assetsaudion<1K1 likes703 downloads2y agoHugging Face13ArtificialAnalysis /Earnings22-Cleaned-AA-chunked Earnings22-Cleaned-AA-chunked Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation. The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.audioautomatic-speech-recognitionn<1K1 likes608 downloads3mo agoHugging Face14vidore /syntheticDocQA_artificial_intelligence_test Dataset Description This dataset is part of a topic-specific retrieval benchmark spanning multiple domains, which evaluates retrieval in more realistic industrial applications. It includes documents about the Artificial Intelligence. Data Collection Thanks to a crawler (see below), we collected 1,000 PDFs from the Internet with the query ('artificial intelligence'). From these documents, we randomly sampled 1000 pages. We associated these with 100 questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/vidore/syntheticDocQA_artificial_intelligence_test.imagedocument-question-answering1K<n<10K2 likes493 downloads1y agoHugging Face15department-of-artificial-intelligence /solar-MID-descriptors0 likes446 downloads13d agoHugging Face16Coder-Dragon /indian-traditional-artificial-jewellery Traditional and Handmade Indian Jewellery Dataset This dataset contains a comprehensive collection of traditional and handmade Indian jewelry, sourced from various e-commerce platforms and manufacturer websites. It provides a rich set of attributes for each jewelry piece, making it a valuable resource for various data analysis, machine learning, and market research tasks. Dataset Overview This dataset is designed to provide detailed information about Indian jewelry… See the full description on the dataset page: https://huggingface.co/datasets/Coder-Dragon/indian-traditional-artificial-jewellery.imageimage-classification1K<n<10K2 likes353 downloads1y agoHugging Face17artificial-memory-lab /imageability-geometry-data0 likes308 downloads1y agoHugging Face18jinaai /docqa_artificial_intelligence_beirThis is a copy of https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence_beir.image1K<n<10K0 likes264 downloads1y agoHugging Face19manidhardevu /indian-traditional-artificial-jewellery Traditional and Handmade Indian Jewellery Dataset This dataset contains a comprehensive collection of traditional and handmade Indian jewelry, sourced from various e-commerce platforms and manufacturer websites. It provides a rich set of attributes for each jewelry piece, making it a valuable resource for various data analysis, machine learning, and market research tasks. Dataset Overview This dataset is designed to provide detailed information about Indian… See the full description on the dataset page: https://huggingface.co/datasets/manidhardevu/indian-traditional-artificial-jewellery.imageimage-classification1K<n<10K0 likes215 downloads1mo agoHugging Face20BAAI /IndustryInstruction_Artificial-Intelligence IndustryInstruction: Artificial Intelligence This repository contains the IndustryInstruction: Artificial Intelligence domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Artificial-Intelligence.tabularquestion-answering100K<n<1M2 likes212 downloads1mo agoHugging Face21department-of-artificial-intelligence /solar-MFI-images0 likes182 downloads2d agoHugging Face22jpamarlphi-byte /Human-Artificial-HAUFlicense: All rights reserved Copyright © 2025/6 JP A-Marl Release Title: HAUF - Human-Artificial Unified Framework by JP A-Marl is now available to Policymakers, Future Civilizational Designers, Future Institutional Governance Designers, AI Researchers, Data Scientists, AI and AGI Developers Version v2.0 27July2026 Version v2.0 published on the 27Jul2025 https://zenodo.org/records/21630286 DOI 10.5281/zenodo.21630286 Version v1.0 originally… See the full description on the dataset page: https://huggingface.co/datasets/jpamarlphi-byte/Human-Artificial-HAUF.imagen<1K0 likes176 downloads7d agoHugging Face23BAAI /IndustryCorpus2_artificial_intelligence_machine_learning IndustryCorpus2: Artificial Intelligence This repository contains the IndustryCorpus2: Artificial Intelligence domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_artificial_intelligence_machine_learning.2 likes144 downloads1mo agoHugging Face24Besedo /artificial_weapon Dataset Card for [Dataset Name] Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/Besedo/artificial_weapon.imageimage-classification1K<n<10K1 likes119 downloads4y agoHugging Face25artificial-memory-lab /ann-datasets0 likes98 downloads1y agoHugging Face26artificial-memory-lab /text-collections-embeddings0 likes95 downloads2y agoHugging Face27artificialheartai /dataset-for-training-onlygated 🛑 DATASET RESTRITO - ACESSO SOB PERMISSÃO EXCLUSIVA Este repositório contém dados estruturados privados destinados estritamente a fins de treinamento controlado. 🔐 Termos da Licença e Regras de Acesso (Gated) Direitos Reservados (License: Other): Este dataset não possui uma licença pública ou de código aberto. É proibida qualquer cópia, distribuição, comercialização ou uso por terceiros sem a autorização expressa do proprietário (artificialheartai). Aprovação… See the full description on the dataset page: https://huggingface.co/datasets/artificialheartai/dataset-for-training-only.text-to-image1 likes94 downloads12d agoHugging Face28akari000 /artificial_variationsetsThis dataset contains artificial variation sets generated by GPT-4o-mini.Variation sets are sets of (mostly consecutive) utterances that convey a similar intent with slight variations in word choice and structure (Küntay and Slobin, 1996). They are a characteristic feature of Child-Directed Speech. All artificial variation sets are available in artificial-variationsets.txt . Described in the following paper: https://arxiv.org/abs/2411.09587 In this paper, artificial variation sets were… See the full description on the dataset page: https://huggingface.co/datasets/akari000/artificial_variationsets.text1M<n<10M1 likes92 downloads2y agoHugging Face29ArtificialZeng /leetcode_code_generationtext1K<n<10K4 likes78 downloads2y agoHugging Face30Artificial-Unintelligence /flachwitze-de Artificial-Unintelligence: Deutsche Flachwitze Ein kuratierter NLP-Datensatz mit deutschen Flachwitzen, Kalauern und Dad Jokes. Jeder Witz ist strukturiert in Setup, Punchline, Wortspiel-Erklärung und einen Cringe-Score (1-5). 📊 Statistik Gesamtanzahl: 7749 Witze Durchschnittlicher Cringe-Score: 3.83 / 5.0 Kategorien: Wortwitz: 7749 📋 Schema Spalte Typ Beschreibung id string Eindeutiger Identifier (AU-DE-XXXXXX) setup string… See the full description on the dataset page: https://huggingface.co/datasets/Artificial-Unintelligence/flachwitze-de.texttext-generation1K<n<10K0 likes75 downloads8d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.