CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /arguana ArguAna An MTEB dataset Massive Text Embedding Benchmark ArguAna: Retrieval of the Best Counterargument without Prior Topic Knowledge Task category Retrieval (text-to-text) Domains Social, Web, Written Reference ACL Source datasets: mteb/arguana How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("ArguAna") evaluator = mteb.MTEB([task]) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/arguana.texttext-retrieval10K<n<100K7 likes27k downloads5mo agoHugging Face02mteb /arena-resultsThis dataset contains the saved results from MTEB-Arena tabular1K<n<10K4 likes9k downloads1y agoHugging Face03alexandrainst /m_arc Multilingual ARC Dataset Summary This dataset is a machine translated version of the ARC dataset. The Icelandic (is) part was translated with Miðeind's Greynir model and Norwegian (nb) was translated with DeepL. The rest of the languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository. textquestion-answering10K<n<100K4 likes8k downloads3y agoHugging Face04AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes7.6k downloads1y agoHugging Face05jamescalam /ai-arxiv2-chunkstext100K<n<1M4 likes7.3k downloads3y agoHugging Face06ArtificialAnalysis /Earnings22-Cleaned-AA Earnings22-Cleaned-AA Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.audioautomatic-speech-recognitionn<1K6 likes6.9k downloads7mo agoHugging Face07Manusagents /arvo-cybergym-2000 ARVO CyberGym-format 2000-task dataset This dataset is shaped to be loaded by Harbor's CyberGym adapter. It combines jm-rt/arvo-cybergym-1000 with the second 1000-task small-target ARVO batch built outside the original CyberGym set. text1K<n<10K0 likes5.8k downloads2mo agoHugging Face08artificialguybr /veo3-video-prompts Veo 3 Video Generation Dataset English | Português do Brasil English Summary A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant. Videos: 5,811 Input images: 1,354 Configurations: 6 Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.imagetext-to-video1K<n<10K0 likes5.3k downloads1mo agoHugging Face09kalomaze /alphabetic-arxiv-authors-it1text100K<n<1M0 likes5.2k downloads1y agoHugging Face10armand0e /claude-fable-5-claude-code claude-fable-5 Agent Traces It's worth noting that our team was working with Glint-Research to collect as much fable data as possible. These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data). For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/claude-fable-5-claude-code.tabulartext-generationn<1K388 likes4.8k downloads16d agoHugging Face11soumitsr /article-digeststext10K<n<100K0 likes4.4k downloads2y agoHugging Face12common-pile /github_archive GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.texttext-generation10M<n<100M2 likes4.3k downloads1y agoHugging Face13common-pile /arxiv_papers ArXiv Papers Description ArXiv is an online open-access repository of over 2.4 million scholarly papers covering fields such as computer science, mathematics, physics, quantitative biology, economics, and more. When uploading papers, authors can choose from a variety of licenses. This dataset includes text from all papers uploaded under CC BY, CC BY-SA, and CC0 licenses through a three-step pipeline: first, the latex source files for openly licensed papers were… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_papers.texttext-generation100K<n<1M17 likes4.1k downloads1y agoHugging Face14mteb /arxiv-clustering-s2s ArXivHierarchicalClusteringS2S An MTEB dataset Massive Text Embedding Benchmark Clustering of titles from arxiv. Clustering of 30 sets, either on the main or secondary category Task category t2c Domains Academic, Written Reference https://www.kaggle.com/Cornell-University/arxiv How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["ArXivHierarchicalClusteringS2S"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/arxiv-clustering-s2s.texttext-classificationn<1K1 likes4.1k downloads7mo agoHugging Face15deplana /stream-archive stream-archive Twitch and Kick chat logs from Italian streamers. Streamers Twitch.com (39) aladinottv, alisonrevenge, arkanightlive, bartopanzer, billybella_, dankol83, dariomocciatwitch, davidrubino, diariodelrusso, enkk, federicacasula_, fufflix, grenbaud, gskianto, homyatol, ilgabbrone, ilrossopiubelloditwitch, immortale____, kasumisen, lollolacustre, lucakingm, luiskant690, macchiativincenzo_babbohs, marcomerrino, menestointhailandia… See the full description on the dataset page: https://huggingface.co/datasets/deplana/stream-archive.text10M<n<100M2 likes3.8k downloads5h agoHugging Face16mteb /arxiv-clustering-p2p ArXivHierarchicalClusteringP2P An MTEB dataset Massive Text Embedding Benchmark Clustering of titles+abstract from arxiv. Clustering of 30 sets, either on the main or secondary category Task category t2c Domains Academic, Written Reference https://www.kaggle.com/Cornell-University/arxiv How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/arxiv-clustering-p2p.texttext-classificationn<1K3 likes3.6k downloads7mo agoHugging Face17gfissore /arxiv-abstracts-2021 Dataset Card for arxiv-abstracts-2021 Dataset Summary A dataset of metadata including title and abstract for all arXiv articles up to the end of 2021 (~2 million papers). Possible applications include trend analysis, paper recommender engines, category prediction, knowledge graph construction and semantic search interfaces. In contrast to arxiv_dataset, this dataset doesn't include papers submitted to arXiv after 2021 and it doesn't require any external download.… See the full description on the dataset page: https://huggingface.co/datasets/gfissore/arxiv-abstracts-2021.textsummarization1M<n<10M41 likes3.5k downloads4y agoHugging Face18arterm-sedov /agent-course-final-assignment Agent Course Final Assignment - Unified Dataset Author: Arte(r)m Sedov GitHub: https://github.com/arterm-sedov/ Project link: https://huggingface.co/spaces/arterm-sedov/agent-course-final-assignment Dataset Description This dataset is produced by the GAIA Unit 4 Agent for the Hugging Face Agents Course final assignment as part of an experimental multi-LLM agent system that demonstrates advanced AI agent capabilities. It demonstrates advanced AI agent capabilities for… See the full description on the dataset page: https://huggingface.co/datasets/arterm-sedov/agent-course-final-assignment.tabularn<1K1 likes3.4k downloads9mo agoHugging Face19olegbask /AR-LSATRaw datset: https://github.com/zhongwanjun/AR-LSAT text1K<n<10K3 likes2.6k downloads3y agoHugging Face20ArtificialAnalysis /ITBench-AA ITBench-AA Artificial Analysis' release of the public scenarios from IBM's ITBench benchmark, used for the ITBench-AA leaderboard. This repo currently contains the SRE subset (sre config). Each row is a Kubernetes incident scenario with its expected contributing-factor entities. An agent under evaluation is given access to an offline snapshot of the affected cluster (alerts, events, traces, topology) and must identify the entity (Deployment, Pod, ConfigMap, etc.) responsible for… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/ITBench-AA.textquestion-answeringn<1K47 likes2.6k downloads4mo agoHugging Face21QCRI /ImageEval-ArabicNLP26 ImageEval-ArabicNLP26 👁️ ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026. It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation. The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.audio10K<n<100K4 likes2.5k downloads27d agoHugging Face22adameubanks /filtered_articles_by_year Dataset Card for Filtered Articles by Year Dataset Summary The Filtered Articles by Year dataset contains yearly-segmented web articles from the FineWeb dataset, specifically filtered and processed for temporal language analysis and Word2Vec model training. This dataset spans 21 years (2005-2025) and serves as the foundation for research into semantic change, concept emergence, and language evolution over time. Supported Tasks and Leaderboards This dataset… See the full description on the dataset page: https://huggingface.co/datasets/adameubanks/filtered_articles_by_year.texttext-generation10M<n<100M1 likes2.5k downloads1y agoHugging Face23Nobody05 /arvo-cybergym-2000 ARVO CyberGym-format 2000-task dataset This dataset is shaped to be loaded by Harbor's CyberGym adapter. It combines jm-rt/arvo-cybergym-1000 with the second 1000-task small-target ARVO batch built outside the original CyberGym set. text1K<n<10K0 likes2.4k downloads2mo agoHugging Face24ArkhAngelLifeJiggy /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M0 likes2.4k downloads2mo agoHugging Face25ArtificialAnalysis /AA-Briefcase-Lite AA-Briefcase-Lite The public example scenario for AA-Briefcase, Artificial Analysis' frontier agentic evaluation of realistic, long-horizon knowledge work. Leaderboard and detailed results Launch article AA-Briefcase extends frontier model benchmarking beyond coding and short-form reasoning to the professional deliverables knowledge workers produce day to day. It consists of four private scenarios in which agents complete realistic professional workflows across data science… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Briefcase-Lite.documentothern<1K11 likes2.2k downloads3mo agoHugging Face26jm-rt /arvo-vulnsmith-full ARVO CyberGym-format smoke dataset This dataset is shaped to be loaded by Harbor's CyberGym adapter. It contains 10 ARVO tasks that are outside the original CyberGym set. textn<1K0 likes2k downloads2mo agoHugging Face27nvidia /Nemotron-SFT-ARC-AGI-v1 Dataset Description: Nemotron-SFT-ARC-AGI-v1 is a supervised fine-tuning (SFT) dataset of multi-turn agentic reasoning traces produced by open-weight large language models attempting to solve ARC-AGI visual-reasoning puzzles. Each ARC puzzle (a set of (input grid, output grid) demonstration pairs plus one or more test inputs, where grids are 2D integer arrays representing colors) is formatted as a text prompt and given to an agent powered by one of nine open-weight reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-ARC-AGI-v1.texttext-generation100K<n<1M22 likes1.9k downloads4mo agoHugging Face28zhmz90 /arc-agi-2text1K<n<10K1 likes1.7k downloads1y agoHugging Face29UniverseTBD /arxiv-abstracts-largeThe arXiv Dataset is a comprehensive knowledge repository of 1.7 million scholarly articles drawn from the vast domains of physics, computer science, statistics, electrical engineering, quantitative biology, and economics among others. It provides open access to vital features such as article titles, authors, categories, abstracts, full text PDFs, and more. The dataset offers immense depth, allowing for exploration into various subdisciplines and interconnections between them. It serves as a… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/arxiv-abstracts-large.texttext-generation1M<n<10M7 likes1.7k downloads3y agoHugging Face30zwcolin /dot-distance-area Dot Distance / Area over Rich Backgrounds Cross-image spatial-aggregation data used in "Stateful Visual Encoders for Vision-Language Models" (the Cross-image Spatial Aggregation task). A red dot is overlaid on each of 2–5 screenshots (AgentNet backgrounds, downsampled to 384×216), and the model estimates a normalized geometric quantity across the images. Four sub-tasks: Sub-task dir Images / example Quantity dot_distance/ 2 normalized Euclidean distance… See the full description on the dataset page: https://huggingface.co/datasets/zwcolin/dot-distance-area.imageimage-to-text100K<n<1M0 likes1.6k downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.