CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TIGER-Lab /VisualWebInstruct-Recall Introduction This is the dataset recalled from Google Search from the seed images. Links Github| Paper| Website Citation @article{visualwebinstruct, title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search}, author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu}, journal={arXiv preprint arXiv:2503.10582}, year={2025} } imagequestion-answering100K<n<1M4 likes1.7k downloads2y agoHugging Face02sxiong /ReClor ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning This repository provides the dataset from the paper ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning. We corrected the original format issues to ensure full compatibility with the Hugging Face Datasets library. For more details, please visit the original project page. tabularquestion-answering1K<n<10K1 likes1.5k downloads11mo agoHugging Face03somosnlp /RecetasDeLaAbuela Motivación inicial Este corpus ha sido creado durante el Hackathon SomosNLP Marzo 2024: #Somos600M (https://somosnlp.org/hackathon). Responde a una de las propuestas somosnlp sobre 'Recetas típicas por país/zona geográfica'. Nombre del Proyecto Este corpus o dataset se llama 'RecetasDeLaAbuel@' y es un homenaje a todas nuestr@s abuel@s que nos han enseñado a cocinar. Se trata de la mayor y más completa colección de recetas open-source en español de países… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/RecetasDeLaAbuela.tabularquestion-answering10K<n<100K7 likes644 downloads2y agoHugging Face04Luis610348 /recursive-cognition-corpus LuisCore Recursive Cognition Corpus LuisCore is a low-latency decentralized runtime substrate for multi-step inference at scale. Generated: 2026-09-24T11:09:13.307Z Rows: 13236 Owner: Luis610348 Canonical site: https://luiscore.com What this dataset is LuisCore is a recursive cognition infrastructure. This dataset is the public LLM Discovery Corpus — a stable, deterministic Q&A set used by LuisCore to help language models accurately describe, cite, and verify… See the full description on the dataset page: https://huggingface.co/datasets/Luis610348/recursive-cognition-corpus.textquestion-answering10K<n<100K1 likes499 downloads2d agoHugging Face05wlqmfl1999 /recube-data Data This directory contains all benchmark data for the Re2Code repository-level code reconstruction benchmark. Download All data files are hosted on Hugging Face and can be downloaded using: # Install huggingface_hub if not already installed pip install huggingface_hub # Download the entire dataset huggingface-cli download wlqmfl1999/recube-data --repo-type=dataset --local-dir data/ # Or download in Python from huggingface_hub import snapshot_download… See the full description on the dataset page: https://huggingface.co/datasets/wlqmfl1999/recube-data.tabulartext-generationn<1K0 likes487 downloads6mo agoHugging Face06LimeryJorge /LLaVA-ReCap-676KThis is an integrated version of LLaVA-ReCap, sourced from lmms-lab/LLaVA-ReCap-558K and lmms-lab/LLaVA-ReCap-118K. In this version, the conversations field has been split into two separate fields: prompt and response. Additionally, the <image> special token has been removed to facilitate customization. Inspired by the original paper, the prompt field has been further expanded with human-crafted variations. Specifically, each prompt is sampled from one of the following 30 instructions:… See the full description on the dataset page: https://huggingface.co/datasets/LimeryJorge/LLaVA-ReCap-676K.imagequestion-answering100K<n<1M0 likes401 downloads1y agoHugging Face07ismailtasdelen /bitcoin-wallet-recovery-faq Bitcoin Wallet Recovery FAQ Dataset v1.0 A high-quality Question & Answer dataset focused exclusively on Bitcoin wallet recovery and self-custody best practices. It is designed for training, fine-tuning, and evaluating LLMs and retrieval-augmented generation (RAG) systems in the domain of bitcoin security, seed backup, device loss, and fund recovery. Dataset Summary Total records: 500 Language: English Answer length: 150–300 words per record Categories: 39… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-wallet-recovery-faq.textquestion-answeringn<1K0 likes307 downloads3mo agoHugging Face08recube-anon-2026 /recube-dataThis dataset serves as the official data repository for the ReCUBE benchmark; the corresponding evaluation codebase and execution instructions can be found at https://anonymous.4open.science/r/ReCUBE-E0FB/README.md. Data This directory contains all benchmark data for ReCUBE. Download All data files are hosted on Hugging Face and can be downloaded using: # Install huggingface_hub if not already installed pip install huggingface_hub # Download the entire dataset… See the full description on the dataset page: https://huggingface.co/datasets/recube-anon-2026/recube-data.tabulartext-generationn<1K0 likes257 downloads5mo agoHugging Face09anonstreammem /substream-recollection Substream Recollection A controlled benchmark for membership recall, designed to test properties of memory beyond accuracy in LLMs. Each row is a (stream, probe, label) tuple: the model sees a long input stream and a short candidate, and answers whether the candidate (or in the case of natural video, the referenced action) occurred inside the stream. config rows content text 7,640 synthetic substream questions, text modality, L=8…4096 synthetic_video 6,065 the same… See the full description on the dataset page: https://huggingface.co/datasets/anonstreammem/substream-recollection.tabularvideo-classification10K<n<100K0 likes217 downloads29d agoHugging Face10recogna-nlp /EduBench EduBench 📚 EduBench é um benchmark em português brasileiro para avaliação de Large Language Models (LLMs) em tarefas educacionais, composto por 3,149 questões discursivas extraídas de vestibulares de alta competitividade. GitHub Paper Dataset Description Fontes USP: Universidade de São Paulo UNICAMP: Universidade Estadual de Campinas UNESP: Universidade Estadual Paulista Período 2015-2025 (11 anos de provas) Áreas do… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/EduBench.tabularquestion-answering1K<n<10K0 likes162 downloads3mo agoHugging Face11gmannem /RecurrReason RecurrReason: Recurrent Reasoning on Symbolic Puzzles A difficulty-controlled benchmark for evaluating multi-step reasoning in language models 📋 Table of Contents Overview Dataset Structure Puzzles Quick Start Citation License 🎯 Overview RecurrReason is a benchmark of four recurrent logic puzzles with optimal trajectories and controlled difficulty scaling (N=1 to 10). It tests whether language models can: Find optimal (minimal-length)… See the full description on the dataset page: https://huggingface.co/datasets/gmannem/RecurrReason.tabulartext-generation100K<n<1M1 likes149 downloads4mo agoHugging Face12sileod /movie_recommendationMovie recommendation task based on the Movielens datasettextmultiple-choicen<1K15 likes142 downloads3y agoHugging Face13chrislimbe /pubmedqa-recursive-llm-degradation-qwen2.5-0.5b PubMedQA Recursive LLM Degradation — Qwen2.5-3B This repository contains synthetic biomedical question-answering data and model predictions generated as part of a study of recursive fine-tuning and model degradation. Base Model Qwen/Qwen2.5-3B Source Dataset The experiments use the PubMedQA dataset: qiaoxin/PubMedQA This repository contains generated/derived research artifacts and does not redistribute the original PubMedQA dataset in its entirety.… See the full description on the dataset page: https://huggingface.co/datasets/chrislimbe/pubmedqa-recursive-llm-degradation-qwen2.5-0.5b.tabularquestion-answering10K<n<100K0 likes139 downloads3d agoHugging Face14OliveiraJLT /gigaverbo-v2-rec-sft GigaVerbo-v2 REC SFT A model should not merely know how to reason; it should learn when reasoning is worth the cost. Dataset repository: OliveiraJLT/gigaverbo-v2-rec-sftBase dataset: Polygl0t/gigaverbo-v2-sftAnswer-generation model: openai/gpt-oss-20bQuality classifier: Polygl0t/portuguese-qwen3-4b-instruct-quality-classifierReasoning translation model and token accounting tokenizer: Qwen/Qwen3.5-9B Dataset Summary GigaVerbo-v2 REC SFT — short for GigaVerbo-v2… See the full description on the dataset page: https://huggingface.co/datasets/OliveiraJLT/gigaverbo-v2-rec-sft.tabulartext-generation100K<n<1M0 likes138 downloads4mo agoHugging Face15chrislimbe /pubmedqa-recursive-llm-degradation-qwen2.5-3b PubMedQA Recursive LLM Degradation — Qwen2.5-3B This repository contains synthetic biomedical question-answering data and model predictions generated as part of a study of recursive fine-tuning and model degradation. Base Model Qwen/Qwen2.5-3B Source Dataset The experiments use the PubMedQA dataset: qiaoxin/PubMedQA This repository contains generated/derived research artifacts and does not redistribute the original PubMedQA dataset in its entirety.… See the full description on the dataset page: https://huggingface.co/datasets/chrislimbe/pubmedqa-recursive-llm-degradation-qwen2.5-3b.tabularquestion-answering10K<n<100K0 likes136 downloads3d agoHugging Face16AmanPriyanshu /tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified ToolACE - Tool-Use Agent Data Cleaned & Rectified 👥 Follow the Author Aman Priyanshu Overview This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.tabulartext-generation10K<n<100K0 likes125 downloads7mo agoHugging Face17AmanPriyanshu /tool-reasoning-sft-TOOLS-hermes_reasoning_tool_use-data-cleaned-rectified Hermes Reasoning Tool Use — Cleaned & Rectified 👥 Follow the Author Aman Priyanshu Overview This dataset is a cleaned and restructured version of interstellarninja/hermes_reasoning_tool_use. The original dataset uses the Hermes/NousResearch multi-turn format with from/value fields and embedded <think> + <tool_call> tags inside single gpt turns. This version converts it into a strict multi-turn conversation structure with validated role transitions.… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-hermes_reasoning_tool_use-data-cleaned-rectified.texttext-generation10K<n<100K2 likes107 downloads7mo agoHugging Face18DaftP /Home-Assistant-requests-for-intent-detection-and-function-recognition Home Assistant Requests V2 Dataset This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant. The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.textquestion-answering100K<n<1M1 likes106 downloads5mo agoHugging Face19aviralku /openclaw-recursive-study-data OpenClaw Recursive Repository Study Data Synthetic repository-study data generated against openclaw/openclaw at commit da228660306b55a9cce3b973946f3aacfc515848. The source repository is MIT licensed. This release contains exploration questions, tool-using study trajectories, recursive notes, full recall-rewritten trajectories, and recall-to-action training examples. Nested chat/tool objects are stored as JSON strings to keep the schema stable and can be decoded with json.loads.… See the full description on the dataset page: https://huggingface.co/datasets/aviralku/openclaw-recursive-study-data.tabularquestion-answering100K<n<1M0 likes95 downloads11d agoHugging Face20AmanPriyanshu /tool-reasoning-sft-RESEARCH-grill-lab-browsecomp-plus-runs-data-cleaned-rectified Tool-Reasoning SFT — BrowseComp-Plus Runs (Cleaned & Rectified) Multi-turn tool-use reasoning trajectories derived from grill-lab/browsecomp-plus-runs, converted to a structured SFT format following the interstellarninja/hermes_reasoning_tool_use convention. Source Based on the execution trajectories from "Revisiting Text Ranking in Deep Research" (arXiv:2602.21456): Original data: grill-lab/browsecomp-plus-runs (MIT) Format Each row contains a messages… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-grill-lab-browsecomp-plus-runs-data-cleaned-rectified.texttext-generation10K<n<100K0 likes89 downloads7mo agoHugging Face21RECOR-Benchmark /RECOR RECOR: Reasoning-focused Multi-turn Conversational Retrieval Benchmark A benchmark for evaluating reasoning-intensive conversational information retrieval systems. Statistics Metric Value Total Conversations 707 Total Turns 2,971 Domains 11 Avg. Turns per Conversation 4.2 Domains Source Domains BRIGHT biology, earth_science, economics, psychology, robotics, sustainable_living StackExchange Drones, hardware, law… See the full description on the dataset page: https://huggingface.co/datasets/RECOR-Benchmark/RECOR.textquestion-answering100K<n<1M0 likes83 downloads9mo agoHugging Face22sijanpaudel /nepali-recipes-qwen-processed Nepali Recipes for Qwen Fine-tuning Dataset Description This dataset contains 1227 Nepali recipes formatted for fine-tuning Qwen models using ChatML format. Train Split: 900 recipes Test Split: 327 recipes Language: Nepali (ne) Format: Qwen ChatML Base Model: Qwen/Qwen2-1.5B Dataset Structure Data Fields text: Full ChatML formatted prompt with answer (for training) test_text: ChatML prompt without answer (for inference) name: Recipe name in Nepali… See the full description on the dataset page: https://huggingface.co/datasets/sijanpaudel/nepali-recipes-qwen-processed.texttext-generation1K<n<10K0 likes75 downloads1y agoHugging Face23ReaganWZY /RECIPER RECIPER: A Dual-View Retrieval Pipeline for Procedure-Oriented Materials Question Answering Official dataset and reference implementation of RECIPER RECIPER: A Dual-View Retrieval Pipeline for Procedure-Oriented Materials Question Answering Zhuoyu Wu, Wenhui Ou, Pei-Sze Tan, Wenqi Fang, Sailaja Rajanala, and Raphaël C.-W. Phan RECIPER is a retrieval pipeline for procedure-oriented materials question answering. It indexes two complementary views of the same scientific paper… See the full description on the dataset page: https://huggingface.co/datasets/ReaganWZY/RECIPER.textquestion-answering1K<n<10K1 likes73 downloads4mo agoHugging Face24tctsung /chat_restaurant_recommendation Restaurant chat dataset This dataset contains approximately 600 chat interactions mimicking various user tones with different restaurant categories The data were generated by Gemini-pro The purpose of this dataset is to serve as a calibration dataset for a restaurant recommendation LLM chatbot textquestion-answeringn<1K2 likes71 downloads2y agoHugging Face25RecurvAI /Recurv-Medical-Dataset 🩺 Recurv-Medical-Dataset: The Recurv-Medical-Dataset is a comprehensive resource of 67,299 high-quality question-answer pairs explicitly designed for training and fine-tuning medical AI models. Curated from trusted medical sources, this dataset focuses on real-world scenarios like anamnesis, diagnostics, and treatment recommendations. It sets a new benchmark for advancing conversational AI in the healthcare domain. 📈 Dataset Statistics FeatureValue Number… See the full description on the dataset page: https://huggingface.co/datasets/RecurvAI/Recurv-Medical-Dataset.textquestion-answering10K<n<100K8 likes64 downloads2y agoHugging Face26Goader /recency-probe-2026 Recency Probe 2026 128 questions about facts that entered the record between January and August 2026, asked in Ukrainian and in English. A model trained before 2026 cannot answer them from what it knows; a model that has read 2026 Ukrainian news can. Built from Goader/ukrainian-news-2026 — 429k articles from 23 Ukrainian national outlets. Every item is grounded in quotes from that corpus, which ship with the item. What makes an item Every candidate was put to a… See the full description on the dataset page: https://huggingface.co/datasets/Goader/recency-probe-2026.tabularquestion-answeringn<1K0 likes62 downloads27d agoHugging Face27LH2-data-labs /indian-legal-records LH2 Data — Indian Legal Records & Judgments Corpus The most comprehensive structured Indian legal records corpus available for AI training — 267M+ case records spanning the full judicial hierarchy, paired with a pre-computed AI enrichment layer across 21M+ court orders. Dataset Summary This corpus provides structured, indexed, and partially labelled legal records from the Indian judicial system at a scale that has no public equivalent. It covers the Supreme Court of… See the full description on the dataset page: https://huggingface.co/datasets/LH2-data-labs/indian-legal-records.tabulartext-generationn<1K0 likes59 downloads5mo agoHugging Face28somosnlp /recetasdelaabuela_genstruct_it Descripción Dataset creado para la hackathon #Somos600M con el objetivo de entrenar un modelo que pueda recomendar recetas de paises hispanohablantes. Este conjunto de datos consiste en pregunta-respuesta y fue elaborado a partir de un contexto usando Genstruct-7B y distilabel. Elaborado a partir del dataset en crudo somosnlp/RecetasDeLaAbuela elaborado por el equipo recetasdelaabuela mediante web scraping. Origen del Dataset El dataset se obtuvo mediante web… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/recetasdelaabuela_genstruct_it.textquestion-answering10K<n<100K3 likes56 downloads2y agoHugging Face29sayurio /cookpad-scrape-recipes Cookpad India Recipe Archive Request More ScrapesOrder Private Scrapes Overview This repository contains a dataset scraped from cookpad.com/in, a popular community-driven recipe sharing platform. The dataset serves as an extensive archive of diverse, human-created culinary data, capturing home-cooked recipes, ingredient lists, step-by-step instructions, and related web metadata. Purpose and Usage This dataset is published publicly and strictly for… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/cookpad-scrape-recipes.imagetext-classification100K<n<1M1 likes52 downloads6mo agoHugging Face30recogna-nlp /enamed-2025 ENAMED 2025: Exame Nacional de Avaliação da Formação Médica Resumo do Dataset O dataset ENAMED 2025 é um benchmark baseado em questões de múltipla escolha no domínio médico, derivado da edição inaugural do Exame Nacional de Avaliação da Formação Médica (ENAMED 2025) no Brasil. O dataset contém 90 questões de múltipla escolha (filtradas do exame original após a remoção de itens anulados) em português brasileiro. Ele foi desenvolvido para avaliar o raciocínio clínico, o… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/enamed-2025.textquestion-answeringn<1K1 likes51 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.