CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Bekhouche /HalluTruthQA-4K HalluTruthQA-4K HalluTruthQA-4K is the official data release for Subtask 2.2 ("Hallucination Detection and Find the Truth") of the HalluScoring 2026 shared task, hosted at ArabicNLP 2026. It extends the HalluTruthQA benchmark from 2,400 to 4,000 expert-annotated Arabic question-answering instances across four knowledge-intensive domains. Dataset Summary The full corpus is 4,000 Arabic question-answering instances, exactly balanced across four domains (1,000… See the full description on the dataset page: https://huggingface.co/datasets/Bekhouche/HalluTruthQA-4K.texttext-classification1K<n<10K0 likes116 downloads1mo agoHugging Face02rqq /GLM-4-Instruct-4K-zh Dataset Card for Dataset Name ❤️欢迎使用rqq/GLM-4-Instruct-4K-zh数据集,本数据集包含了4000条高质量的glm4回复。 该数据集的提问数据源自高质量的Sao10K/Claude-3-Opus-Instruct-5K数据集,我们把它的问题翻译成了中文,使用glm-4进行了重新回答。 该数据集使用alpaca格式,可以直接用在llama-factory项目中进行训练! 文件如下: GLM-4-Instruct-4K-zh.json 问答数据集,alpaca格式 GLM-4-question-translate-5K-zh 翻译-对话数据集,记录了把Sao10K/Claude-3-Opus-Instruct-5K问题翻译成中文的数据 Welcome to the rqq/GLM-4-Instruct-4K-zh dataset! This dataset includes 4,000 high-quality responses from the GLM-4 model. The question data… See the full description on the dataset page: https://huggingface.co/datasets/rqq/GLM-4-Instruct-4K-zh.texttranslation1K<n<10K33 likes79 downloads2y agoHugging Face03nphearum /grpo-4k-reasoning-tools 🧠 GRPO 4K Reasoning Tools A compact, high-quality dataset designed to train and evaluate reasoning-capable LLMs with tool usage and function calling. This dataset focuses on structured thinking, multi-step reasoning, and practical tool integration across diverse NLP tasks. 📌 Overview GRPO 4K Reasoning Tools is a curated dataset of ~4K examples that combines: 🧩 Step-by-step reasoning (“thinking” traces) 🛠️ Tool usage / function calling patterns 🧠 Multi-task learning… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/grpo-4k-reasoning-tools.textquestion-answering1K<n<10K0 likes45 downloads6mo agoHugging Face04ShaoRun /RS-EoT-4K Asking like Socrates: RS-EoT-4K Dataset 🌐 Project Website | 💻 GitHub Repository | 📄 Paper (ArXiv) | 🤗 Model (RS-EoT-7B) 📖 Introduction RS-EoT-4K is a multimodal instruction-tuning dataset specifically designed to instill Evidence-of-Thought (EoT) reasoning capabilities into Vision-Language Models (VLMs) for Remote Sensing (RS) tasks. This dataset was introduced in the paper "Asking like Socrates: Socrates helps VLMs understand remote sensing images". To… See the full description on the dataset page: https://huggingface.co/datasets/ShaoRun/RS-EoT-4K.imagequestion-answering1K<n<10K4 likes44 downloads10mo agoHugging Face05Rakancorle1 /hans-sft-4k Hans-SFT-4K · SFT recipe for the audio-visual Clever Hans Supervised fine-tuning (SFT) data accompanying the paper When Vision Speaks for Sound. Like the original Clever Hans — the horse that looked like he could do arithmetic but was actually reading his trainer's body language — video-capable MLLMs often look like they can hear: they answer audio questions by reading visual cues and never verifying the audio stream. Hans-SFT-4K is the 3,834-sample SFT mix that teaches models to… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-sft-4k.audioaudio-classification1K<n<10K1 likes24 downloads4mo agoHugging Face06SimonSun /PyMath-AgentRL-4K 🧮 PyMath-AgentRL-4K A supervised fine-tuning (SFT) dataset for training language models to solve mathematical problems using Python code interpreters. Overview Total Examples: 4,000 (deduplicated) Format: Parquet Language: English License: Apache 2.0 Source Datasets Merged from two high-quality sources: Gen-Verse/Open-AgentRL-SFT-3K (3,000 examples) JoeYing/ReTool-SFT (2,000 examples) Deduplication: 20% overlap removed (1,000 duplicates) Key… See the full description on the dataset page: https://huggingface.co/datasets/SimonSun/PyMath-AgentRL-4K.textquestion-answering1K<n<10K0 likes22 downloads9mo agoHugging Face07Ocheretny /skolkovo-coaching-qa-4k Skolkovo Executive Coaching Q&A Dataset Описание Датасет Q&A пар для обучения ИИ-коуча по программе Executive Coaching от Skolkovo. Размер: 3,901 пар (893 оригинальных + 3,008 augmented) Источники: Административные документы Skolkovo Лекции по коучингу Книги по коучингу и лидерству (18 книг) Кейсы и практические материалы Структура данных { "question": "Как мотивировать команду?", "answer": "Для мотивации команды важно...", "role": "coach", "source":… See the full description on the dataset page: https://huggingface.co/datasets/Ocheretny/skolkovo-coaching-qa-4k.textquestion-answering1K<n<10K0 likes17 downloads11mo agoHugging Face08stindardlogic /rag-grounding-dpo-4k RAG Grounding DPO Pairs (4K) DPO preference pairs for training LLMs to faithfully use retrieved context in RAG pipelines. Motivation RAG is the dominant LLM deployment pattern in production. The core failure mode: models that ignore retrieved context and hallucinate answers, or that contradict documents with confidently-stated fabrications. This dataset trains models to ground answers in provided context. Dataset Description 4,000 preference pairs… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/rag-grounding-dpo-4k.texttext-generation1K<n<10K0 likes13 downloads2mo agoHugging Face09Vishva007 /Databricks-Dolly-4k Databricks-Dolly-4k The resulting dataset contains 4000 samples of the databricks/databricks-dolly-15k dataset. This split of an even smaller subset is provided for very fast experimentation and evaluation of models when computational resources are highly limited or for quick prototyping. Dataset Structure The dataset is provided as a DatasetDict with the following splits: train: Contains 4000 samples. Each split contains the following features, identical to the… See the full description on the dataset page: https://huggingface.co/datasets/Vishva007/Databricks-Dolly-4k.texttable-question-answering1K<n<10K0 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.