CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hw-hwei /MedThoughts-8K MedThoughts-8K English|中文 This dataset is distilled from the full-scale DeepSeek-R1 (671B) in the medical domain. For more detailed information, please refer to our GitHub project MedR1. 1. Original Dataset The data in this dataset is sourced from the US/train partition of MedQA (5 options). 2. Dataset Format The keys in the dataset are explained as follows: "question_id": The unique identifier for the question, "question": The question itself, "options": The… See the full description on the dataset page: https://huggingface.co/datasets/hw-hwei/MedThoughts-8K.textquestion-answering1K<n<10K4 likes110 downloads2y agoHugging Face02oddadmix /arabic-rag-chat-8k-eval arabic-rag-chat-8k-eval Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models: the test split, every model's raw replies, every judge verdict, and the rendered report for each. Thirteen judged models, all scored on the same 1,651 prompts by the same judge at temperature 0.0, so the comparison below is like-for-like and can be recomputed offline without a GPU or a judge server. This is the measurement half of oddadmix/100M-8192-Nawah-dsv4; the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.tabularquestion-answeringn<1K0 likes72 downloads1mo agoHugging Face03Jakumetsu /A-MMK12-8K A-MMK12-8K This dataset is created using the SynthRL pipeline to synthesize 3,380 challenging questions from 8,072 seed samples. Dataset Details Synthesis Method: SynthRL pipeline Total Samples: 11,452 (8,072 seed + 3,380 synthesized) Seed Data: MMK12 dataset Purpose: Training data for VLM reinforcement learning with verifiable rewards (RLVR) Data Sources Original MMK12: Proposed in MM-EUREKA (thanks to the MM-EUREKA authors) Processed Seed Data: Our… See the full description on the dataset page: https://huggingface.co/datasets/Jakumetsu/A-MMK12-8K.imagequestion-answering10K<n<100K2 likes33 downloads1y agoHugging Face04ceadar-ie /AIVision360-8k Dataset Card for AIVision360-8k Dataset Description AIVision360 is the pioneering domain-specific dataset tailor-made for media and journalism, designed expressly for the instruction fine-tuning of Large Language Models (LLMs).The AIVision360-8k dataset is a curated collection sourced from "ainewshub.ie", a platform dedicated to Artificial Intelligence news from quality-controlled publishers. It is designed to provide a comprehensive representation of AI-related… See the full description on the dataset page: https://huggingface.co/datasets/ceadar-ie/AIVision360-8k.textquestion-answering1K<n<10K4 likes26 downloads3y agoHugging Face05sardukar /physiology-mcqa-8kThis dataset is a subset of MedMCQA textquestion-answering1K<n<10K2 likes26 downloads2y agoHugging Face06GaborMadarasz /HuHotpotQA_8k HuHotpotQA HuHotpotQA is a Hungarian-language multi-hop question answering dataset designed in the style of HotpotQA. It contains approximately 2,000 question-answer pairs based on articles from Hungarian Wikipedia. The dataset is designed to evaluate and train models on questions that require combining information from multiple documents rather than retrieving an answer from a single context. This is a truncated version of the original dataset. To count the… See the full description on the dataset page: https://huggingface.co/datasets/GaborMadarasz/HuHotpotQA_8k.textquestion-answering1K<n<10K0 likes18 downloads1mo agoHugging Face07jtatman /databricks-dolly-8k-qa-open-closetextsummarization1K<n<10K0 likes17 downloads3y agoHugging Face08Vishva007 /Databricks-Dolly-8k Databricks-Dolly-8k The resulting dataset contains 8000 samples of the databricks/databricks-dolly-15k dataset. This split of an even smaller subset is provided for very fast experimentation and evaluation of models when computational resources are highly limited or for quick prototyping. Dataset Structure The dataset is provided as a DatasetDict with the following splits: train: Contains 8000 samples. Each split contains the following features, identical to the… See the full description on the dataset page: https://huggingface.co/datasets/Vishva007/Databricks-Dolly-8k.texttable-question-answering1K<n<10K0 likes17 downloads1y agoHugging Face09alex-chiu /OpenMathInstruct-2-Boxed-8k alex-chiu/OpenMathInstruct-2-Boxed-8k A deterministic 8K subset of nvidia/OpenMathInstruct-2 for short non-thinking math SFT before RL. Splits train: 8,192 rows validation: 256 rows Construction Every generated_solution contains a complete final \boxed{...} Duplicate problems and explicit <think> / </think> outputs are excluded No answer verification against expected_answer is performed The original generated_solution is preserved verbatim… See the full description on the dataset page: https://huggingface.co/datasets/alex-chiu/OpenMathInstruct-2-Boxed-8k.tabulartext-generation1K<n<10K0 likes11 downloads3mo agoHugging Face10Vishva007 /Flickr-Dataset-8k Flickr1k This dataset is a subset of the Original Flickr30k Dataset, containing [Total number of samples, e.g., 8000] image-caption pairs. It has been specifically created for [Briefly state the purpose, e.g., faster experimentation with image captioning models or a specific research focus]. Dataset Details Original Dataset: Original Flickr30k dataset on Hugging Face Hub Subset Size: 8000 Data Format: Each example contains the following fields: image: The raw bytes of… See the full description on the dataset page: https://huggingface.co/datasets/Vishva007/Flickr-Dataset-8k.imagetext-generation1K<n<10K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.