datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MedThoughts-8K
MedThoughts-8K
English|中文
This dataset is distilled from the full-scale DeepSeek-R1 (671B) in the medical domain. For more detailed information, please refer to our GitHub project MedR1.
1. Original Dataset
The data in this dataset is sourced from the US/train partition of MedQA (5 options).
2. Dataset Format
The keys in the dataset are explained as follows:
"question_id": The unique identifier for the question,
"question": The question itself,
"options": The… See the full description on the dataset page: https://huggingface.co/datasets/hw-hwei/MedThoughts-8K.arabic-rag-chat-8k-eval
arabic-rag-chat-8k-eval
Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models:
the test split, every model's raw replies, every judge verdict, and the rendered
report for each. Thirteen judged models, all scored on the same 1,651 prompts
by the same judge at temperature 0.0, so the comparison below is like-for-like
and can be recomputed offline without a GPU or a judge server.
This is the measurement half of
oddadmix/100M-8192-Nawah-dsv4;
the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.A-MMK12-8K
A-MMK12-8K
This dataset is created using the SynthRL pipeline to synthesize 3,380 challenging questions from 8,072 seed samples.
Dataset Details
Synthesis Method: SynthRL pipeline
Total Samples: 11,452 (8,072 seed + 3,380 synthesized)
Seed Data: MMK12 dataset
Purpose: Training data for VLM reinforcement learning with verifiable rewards (RLVR)
Data Sources
Original MMK12: Proposed in MM-EUREKA (thanks to the MM-EUREKA authors)
Processed Seed Data: Our… See the full description on the dataset page: https://huggingface.co/datasets/Jakumetsu/A-MMK12-8K.AIVision360-8k
Dataset Card for AIVision360-8k
Dataset Description
AIVision360 is the pioneering domain-specific dataset tailor-made for media and journalism, designed expressly for the instruction fine-tuning of Large Language Models (LLMs).The AIVision360-8k dataset is a curated collection sourced from "ainewshub.ie", a platform dedicated to Artificial Intelligence news from quality-controlled publishers. It is designed to provide a comprehensive representation of AI-related… See the full description on the dataset page: https://huggingface.co/datasets/ceadar-ie/AIVision360-8k.physiology-mcqa-8kThis dataset is a subset of MedMCQA
HuHotpotQA_8k
HuHotpotQA
HuHotpotQA is a Hungarian-language multi-hop question answering dataset designed in the style of HotpotQA. It contains approximately 2,000 question-answer pairs based on articles from Hungarian Wikipedia.
The dataset is designed to evaluate and train models on questions that require combining information from multiple documents rather than retrieving an answer from a single context.
This is a truncated version of the original dataset.
To count the… See the full description on the dataset page: https://huggingface.co/datasets/GaborMadarasz/HuHotpotQA_8k.databricks-dolly-8k-qa-open-closeDatabricks-Dolly-8k
Databricks-Dolly-8k
The resulting dataset contains 8000 samples of the databricks/databricks-dolly-15k dataset.
This split of an even smaller subset is provided for very fast experimentation and evaluation of models when computational resources are highly limited or for quick prototyping.
Dataset Structure
The dataset is provided as a DatasetDict with the following splits:
train: Contains 8000 samples.
Each split contains the following features, identical to the… See the full description on the dataset page: https://huggingface.co/datasets/Vishva007/Databricks-Dolly-8k.OpenMathInstruct-2-Boxed-8k
alex-chiu/OpenMathInstruct-2-Boxed-8k
A deterministic 8K subset of nvidia/OpenMathInstruct-2 for short non-thinking math SFT before RL.
Splits
train: 8,192 rows
validation: 256 rows
Construction
Every generated_solution contains a complete final \boxed{...}
Duplicate problems and explicit <think> / </think> outputs are excluded
No answer verification against expected_answer is performed
The original generated_solution is preserved verbatim… See the full description on the dataset page: https://huggingface.co/datasets/alex-chiu/OpenMathInstruct-2-Boxed-8k.Flickr-Dataset-8k
Flickr1k
This dataset is a subset of the Original Flickr30k Dataset, containing [Total number of samples, e.g., 8000] image-caption pairs. It has been specifically created for [Briefly state the purpose, e.g., faster experimentation with image captioning models or a specific research focus].
Dataset Details
Original Dataset: Original Flickr30k dataset on Hugging Face Hub
Subset Size: 8000
Data Format: Each example contains the following fields:
image: The raw bytes of… See the full description on the dataset page: https://huggingface.co/datasets/Vishva007/Flickr-Dataset-8k.
