datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LIMODataset for LIMO: Less is More for Reasoning
Usage
from datasets import load_dataset
dataset = load_dataset("GAIR/LIMO", split="train")
Citation
If you find our dataset useful, please cite:
@misc{ye2025limoreasoning,
title={LIMO: Less is More for Reasoning},
author={Yixin Ye and Zhen Huang and Yang Xiao and Ethan Chern and Shijie Xia and Pengfei Liu},
year={2025},
eprint={2502.03387},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/GAIR/LIMO.Sci-Fi-ZH一份 VeejaLiu 正在手工清洗的数据:https://github.com/VeejaLiu/ScienceFictionCollection
LIMO-v2
LIMO: Less is More for Reasoning
Dataset for LIMO: Less is More for Reasoning
This is the updated version (v2) of the LIMO dataset, corresponding to the latest paper version as of July 30, 2025.
Usage
from datasets import load_dataset
# Load LIMO-v2 dataset
dataset = load_dataset("GAIR/LIMO-v2", split="train")
Previous Version
If you need the original LIMO dataset (corresponding to the initial paper version), you can access it at:
LIMO v1: GAIR/LIMO
# To… See the full description on the dataset page: https://huggingface.co/datasets/GAIR/LIMO-v2.b-corpus
📊 Statistic
约 7.86 million 行中文对话,183 million tokens
⚠️注意
请注意,数据来自 R18 的视觉小说,并且包含可能被认为是不适当、令人震惊、令人不安、令人反感和极端的主题。如果您不确定在您的国家拥有任何形式的虚构文字内容的法律后果,请不要下载。
本项目内的所有数据及基于这些数据的衍生作品禁止用作商业性目的。
⚠️Warning
Please note that the data comes from R18 visual novels and contains themes that may be considered inappropriate, shocking, disturbing, offensive, and extreme. If you are unsure about the legal implications of possessing any form of fictional written content… See the full description on the dataset page: https://huggingface.co/datasets/Limour/b-corpus.testeqwnh-corpus-raw未清洗的中文H小说
仅供科学研究使用!
LIMOlimo-math8k-distill-16k-reasoning-tracesG2Retrieval视觉小说 领域的 Retrieval 评价数据集。
Leaderboard
data_sample2k
https://www.kaggle.com/code/reginliu/g2retrieval
Model
NDCG@3
NDCG@10
NDCG@50
NDCG@100
NDCG@200
acge_text_embedding
83.53±17.86
76.97±17.79
61.52±20.61
52.07±20.87
42.49±19.83
IYun-large-zh
80.53±20.53
71.40±20.87
52.93±21.96
43.40±20.72
34.88±18.50
bce-embedding-base_v1
77.08±23.44
68.39±22.61
51.95±22.85
43.36±21.51
35.31±19.09
Dmeta-embedding
77.56±22.12
68.62±21.96
51.58±22.29
42.71±21.04… See the full description on the dataset page: https://huggingface.co/datasets/Limour/G2Retrieval.MSMS-LIMO-v2-SFT
A Multi-Source Multi-Solution Long CoT SFT Dataset from LIMO-v2
H2Retrievalh-corpus 领域的 Retrieval 评价数据集。
Leaderboard
new/data_sample1k
https://www.kaggle.com/code/reginliu/h2retrieval
Model
NDCG@5
NDCG@10
NDCG@15
NDCG@20
NDCG@30
IYun-large-zh
66.70±27.29
59.67±26.05
56.69±25.36
56.58±25.32
57.97±25.48
acge_text_embedding
64.60±28.04
57.80±25.88
55.54±25.166
55.77±25.17
57.31±25.18
bce-embedding-base_v1
60.66±28.37
53.44±26.1351.11±25.10
51.18±25.16
52.84±25.45
Dmeta-embedding
52.12±29.83
45.38±26.65
43.20±25.33
43.41±25.10… See the full description on the dataset page: https://huggingface.co/datasets/Limour/H2Retrieval.LIMO-v2-harmonized
LIMO-v2-harmonized
Direct-ingest Harmony mirror of GAIR/LIMO-v2.
License note: the source card publishes Apache-2.0.
LIMO-best_of_n-VLLM-Skywork-o1-Open-PRM-Qwen-2.5-7B-completionsarchvieempathy-affective-datasetsEmpathy and Affective Computing Datasets Summary
This repository is a curated summary of existing datasets for empathy and affective computing research. It distinguishes between empathy-focused datasets (directly measuring empathic processes) and general affective computing datasets (emotion recognition, valence/arousal, etc.). This is not a new dataset but a reference guide—please access original datasets via provided links and cite their sources.
Empathy and Affective Computing… See the full description on the dataset page: https://huggingface.co/datasets/Limorgu/empathy-affective-datasets.LIMO_QFFT
📘 LIMO–QFFT
LIMO–QFFT is a question-free variant of the original GAIR/LIMO dataset, tailored for use in QFFT (Question-Free Fine-Tuning) pipelines.
🔍 Description
This dataset removes the original input questions and system prompts from the LIMO dataset, and keeps only the long-form reasoning responses. The goal is to enable training large language models to learn from reasoning traces alone, without depending on task-specific questions.
All entries are converted into… See the full description on the dataset page: https://huggingface.co/datasets/lwl-uestc/LIMO_QFFT.LTXmrm8488__phi-4-14B-grpo-limo-details
Dataset Card for Evaluation run of mrm8488/phi-4-14B-grpo-limo
Dataset automatically created during the evaluation run of model mrm8488/phi-4-14B-grpo-limo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mrm8488__phi-4-14B-grpo-limo-details.LIMO-eval-datasetsLIMOlimo-indonesian-preferenceHO_LiMoNiTi_NPJCM_2020_LiMoNiTi_validation
Cite this dataset Cooper, A. M., Kästner, J., Urban, A., and Artrith, N. HO LiMoNiTi NPJCM 2020 LiMoNiTi validation. ColabFit, 2023. https://doi.org/10.60732/40d9f4e8
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_fgjil336jos6_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/HO_LiMoNiTi_NPJCM_2020_LiMoNiTi_validation.HO_LiMoNiTi_NPJCM_2020_water_clusters
Cite this dataset Cooper, A. M., Kästner, J., Urban, A., and Artrith, N. HO LiMoNiTi NPJCM 2020 water clusters. ColabFit, 2023. https://doi.org/10.60732/b633b325
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_or3nu4t64mvk_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/HO_LiMoNiTi_NPJCM_2020_water_clusters.limo-newlimo-cod此数据集为 limo 数据集的强化版本,将 cot => cod,其中 cod 指的是 Chain of Draft 提出的节约 token 的 cot。
files
limo-cod.jsonl 使用 deepseek-v3-0324 生成
limo-cod-g25p.jsonl 使用一个模型名称中包含 g-2-5-p 的模型生成
cot 字段来源于源数据集 GAIR/LIMO 采样于 deepseek-r1
refs
LIMO: https://huggingface.co/datasets/GAIR/LIMO https://arxiv.org/abs/2502.03387
COD: https://arxiv.org/abs/2502.18600
MixChain-C-LIMO
MixChian-C-LIMO
MixChain-C-LIMO contains two distinct solutions for each question from the LIMO dataset.
These solutions vary in the number of samples and the average length of their CoT.
Solution 1: 474 samples, avg. CoT length = 2994.7
Solution 2: 564 samples, avg. CoT length = 4890.6
Usage
To load the dataset using the 🤗 datasets library:
from datasets import load_dataset
ds = load_dataset("horseee/MixChain-C-LIMO", "solution_1") # Or solution_2
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/horseee/MixChain-C-LIMO.HO_LiMoNiTi_NPJCM_2020_bulk_water_train_test
Cite this dataset Cooper, A. M., Kästner, J., Urban, A., and Artrith, N. HO LiMoNiTi NPJCM 2020 bulk water train test. ColabFit, 2023. https://doi.org/10.60732/7f3ffd0b
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_jjiywlyfpnme_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/HO_LiMoNiTi_NPJCM_2020_bulk_water_train_test.limo-deepseek32b-responsesReflect_LIMOko-limo
Dataset Card for Ko-LIMO
Dataset Description
Ko-LIMO는 LIMO: Less is More for Reasoning의 학습 데이터를 한국어로 번역한 데이터셋입니다.
번역은 초벌 번역으로 DeepL를 활용하였고, 수학 용어들을 최대한 잘 표현할 수 있도록 용어집을 만들어 번역하였습니다. 용어집은 하단 토글에서 확인하실 수 있습니다.
수식이나 그림을 나타내는 특수문자 사이의 텍스트는 최대한 원문을 유지하는 형태로 번역을 진행하였으며, question, solution, answer 로 이루어진 총 817건의 데이터를 활용하실 수 있습니다.
데이터셋 관련하여 문의가 있으신 경우 메일을 통해 연락주세요! 🥰
차후 데이터 품질 검사 후 추가 업데이트가 있을 예정입니다.
This is Korean LIMO dataset, which is translated from the LIMO dataset.
DeepL… See the full description on the dataset page: https://huggingface.co/datasets/junnei/ko-limo.
