datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WISA-80K
WISA-80K
Dataset Description
WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation
Jing Wang*, Ao Ma*, Ke Cao*, Jun Zheng, Zhanjie Zhang, Jiasong Feng, Shanyuan Liu, Yuhang Ma, Bo Cheng, Dawei Leng‡, Yuhui Yin, Xiaodan Liang‡(*Equal Contribution, ‡Corresponding Authors)
BibTeX
@article{wang2025wisa,
title={WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation},
author={Wang, Jing and Ma, Ao and Cao… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/WISA-80K.Light-R1-SFTData
Light-R1: Surpassing R1-Distill from Scratch* with $1000 through Curriculum SFT & DPO
*from models without long COT
technical report
GitHub page
Here are the two-stage SFT data we used to train Light-R1-32B.
Simply refer to stage1-76k.json and stage2-3k.json
Model
Trained From
Release Date
AIME24
AIME25
DeepSeek-R1-Distill-Llama-70B
Llama-3.3-70B-Instruct
25.1.20
70.0
54.1
DeepSeek-R1-Distill-Qwen-32B
Qwen2.5-32B
25.1.20
72.6
54.9
LIMO (32B)
Qwen2.5-32B-Instruct
25.2.4… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/Light-R1-SFTData.InduOCRBenchInduOCRBench
English | 简体中文
[📜 arXiv] | [Dataset (🤗Hugging Face)]
News
[2026-04] InduOCRBench paper accepted to ACL 2026 Industry Track. Dataset released.
📖 Introduction
InduOCRBench is an OCR benchmark for industrial RAG systems, covering 11 challenging document types observed in real-world enterprise workflows. It addresses the gap between traditional character-level OCR metrics and actual downstream RAG utility, evaluating OCR robustness in terms of both… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/InduOCRBench.FineHARD-CN
FineHARD-CN
FineHARD-CN is a large-scale Chinese image-region grounding dataset.
annotations: 13,818,563 original annotation records.
image_urls: 13,829,272 image download URLs.
Images are not included and should be downloaded from the image_urls configuration.
Annotation field names and values are preserved from the original data.
See metadata/manifest.json for file checksums and release statistics.
Light-R1-DPOData
Light-R1: Surpassing R1-Distill from Scratch* with $1000 through Curriculum SFT & DPO
*from models without long COT
technical report
GitHub page
Here is the DPO data we used to train Light-R1-32B.
Simply refer to dpo-pairs.json
Model
Trained From
Release Date
AIME24
AIME25
DeepSeek-R1-Distill-Llama-70B
Llama-3.3-70B-Instruct
25.1.20
70.0
54.1
DeepSeek-R1-Distill-Qwen-32B
Qwen2.5-32B
25.1.20
72.6
54.9
LIMO (32B)
Qwen2.5-32B-Instruct25.2.4
56.3
47.1
s1.1-32B… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/Light-R1-DPOData.Light-IF-SFTData
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking
Here are the cold start data we used to train Light-IF-32B.
Simply refer to cold-start.json
🧪 Benchmarks
Model
SuperClue
IFEval
CFBench
IFBench
Qwen3-4B
0.225
0.888
0.787
0.382
Qwen3-8B
0.225
0.888
0.813
0.417
Qwen3-32B
0.234
0.877
0.823
0.384
Qwen3-235B-A22B
0.244
0.882
0.834
0.423
Qwen3-235B-A22B-Thinking-2507
0.434
0.916
0.843
0.475
DeepSeek-R1-0528… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/Light-IF-SFTData.
