datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WISA-80K
WISA-80K
Dataset Description
WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation
Jing Wang*, Ao Ma*, Ke Cao*, Jun Zheng, Zhanjie Zhang, Jiasong Feng, Shanyuan Liu, Yuhang Ma, Bo Cheng, Dawei Leng‡, Yuhui Yin, Xiaodan Liang‡(*Equal Contribution, ‡Corresponding Authors)
BibTeX
@article{wang2025wisa,
title={WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation},
author={Wang, Jing and Ma, Ao and Cao… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/WISA-80K.RevealLayer-100K
RevealLayer Open Dataset
RevealLayer Open is the open-source dataset accompanying RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition.
Paper: https://arxiv.org/html/2605.11818v1 Accepted by ICML 2026
RevealLayer studies box-guided layered image decomposition for natural images. Given an RGB image and instance bounding boxes, the task is to decompose the scene into a clean background and object-level foreground layers, where each… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/RevealLayer-100K.Light-R1-SFTData
Light-R1: Surpassing R1-Distill from Scratch* with $1000 through Curriculum SFT & DPO
*from models without long COT
technical report
GitHub page
Here are the two-stage SFT data we used to train Light-R1-32B.
Simply refer to stage1-76k.json and stage2-3k.json
Model
Trained From
Release Date
AIME24
AIME25
DeepSeek-R1-Distill-Llama-70B
Llama-3.3-70B-Instruct
25.1.20
70.0
54.1
DeepSeek-R1-Distill-Qwen-32B
Qwen2.5-32B
25.1.20
72.6
54.9
LIMO (32B)
Qwen2.5-32B-Instruct
25.2.4… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/Light-R1-SFTData.InduOCRBenchInduOCRBench
English | 简体中文
[📜 arXiv] | [Dataset (🤗Hugging Face)]
News
[2026-04] InduOCRBench paper accepted to ACL 2026 Industry Track. Dataset released.
📖 Introduction
InduOCRBench is an OCR benchmark for industrial RAG systems, covering 11 challenging document types observed in real-world enterprise workflows. It addresses the gap between traditional character-level OCR metrics and actual downstream RAG utility, evaluating OCR robustness in terms of both… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/InduOCRBench.FineHARD-CN
FineHARD-CN
FineHARD-CN is a large-scale Chinese image-region grounding dataset.
annotations: 13,818,563 original annotation records.
image_urls: 13,829,272 image download URLs.
Images are not included and should be downloaded from the image_urls configuration.
Annotation field names and values are preserved from the original data.
See metadata/manifest.json for file checksums and release statistics.
FineHARD
FG-CLIP: Fine-Grained Visual and Textual Alignment
FG-CLIP: Fine-Grained Visual and Textual Alignment
Chunyu Xie*, Bin Wang*, Fanjing Kong, Jincheng Li, Dawei Liang, Gengshen Zhang, Dawei Leng†, Yuhui Yin(*Equal Contribution, ✝Corresponding Author)
Model Framework
FG-CLIP’s training proceeds in two stages: the first stage leverages
global-level caption-image pairs to achieve initial fine-grained alignment, while the second stage supplements these with… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/FineHARD.DCI-CN
FG-CLIP 2: A Bilingual Fine-grained Vision-language Alignment Model
Code: https://github.com/360CVGroup/FG-CLIP
FG-CLIP 2 is the foundation model for fine-grained vision-language understanding in both English and Chinese.
Across 29 datasets and 8 diverse tasks, it consistently surpasses recent strong baselines such as SigLIP 2 and MetaCLIP 2, achieving the best reported performance to date in both languages.
FG-CLIP 2: A Bilingual Fine-grained Vision-language Alignment Model… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/DCI-CN.DOCCI-CN
FG-CLIP 2: A Bilingual Fine-grained Vision-language Alignment Model
Code: https://github.com/360CVGroup/FG-CLIP
FG-CLIP 2 is the foundation model for fine-grained vision-language understanding in both English and Chinese.
Across 29 datasets and 8 diverse tasks, it consistently surpasses recent strong baselines such as SigLIP 2 and MetaCLIP 2, achieving the best reported performance to date in both languages.
FG-CLIP 2: A Bilingual Fine-grained Vision-language Alignment Model… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/DOCCI-CN.details_qihoo360__TinyR1-32B-Preview_v2
Dataset Card for Evaluation run of qihoo360/TinyR1-32B-Preview
Dataset automatically created during the evaluation run of model qihoo360/TinyR1-32B-Preview.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_qihoo360__TinyR1-32B-Preview_v2.BoxClass-CN
FG-CLIP 2: A Bilingual Fine-grained Vision-language Alignment Model
Code: https://github.com/360CVGroup/FG-CLIP
FG-CLIP 2 is the foundation model for fine-grained vision-language understanding in both English and Chinese.
Across 29 datasets and 8 diverse tasks, it consistently surpasses recent strong baselines such as SigLIP 2 and MetaCLIP 2, achieving the best reported performance to date in both languages.
FG-CLIP 2: A Bilingual Fine-grained Vision-language Alignment Model… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/BoxClass-CN.LIT-CN
FG-CLIP 2: A Bilingual Fine-grained Vision-language Alignment Model
Code: https://github.com/360CVGroup/FG-CLIP
FG-CLIP 2 is the foundation model for fine-grained vision-language understanding in both English and Chinese.
Across 29 datasets and 8 diverse tasks, it consistently surpasses recent strong baselines such as SigLIP 2 and MetaCLIP 2, achieving the best reported performance to date in both languages.
FG-CLIP 2: A Bilingual Fine-grained Vision-language Alignment Model… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/LIT-CN.VRF-datasetsLight-R1-DPOData
Light-R1: Surpassing R1-Distill from Scratch* with $1000 through Curriculum SFT & DPO
*from models without long COT
technical report
GitHub page
Here is the DPO data we used to train Light-R1-32B.
Simply refer to dpo-pairs.json
Model
Trained From
Release Date
AIME24
AIME25
DeepSeek-R1-Distill-Llama-70B
Llama-3.3-70B-Instruct
25.1.20
70.0
54.1
DeepSeek-R1-Distill-Qwen-32B
Qwen2.5-32B
25.1.20
72.6
54.9
LIMO (32B)
Qwen2.5-32B-Instruct25.2.4
56.3
47.1
s1.1-32B… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/Light-R1-DPOData.Light-IF-SFTData
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking
Here are the cold start data we used to train Light-IF-32B.
Simply refer to cold-start.json
🧪 Benchmarks
Model
SuperClue
IFEval
CFBench
IFBench
Qwen3-4B
0.225
0.888
0.787
0.382
Qwen3-8B
0.225
0.888
0.813
0.417
Qwen3-32B
0.234
0.877
0.823
0.384
Qwen3-235B-A22B
0.244
0.882
0.834
0.423
Qwen3-235B-A22B-Thinking-2507
0.434
0.916
0.843
0.475
DeepSeek-R1-0528… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/Light-IF-SFTData.TinyR1-32B-Preview-datasetsPlanGen_dataLight-IF-cold-start-datasets.jsonHere are the cold start SFT data we used to train Light-IF-32B.
Simply refer to Light-IF-cold-start-datasets.json
