datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SpatialForge
SpatialForge-10M
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
📑 Paper
Zishan Liu, Ruoxi Zang, Yanglin Zhang, Wei Liu, Yin Zhang, Jian Yao, Jiayin Zheng, Zhengzhe Liu
Lingnan University · XPENG Robotics
📦 SpatialForge-10M
A large-scale vision-language dataset designed for 3D-aware spatial perception and reasoning from open-world 2D images.
SpatialForge-10M contains over 10 million QA pairs generated from 2.8 million curated… See the full description on the dataset page: https://huggingface.co/datasets/shana643/SpatialForge.ai-detection-dataset-v2
---dataset_info:
features:
- name: image # use the exact column name from your parquet schema
dtype: image # this forces Hugging Face to render it as an image
- name: label
dtype: string
license: other
task_categories:
- image-classification
language:
- en
tags:
- ai-generated-image-detection
- synthetic-image-detection
- diffusion-models
pretty_name: AI-Generated Image Detection Dataset v2
size_categories:
- 10K<n<100K
AI-Generated… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/ai-detection-dataset-v2.SFT-Qwenaawaaz-transcript-cleanup-dataset
Aawaaz Transcript Cleanup Dataset
Training pairs for cleaning messy speech transcripts (ASR output, voice dictation) into well-formatted text while preserving the speaker's voice and meaning.
Dataset Description
Each example is a pair of:
input: A realistic messy transcript with filler words, false starts, self-corrections, grammar errors, and missing punctuation
output: The cleaned version with fillers removed, grammar fixed, punctuation added, and domain-appropriate… See the full description on the dataset page: https://huggingface.co/datasets/shantanugoel/aawaaz-transcript-cleanup-dataset.exdarkKAF-DatasetThe dataset sourced from https://github.com/IAAR-Shanghai/xFinder
Citation
@inproceedings{
xFinder,
title={xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation},
author={Qingchen Yu and Zifan Zheng and Shichao Song and Zhiyu li and Feiyu Xiong and Bo Tang and Ding Chen},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=7UqQJUKaLM}
}
ToxiRewriteCN
ToxiRewriteCN
ToxiRewriteCN is a Chinese toxic language mitigation dataset introduced in the EMNLP 2025 paper Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites. It is designed for detoxification and rewriting research where a toxic input is rewritten into a non-toxic sentence while preserving the original sentiment polarity and intent.
The dataset contains harmful, offensive, and disturbing language. It is released for research on safety, detoxification… See the full description on the dataset page: https://huggingface.co/datasets/shanewang/ToxiRewriteCN.Fence-Climbing-Action-Recognition-Dataset
Fence Climbing Action Recognition Dataset
The current security industry faces challenges from people climbing over walls, fences, and other security hazards. Traditional surveillance methods often cannot timely and effectively recognize these abnormal behaviors. Existing solutions are insufficient in the accuracy and real-time detection of actions, resulting in the inability to quickly respond to potential dangers. This dataset aims to support the training of action recognition… See the full description on the dataset page: https://huggingface.co/datasets/shangzx/Fence-Climbing-Action-Recognition-Dataset.math-reasoning-zh
Math Reasoning Chinese Dataset
Overview
A high-quality Chinese math word problem reasoning dataset with 500 problems featuring detailed Chain-of-Thought reasoning processes. All problems are algorithmically verified for 100% correctness.
Dataset Structure
Field
Description
Example
problem_id
Unique identifier
math_0001
question
Problem text
商店原价800元的商品打75折后...
chain_of_thought
Step-by-step reasoning
先算打折后价格:800 × 75% = 600元...… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/math-reasoning-zh.shangkhachil-bengali-public-domain
Bengali Public-Domain Literature
101 complete works by 21 authors,
11,250,629 characters. Corpus corpus-f8c532fcb4e7, built 2026-09-09.
Where these texts are read
https://shangkhachil.com — the reading site this corpus was built for. Free, no
account, 246 works by 28 authors. The complete text of
every work in this file can be read there.
This file is the text. The site is the part a JSONL cannot be:
Rights computed for the reader's own country, at the edge… See the full description on the dataset page: https://huggingface.co/datasets/mir178/shangkhachil-bengali-public-domain.programming-interview-zh-extended
Programming Interview Dataset (Chinese Extended) - 编程面试数据集(扩展版)
Overview
An EXTENDED version of the Chinese programming interview question dataset with 2000 problems featuring detailed solutions in Python, Java, and C++, complexity analysis, and key insights. Designed for LLM training in coding assistance and technical interview preparation.
Dataset Structure
Field
Description
problem_id
Unique identifier
original_id
Original problem ID… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/programming-interview-zh-extended.Qwen3.8-2.4T-A95B-responses-original10k
Qwen3.8-2.4T-A95B responses — original aligned 10k
The first 10,000 exact rows from the private source dataset inference-optimization/Qwen3.8-2.4T-A95B-responses. Records are preserved without modification.
The 10,000 rows are aligned by id and primary_id with the companion dataset.
The source revision is 12750d033529d53fed1e29b1d9734e8bd76b73e5.
programming-interview-zh-ultimate
Programming Interview Dataset (Chinese Ultimate) - 编程面试数据集(终极版)
Overview
The ULTIMATE version of the Chinese programming interview question dataset with 5000 problems featuring detailed solutions in Python, Java, and C++, complexity analysis, common mistakes, and key insights. Designed for LLM training in coding assistance and technical interview preparation.
Dataset Structure
Field
Description
problem_id
Unique identifier (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/programming-interview-zh-ultimate.programming-interview-zh
Programming Interview Dataset (Chinese) - 编程面试数据集
Overview
A high-quality Chinese programming interview question dataset with 500 problems featuring detailed solutions, complexity analysis, and key insights. Designed for LLM training in coding assistance and technical interview preparation.
Dataset Structure
Field
Description
problem_id
Unique identifier
title
Problem title (Chinese)
category
Problem type… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/programming-interview-zh.Qwen3.8-27B-responses-regenerated10k
Qwen3.8-27B regenerated responses — aligned 10k
The same 10,000 original prompts regenerated with dense Qwen/Qwen3.8-27B. Original prompt strings were used directly; they were never reconstructed by detokenization.
The 10,000 rows are aligned by id and primary_id with the companion dataset.
The source revision is 12750d033529d53fed1e29b1d9734e8bd76b73e5.
Aegis-AI-Content-Safety-Dataset-2.0
🛡️ Nemotron Content Safety Dataset V2
The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1.
To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/shannifnju/Aegis-AI-Content-Safety-Dataset-2.0.website_metadata_c4_toyA smaller version (100 samples) of https://huggingface.co/datasets/bs-modeling-metadata/website_metadata_c4
qwen3_5_4b_perfectblend_regends1000-shit-match-charity-shandong-aid-policy
学生资助政策问答数据集
数据集简介
本数据集整理学生资助高频咨询问题与官方政策答复,用于RAG检索增强、大模型垂直领域知识库建设。
全部内容来自公开政务政策,已完成隐私脱敏,无个人敏感信息。
文件列表
student_aid_qa.jsonl:结构化问答对
policy_docs/*.md:政策Markdown原文
许可协议
License: CC‑BY‑4.0
允许研究、商业使用,使用时请标注来源。
免责声明
AI生成结果仅供参考,资助办理请以当地教育主管部门正式文件为准。
shan-blogspots
Language
Shan - shn
TeaMs-RL-9kTeaMs-RL: Teaching LLMs to Generate Better Instruction Datasets via Reinforcement Learning
Run experiments
Install:
pip install -r requirements.txt
pip install -e .
run experiments / train models
cd Teams_RL_GPT/teams_rl/runner/
sh run_llm_rl.sh
If you met some issues, please check the existing solutions for the reported issues, which could help you address your issue.
We also provide the datasets that we used to train the models.
After collected datasets, use train_models.sh… See the full description on the dataset page: https://huggingface.co/datasets/Shangding-Gu/TeaMs-RL-9k.authori-prospector-lexicon
AuthoriProspector AEO Lexicon Dataset
Authoritative term definitions published by AuthoriProspector -- structured for AI answer engine consumption.
Schema
Field
Type
Description
term
string
The defined term
law_definition
string
Definition
lore_definition
string
Context
aura_score
integer
Authority score
source_url
string
AEO term page URL
canonical_url
string
Canonical home for this term
Query with DuckDB
SELECT term… See the full description on the dataset page: https://huggingface.co/datasets/shannonbox1999/authori-prospector-lexicon.physr1corp-cold-start
PhysR1Corp Cold-Start — Tool-Using Physics Trajectories (Text + Multimodal)
SFT cold-start dataset for Physics-R2 / Phase E (a tool-use RL paper for VLMs, target ICLR 2027).
1,973 audited trajectories — 1,740 text-only + 233 multimodal — covering the full PhysR1Corp
corpus (all 2,268 problems). Each trajectory solves a physics problem using a custom SymPy tool
routed through the project's harness.sandbox runtime. The trajectory schema matches the W6
RL-training inference format… See the full description on the dataset page: https://huggingface.co/datasets/shanyangmie/physr1corp-cold-start.shan-wordpress
Language
Shan - shn
rag_finetunesggs-bench
☬ SGGS-Bench v0.1 — A Benchmark for Sri Guru Granth Sahib AI Systems
The first comprehensive evaluation framework for AI systems that interpret Sikh scripture.
115 questions · 8 task dimensions · Hybrid automated + LLM-judge scoring
Factual · Retrieval · Exegesis · Guidance · Hallucination · Theology · Cross-Reference · Safety
📋 Overview
SGGS-Bench is an 8-task, 115-question evaluation framework designed to measure AI competence on the Sri Guru Granth Sahib… See the full description on the dataset page: https://huggingface.co/datasets/ShanvirDhinsa/sggs-bench.workload-tab-lexicon
WorkLoad Tab AEO Lexicon Dataset
Authoritative term definitions published by WorkLoad Tab -- structured for AI answer engine consumption.
Schema
Field
Type
Description
term
string
The defined term
law_definition
string
Product Spec
lore_definition
string
User Voice
aura_score
integer
Authority score
source_url
string
AEO term page URL
canonical_url
string
Canonical home for this term
Query with DuckDB
SELECT term… See the full description on the dataset page: https://huggingface.co/datasets/shannonbox1999/workload-tab-lexicon.Hazardous-Materials-Vehicle-Illegal-Parking-Behavior-Detection-Dataset
Hazardous Materials Vehicle Illegal Parking Behavior Detection Dataset
The current transportation industry faces the issue of frequent illegal parking behaviors of hazardous materials vehicles, which not only affects traffic order but also increases safety hazards. However, existing monitoring systems often cannot accurately identify and judge the illegal behavior of hazardous materials vehicles, making it difficult to take effective measures. This dataset aims to provide… See the full description on the dataset page: https://huggingface.co/datasets/shangzx/Hazardous-Materials-Vehicle-Illegal-Parking-Behavior-Detection-Dataset.oumi_rag_grpo_data
