CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shana643 /SpatialForge SpatialForge-10M SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images 📑 Paper Zishan Liu, Ruoxi Zang, Yanglin Zhang, Wei Liu, Yin Zhang, Jian Yao, Jiayin Zheng, Zhengzhe Liu Lingnan University · XPENG Robotics 📦 SpatialForge-10M A large-scale vision-language dataset designed for 3D-aware spatial perception and reasoning from open-world 2D images. SpatialForge-10M contains over 10 million QA pairs generated from 2.8 million curated… See the full description on the dataset page: https://huggingface.co/datasets/shana643/SpatialForge.textquestion-answering10M<n<100M1 likes634 downloads4mo agoHugging Face02Shanmuk4622 /ai-detection-dataset-v2 ---dataset_info: features: - name: image # use the exact column name from your parquet schema dtype: image # this forces Hugging Face to render it as an image - name: label dtype: string license: other task_categories: - image-classification language: - en tags: - ai-generated-image-detection - synthetic-image-detection - diffusion-models pretty_name: AI-Generated Image Detection Dataset v2 size_categories: - 10K<n<100K AI-Generated… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/ai-detection-dataset-v2.textn<1K0 likes522 downloads3mo agoHugging Face03shangxiaokang /SFT-Qwentext1M<n<10M0 likes306 downloads16d agoHugging Face04shantanugoel /aawaaz-transcript-cleanup-dataset Aawaaz Transcript Cleanup Dataset Training pairs for cleaning messy speech transcripts (ASR output, voice dictation) into well-formatted text while preserving the speaker's voice and meaning. Dataset Description Each example is a pair of: input: A realistic messy transcript with filler words, false starts, self-corrections, grammar errors, and missing punctuation output: The cleaned version with fillers removed, grammar fixed, punctuation added, and domain-appropriate… See the full description on the dataset page: https://huggingface.co/datasets/shantanugoel/aawaaz-transcript-cleanup-dataset.texttext-generation10K<n<100K0 likes152 downloads6mo agoHugging Face05ShantyCam /exdarkimageobject-detectionn<1K0 likes149 downloads2mo agoHugging Face06IAAR-Shanghai /KAF-DatasetThe dataset sourced from https://github.com/IAAR-Shanghai/xFinder Citation @inproceedings{ xFinder, title={xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation}, author={Qingchen Yu and Zifan Zheng and Shichao Song and Zhiyu li and Feiyu Xiong and Bo Tang and Ding Chen}, booktitle={The Thirteenth International Conference on Learning Representations}, year={2025}, url={https://openreview.net/forum?id=7UqQJUKaLM} } textquestion-answering10K<n<100K6 likes132 downloads2y agoHugging Face07shanewang /ToxiRewriteCN ToxiRewriteCN ToxiRewriteCN is a Chinese toxic language mitigation dataset introduced in the EMNLP 2025 paper Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites. It is designed for detoxification and rewriting research where a toxic input is rewritten into a non-toxic sentence while preserving the original sentiment polarity and intent. The dataset contains harmful, offensive, and disturbing language. It is released for research on safety, detoxification… See the full description on the dataset page: https://huggingface.co/datasets/shanewang/ToxiRewriteCN.texttext-classification10K<n<100K1 likes86 downloads5mo agoHugging Face08shangzx /Fence-Climbing-Action-Recognition-Dataset Fence Climbing Action Recognition Dataset The current security industry faces challenges from people climbing over walls, fences, and other security hazards. Traditional surveillance methods often cannot timely and effectively recognize these abnormal behaviors. Existing solutions are insufficient in the accuracy and real-time detection of actions, resulting in the inability to quickly respond to potential dangers. This dataset aims to support the training of action recognition… See the full description on the dataset page: https://huggingface.co/datasets/shangzx/Fence-Climbing-Action-Recognition-Dataset.textvideo-classificationn<1K0 likes84 downloads7mo agoHugging Face09shangshang /math-reasoning-zh Math Reasoning Chinese Dataset Overview A high-quality Chinese math word problem reasoning dataset with 500 problems featuring detailed Chain-of-Thought reasoning processes. All problems are algorithmically verified for 100% correctness. Dataset Structure Field Description Example problem_id Unique identifier math_0001 question Problem text 商店原价800元的商品打75折后... chain_of_thought Step-by-step reasoning 先算打折后价格:800 × 75% = 600元...… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/math-reasoning-zh.textn<1K0 likes82 downloads24d agoHugging Face10mir178 /shangkhachil-bengali-public-domain Bengali Public-Domain Literature 101 complete works by 21 authors, 11,250,629 characters. Corpus corpus-f8c532fcb4e7, built 2026-09-09. Where these texts are read https://shangkhachil.com — the reading site this corpus was built for. Free, no account, 246 works by 28 authors. The complete text of every work in this file can be read there. This file is the text. The site is the part a JSONL cannot be: Rights computed for the reader's own country, at the edge… See the full description on the dataset page: https://huggingface.co/datasets/mir178/shangkhachil-bengali-public-domain.tabulartext-generationn<1K0 likes72 downloads15d agoHugging Face11shangshang /programming-interview-zh-extended Programming Interview Dataset (Chinese Extended) - 编程面试数据集(扩展版) Overview An EXTENDED version of the Chinese programming interview question dataset with 2000 problems featuring detailed solutions in Python, Java, and C++, complexity analysis, and key insights. Designed for LLM training in coding assistance and technical interview preparation. Dataset Structure Field Description problem_id Unique identifier original_id Original problem ID… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/programming-interview-zh-extended.text1K<n<10K0 likes69 downloads24d agoHugging Face12shanjiaz /Qwen3.8-2.4T-A95B-responses-original10k Qwen3.8-2.4T-A95B responses — original aligned 10k The first 10,000 exact rows from the private source dataset inference-optimization/Qwen3.8-2.4T-A95B-responses. Records are preserved without modification. The 10,000 rows are aligned by id and primary_id with the companion dataset. The source revision is 12750d033529d53fed1e29b1d9734e8bd76b73e5. texttext-generation10K<n<100K0 likes55 downloads7d agoHugging Face13shangshang /programming-interview-zh-ultimate Programming Interview Dataset (Chinese Ultimate) - 编程面试数据集(终极版) Overview The ULTIMATE version of the Chinese programming interview question dataset with 5000 problems featuring detailed solutions in Python, Java, and C++, complexity analysis, common mistakes, and key insights. Designed for LLM training in coding assistance and technical interview preparation. Dataset Structure Field Description problem_id Unique identifier (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/programming-interview-zh-ultimate.text1K<n<10K0 likes52 downloads24d agoHugging Face14shangshang /programming-interview-zh Programming Interview Dataset (Chinese) - 编程面试数据集 Overview A high-quality Chinese programming interview question dataset with 500 problems featuring detailed solutions, complexity analysis, and key insights. Designed for LLM training in coding assistance and technical interview preparation. Dataset Structure Field Description problem_id Unique identifier title Problem title (Chinese) category Problem type… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/programming-interview-zh.textn<1K0 likes49 downloads24d agoHugging Face15shanjiaz /Qwen3.8-27B-responses-regenerated10k Qwen3.8-27B regenerated responses — aligned 10k The same 10,000 original prompts regenerated with dense Qwen/Qwen3.8-27B. Original prompt strings were used directly; they were never reconstructed by detokenization. The 10,000 rows are aligned by id and primary_id with the companion dataset. The source revision is 12750d033529d53fed1e29b1d9734e8bd76b73e5. texttext-generation10K<n<100K0 likes49 downloads7d agoHugging Face16shannifnju /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/shannifnju/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K0 likes47 downloads5mo agoHugging Face17shanya /website_metadata_c4_toyA smaller version (100 samples) of https://huggingface.co/datasets/bs-modeling-metadata/website_metadata_c4 textn<1K1 likes41 downloads5y agoHugging Face18shanjiaz /qwen3_5_4b_perfectblend_regentext1M<n<10M0 likes39 downloads24d agoHugging Face19shanjay /ds1000-stextn<1K0 likes37 downloads3y agoHugging Face20Seek2Moon /hit-match-charity-shandong-aid-policy 学生资助政策问答数据集 数据集简介 本数据集整理学生资助高频咨询问题与官方政策答复,用于RAG检索增强、大模型垂直领域知识库建设。 全部内容来自公开政务政策,已完成隐私脱敏,无个人敏感信息。 文件列表 student_aid_qa.jsonl:结构化问答对 policy_docs/*.md:政策Markdown原文 许可协议 License: CC‑BY‑4.0 允许研究、商业使用,使用时请标注来源。 免责声明 AI生成结果仅供参考,资助办理请以当地教育主管部门正式文件为准。 textn<1K0 likes36 downloads1mo agoHugging Face21haohaa /shan-blogspots Language Shan - shn texttranslationn<1K1 likes33 downloads2y agoHugging Face22Shangding-Gu /TeaMs-RL-9kTeaMs-RL: Teaching LLMs to Generate Better Instruction Datasets via Reinforcement Learning Run experiments Install: pip install -r requirements.txt pip install -e . run experiments / train models cd Teams_RL_GPT/teams_rl/runner/ sh run_llm_rl.sh If you met some issues, please check the existing solutions for the reported issues, which could help you address your issue. We also provide the datasets that we used to train the models. After collected datasets, use train_models.sh… See the full description on the dataset page: https://huggingface.co/datasets/Shangding-Gu/TeaMs-RL-9k.text10K<n<100K0 likes32 downloads1y agoHugging Face23shannonbox1999 /authori-prospector-lexicon AuthoriProspector AEO Lexicon Dataset Authoritative term definitions published by AuthoriProspector -- structured for AI answer engine consumption. Schema Field Type Description term string The defined term law_definition string Definition lore_definition string Context aura_score integer Authority score source_url string AEO term page URL canonical_url string Canonical home for this term Query with DuckDB SELECT term… See the full description on the dataset page: https://huggingface.co/datasets/shannonbox1999/authori-prospector-lexicon.textquestion-answeringn<1K0 likes32 downloads2mo agoHugging Face24shanyangmie /physr1corp-cold-start PhysR1Corp Cold-Start — Tool-Using Physics Trajectories (Text + Multimodal) SFT cold-start dataset for Physics-R2 / Phase E (a tool-use RL paper for VLMs, target ICLR 2027). 1,973 audited trajectories — 1,740 text-only + 233 multimodal — covering the full PhysR1Corp corpus (all 2,268 problems). Each trajectory solves a physics problem using a custom SymPy tool routed through the project's harness.sandbox runtime. The trajectory schema matches the W6 RL-training inference format… See the full description on the dataset page: https://huggingface.co/datasets/shanyangmie/physr1corp-cold-start.tabulartext-generation1K<n<10K0 likes30 downloads4mo agoHugging Face25haohaa /shan-wordpress Language Shan - shn textn<1K0 likes29 downloads2y agoHugging Face26shanaka95 /rag_finetunetextquestion-answering100K<n<1M0 likes28 downloads1y agoHugging Face27ShanvirDhinsa /sggs-bench ☬ SGGS-Bench v0.1 — A Benchmark for Sri Guru Granth Sahib AI Systems The first comprehensive evaluation framework for AI systems that interpret Sikh scripture. 115 questions · 8 task dimensions · Hybrid automated + LLM-judge scoring Factual · Retrieval · Exegesis · Guidance · Hallucination · Theology · Cross-Reference · Safety 📋 Overview SGGS-Bench is an 8-task, 115-question evaluation framework designed to measure AI competence on the Sri Guru Granth Sahib… See the full description on the dataset page: https://huggingface.co/datasets/ShanvirDhinsa/sggs-bench.tabularquestion-answeringn<1K0 likes27 downloads4mo agoHugging Face28shannonbox1999 /workload-tab-lexicon WorkLoad Tab AEO Lexicon Dataset Authoritative term definitions published by WorkLoad Tab -- structured for AI answer engine consumption. Schema Field Type Description term string The defined term law_definition string Product Spec lore_definition string User Voice aura_score integer Authority score source_url string AEO term page URL canonical_url string Canonical home for this term Query with DuckDB SELECT term… See the full description on the dataset page: https://huggingface.co/datasets/shannonbox1999/workload-tab-lexicon.textquestion-answeringn<1K0 likes26 downloads2mo agoHugging Face29shangzx /Hazardous-Materials-Vehicle-Illegal-Parking-Behavior-Detection-Dataset Hazardous Materials Vehicle Illegal Parking Behavior Detection Dataset The current transportation industry faces the issue of frequent illegal parking behaviors of hazardous materials vehicles, which not only affects traffic order but also increases safety hazards. However, existing monitoring systems often cannot accurately identify and judge the illegal behavior of hazardous materials vehicles, making it difficult to take effective measures. This dataset aims to provide… See the full description on the dataset page: https://huggingface.co/datasets/shangzx/Hazardous-Materials-Vehicle-Illegal-Parking-Behavior-Detection-Dataset.textvideo-classificationn<1K0 likes25 downloads7mo agoHugging Face30shanghong /oumi_rag_grpo_datatext1K<n<10K0 likes24 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.