CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Shanmuk4622 /jeb-rag JEB-Bench Charging the Gate Rent: Measured-Energy Accounting for Adaptive Retrieval-Augmented Generation ⚠️ Status: under construction. Phase 0 (measurement validation) and Phase 1 (index construction) are landing now. The oracle matrix (bench/oracle/) is populated in Phase 2 and this card will be revised when it is complete. Do not cite numbers from this repository until the status line says complete. What this is The first public per-query × per-configuration… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/jeb-rag.tabularquestion-answering10K<n<100K2 likes1.9k downloads1mo agoHugging Face02shangrilar /ko_text2sqltext10K<n<100K19 likes1.6k downloads3y agoHugging Face03Shanmuk4622 /E2AM_ResNet50 E2AM Ablation Results: ResNet-50 Energy-aware training ablation study for ResNet-50 across three image-classification datasets: CIFAR-10, CIFAR-100, and Tiny-ImageNet. Each dataset has 15 training variants (8 individual-method M0..M7, 7 cumulative ablation C0..C6) at 50 epochs, plus a 5-variant deployment pipeline (FP32 baseline, structured pruning, pruning+finetune, INT8 quantization, pruned+INT8). Status: 45 completed variants, 0 partial. Quick links… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/E2AM_ResNet50.imagen<1K2 likes1.1k downloads3mo agoHugging Face04shanxianzheng /SWE-smith-jstext1K<n<10K0 likes859 downloads2mo agoHugging Face05shana643 /SpatialForge SpatialForge-10M SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images 📑 Paper Zishan Liu, Ruoxi Zang, Yanglin Zhang, Wei Liu, Yin Zhang, Jian Yao, Jiayin Zheng, Zhengzhe Liu Lingnan University · XPENG Robotics 📦 SpatialForge-10M A large-scale vision-language dataset designed for 3D-aware spatial perception and reasoning from open-world 2D images. SpatialForge-10M contains over 10 million QA pairs generated from 2.8 million curated… See the full description on the dataset page: https://huggingface.co/datasets/shana643/SpatialForge.textquestion-answering10M<n<100M1 likes817 downloads4mo agoHugging Face06BangumiBase /shangrilafrontier Bangumi Image Base of Shangri-la Frontier This is the image base of bangumi Shangri-La Frontier, we detected 48 characters, 2678 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/shangrilafrontier.image1K<n<10K0 likes735 downloads3y agoHugging Face07Shanmuk4622 /ai-detection-dataset-v2 ---dataset_info: features: - name: image # use the exact column name from your parquet schema dtype: image # this forces Hugging Face to render it as an image - name: label dtype: string license: other task_categories: - image-classification language: - en tags: - ai-generated-image-detection - synthetic-image-detection - diffusion-models pretty_name: AI-Generated Image Detection Dataset v2 size_categories: - 10K<n<100K AI-Generated… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/ai-detection-dataset-v2.textn<1K0 likes516 downloads3mo agoHugging Face08shanxianzheng /SWE-smith-cpptext1K<n<10K0 likes393 downloads2mo agoHugging Face09shanchen /aime_2025_multilingualWhen Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy https://arxiv.org/abs/2505.22888 Jirui Qi, Shan Chen, Zidi Xiong, Raquel Fernández, Danielle S. Bitterman, Arianna Bisazza Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studied. This capability is as important as answer accuracy for real world applications because… See the full description on the dataset page: https://huggingface.co/datasets/shanchen/aime_2025_multilingual.tabularn<1K0 likes390 downloads1y agoHugging Face10shanxianzheng /SWE-smith-tstext1K<n<10K0 likes381 downloads2mo agoHugging Face11shanchen /gpqa_diamond_mc_multilingualWhen Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy https://arxiv.org/abs/2505.22888 Jirui Qi, Shan Chen, Zidi Xiong, Raquel Fernández, Danielle S. Bitterman, Arianna Bisazza Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studied. This capability is as important as answer accuracy for real world applications because… See the full description on the dataset page: https://huggingface.co/datasets/shanchen/gpqa_diamond_mc_multilingual.text1K<n<10K2 likes368 downloads1y agoHugging Face12shanxianzheng /SWE-smith-javatext1K<n<10K0 likes352 downloads2mo agoHugging Face13ShantyCam /kitti-objectimage10K<n<100K0 likes337 downloads2mo agoHugging Face14Shanmuk4622 /ai-image-detection-datasetgated AI-Image Detection Dataset Paired real / AI images, with shared image-grounded captions, for training and evaluating AI-generated-image detectors. Each of 10,000 real photos is captioned once (BLIP-2) and paired with one synthetic partner per generator (6 generators → 60,000 AI images). A real image and all of its AI partners share the same prompt, so the only systematic difference between the classes is the generative process itself. A detector trained here is pushed toward the… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/ai-image-detection-dataset.imageimage-classification100K<n<1M4 likes324 downloads3mo agoHugging Face15shanthini33 /ASVspoof5 ASVspoof 5 (track 1, eval) Benchmark-ready packaging of the Track 1 (spoofing / deepfake detection) evaluation partition of the ASVspoof 5 challenge, for speech anti-spoofing and synthetic / deepfake voice detection. Overview Track 1 is binary classification: bonafide (genuine human speech) vs. spoof (synthetic / converted speech). This packaging contains the full track_1 evaluation set. The original challenge is at https://www.asvspoof.org/. License &… See the full description on the dataset page: https://huggingface.co/datasets/shanthini33/ASVspoof5.audioaudio-classification100K<n<1M0 likes319 downloads16d agoHugging Face16shanxianzheng /SWE-smith-gotext1K<n<10K0 likes306 downloads2mo agoHugging Face17shangzhu /ChemQA Dataset Card for ChemQA Introducing ChemQA: a Multimodal Question-and-Answering Dataset on Chemistry Reasoning. This work is inspired by IsoBench[1] and ChemLLMBench[2]. Content There are 5 QA Tasks in total: Counting Numbers of Carbons and Hydrogens in Organic Molecules: adapted from the 600 PubChem molecules created from [2], evenly divided into validation and evaluation datasets. Calculating Molecular Weights in Organic Molecules: adapted from the 600 PubChem… See the full description on the dataset page: https://huggingface.co/datasets/shangzhu/ChemQA.image10K<n<100K10 likes298 downloads2y agoHugging Face18shannonnonshan /SpeechTextMatching_Tedlium2Train Dataset Card for "SpeechTextMatching_TEDLIUM2Train" More Information needed audio10K<n<100K0 likes293 downloads7mo agoHugging Face19shangxiaokang /SFT-Qwentext1M<n<10M0 likes279 downloads14d agoHugging Face20jinaai /shanghai_master_plan_beirThis is a copy of https://huggingface.co/datasets/jinaai/shanghai_master_plan reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/shanghai_master_plan_beir.imagen<1K0 likes268 downloads1y agoHugging Face21shangeth /expresso Expresso — audio + text A faithful re-publication of the official Expresso dataset (Nguyen et al., Interspeech 2023) as a loadable HuggingFace audio dataset, sourced directly from FAIR's official tar. ⚠️ License: CC-BY-NC-4.0 — non-commercial use only. Configs read — 11.6k mono read-speech utterances with human transcripts. conversational — ~15.9k mono per-utterance turns derived from the stereo conversational dialogues, transcribed with Whisper Large V3 Turbo.… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/expresso.audiotext-to-speech10K<n<100K0 likes262 downloads5mo agoHugging Face22shanxianzheng /SWE-smith-pytext1K<n<10K0 likes249 downloads2mo agoHugging Face23shantezhou /pokemon_data Pokémon TCG Tournament Game Logs Full game logs from an agent Pokémon TCG tournament ("Limited Card Battle"), 122,542 recorded games collected across 26 daily archives — an initial batch of 15 undated archives (top/, medium_high/) plus dated days 2026-07-21 → 2026-07-31 (days_0721_0731/). Both batches use the same per-day split rule and load together via the split config above. Splits Games in each daily archive were ranked by avg_score (per-agent average rating… See the full description on the dataset page: https://huggingface.co/datasets/shantezhou/pokemon_data.text100K<n<1M0 likes241 downloads2mo agoHugging Face24KTAEHWA /shanghaitech-crowd-countingimage1K<n<10K1 likes232 downloads10mo agoHugging Face25ShantanuT01 /DACTYL DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large language models Dataset The DACTYL dataset is an AI-generated text detection dataset focusing primarily on one-shot or few-shot examples. We also include texts from continued pre-trained small language models. For more information, refer to our paper. Models Used We used the following LLMs to generate texts. OpenAI’s GPT-4o-mini and GPT-4o Anthropic’s Claude Haiku and Sonnet 3.5 Mistral Small (24B)and… See the full description on the dataset page: https://huggingface.co/datasets/ShantanuT01/DACTYL.tabulartext-classification100K<n<1M0 likes201 downloads1y agoHugging Face26shannonnonshan /ViMedCSS-Cop 🩺 ViMedCSS: A Vietnamese Medical Code-Switching Speech Dataset (LREC 2026) 📖 Overview ViMedCSS is a Vietnamese medical speech dataset for code-switching ASR, where each utterance contains at least one non-Vietnamese (mainly English) medical term embedded in Vietnamese speech. 📊 Dataset Statistics Split Statistics (from ViMedCSS-Metadata) Split # Rows Duration (hours) Avg duration (s) Total CS terms train 11,832 24.30 7.39 12,314… See the full description on the dataset page: https://huggingface.co/datasets/shannonnonshan/ViMedCSS-Cop.audioautomatic-speech-recognition10K<n<100K1 likes198 downloads7mo agoHugging Face27TingChen-ppmc /Shanghai_Dialect_Conversational_Speech_Corpus Corpus This dataset is built from Magicdata ASR-CZDIACSC: A CHINESE SHANGHAI DIALECT CONVERSATIONAL SPEECH CORPUS This corpus is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. Please refer to the license for further information. Modifications: The audio is split in sentences based on the time span on the transcription file. Sentences that span less than 1 second is discarded. Topics of conversation is removed. Usage… See the full description on the dataset page: https://huggingface.co/datasets/TingChen-ppmc/Shanghai_Dialect_Conversational_Speech_Corpus.audio1K<n<10K12 likes160 downloads2y agoHugging Face28ShantyCam /exdarkimageobject-detectionn<1K0 likes150 downloads2mo agoHugging Face29shantanugoel /aawaaz-transcript-cleanup-dataset Aawaaz Transcript Cleanup Dataset Training pairs for cleaning messy speech transcripts (ASR output, voice dictation) into well-formatted text while preserving the speaker's voice and meaning. Dataset Description Each example is a pair of: input: A realistic messy transcript with filler words, false starts, self-corrections, grammar errors, and missing punctuation output: The cleaned version with fillers removed, grammar fixed, punctuation added, and domain-appropriate… See the full description on the dataset page: https://huggingface.co/datasets/shantanugoel/aawaaz-transcript-cleanup-dataset.texttext-generation10K<n<100K0 likes143 downloads6mo agoHugging Face30shanthropic /linux-commands LinLM Dataset A curated synthetic dataset for Linux command inference Natural language description -> shell commands Features: Supports 10 languages Arch Linux commands recognition Fine-tune LLM for development, system administration, file operations, Git, Docker, and more Usage from datasets import load_dataset dataset = load_dataset("missvector/linux-commands") def format_for_training(example): return { "prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/shanthropic/linux-commands.texttext-generation10K<n<100K1 likes138 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.