datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jeb-rag
JEB-Bench
Charging the Gate Rent: Measured-Energy Accounting for Adaptive Retrieval-Augmented Generation
⚠️ Status: under construction. Phase 0 (measurement validation) and Phase 1
(index construction) are landing now. The oracle matrix (bench/oracle/) is
populated in Phase 2 and this card will be revised when it is complete. Do not
cite numbers from this repository until the status line says complete.
What this is
The first public per-query × per-configuration… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/jeb-rag.ko_text2sqlE2AM_ResNet50
E2AM Ablation Results: ResNet-50
Energy-aware training ablation study for ResNet-50 across three image-classification datasets: CIFAR-10, CIFAR-100, and Tiny-ImageNet.
Each dataset has 15 training variants (8 individual-method M0..M7, 7 cumulative ablation C0..C6) at 50 epochs, plus a 5-variant deployment pipeline (FP32 baseline, structured pruning, pruning+finetune, INT8 quantization, pruned+INT8).
Status: 45 completed variants, 0 partial.
Quick links… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/E2AM_ResNet50.SWE-smith-jsSpatialForge
SpatialForge-10M
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
📑 Paper
Zishan Liu, Ruoxi Zang, Yanglin Zhang, Wei Liu, Yin Zhang, Jian Yao, Jiayin Zheng, Zhengzhe Liu
Lingnan University · XPENG Robotics
📦 SpatialForge-10M
A large-scale vision-language dataset designed for 3D-aware spatial perception and reasoning from open-world 2D images.
SpatialForge-10M contains over 10 million QA pairs generated from 2.8 million curated… See the full description on the dataset page: https://huggingface.co/datasets/shana643/SpatialForge.shangrilafrontier
Bangumi Image Base of Shangri-la Frontier
This is the image base of bangumi Shangri-La Frontier, we detected 48 characters, 2678 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/shangrilafrontier.ai-detection-dataset-v2
---dataset_info:
features:
- name: image # use the exact column name from your parquet schema
dtype: image # this forces Hugging Face to render it as an image
- name: label
dtype: string
license: other
task_categories:
- image-classification
language:
- en
tags:
- ai-generated-image-detection
- synthetic-image-detection
- diffusion-models
pretty_name: AI-Generated Image Detection Dataset v2
size_categories:
- 10K<n<100K
AI-Generated… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/ai-detection-dataset-v2.SWE-smith-cppaime_2025_multilingualWhen Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy
https://arxiv.org/abs/2505.22888
Jirui Qi, Shan Chen, Zidi Xiong, Raquel Fernández, Danielle S. Bitterman, Arianna Bisazza
Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studied. This capability is as important as answer accuracy for real world applications because… See the full description on the dataset page: https://huggingface.co/datasets/shanchen/aime_2025_multilingual.SWE-smith-tsgpqa_diamond_mc_multilingualWhen Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy
https://arxiv.org/abs/2505.22888
Jirui Qi, Shan Chen, Zidi Xiong, Raquel Fernández, Danielle S. Bitterman, Arianna Bisazza
Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studied. This capability is as important as answer accuracy for real world applications because… See the full description on the dataset page: https://huggingface.co/datasets/shanchen/gpqa_diamond_mc_multilingual.SWE-smith-javakitti-objectai-image-detection-dataset
AI-Image Detection Dataset
Paired real / AI images, with shared image-grounded captions, for training and
evaluating AI-generated-image detectors.
Each of 10,000 real photos is captioned once (BLIP-2) and paired with one synthetic
partner per generator (6 generators → 60,000 AI images). A real image and all of its
AI partners share the same prompt, so the only systematic difference between the
classes is the generative process itself. A detector trained here is pushed toward the… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/ai-image-detection-dataset.ASVspoof5
ASVspoof 5 (track 1, eval)
Benchmark-ready packaging of the Track 1 (spoofing / deepfake detection) evaluation partition of the ASVspoof 5 challenge, for speech anti-spoofing and synthetic / deepfake voice detection.
Overview
Track 1 is binary classification: bonafide (genuine human speech) vs. spoof (synthetic / converted speech). This packaging contains the full track_1 evaluation set. The original challenge is at https://www.asvspoof.org/.
License &… See the full description on the dataset page: https://huggingface.co/datasets/shanthini33/ASVspoof5.SWE-smith-goChemQA
Dataset Card for ChemQA
Introducing ChemQA: a Multimodal Question-and-Answering Dataset on Chemistry Reasoning. This work is inspired by IsoBench[1] and ChemLLMBench[2].
Content
There are 5 QA Tasks in total:
Counting Numbers of Carbons and Hydrogens in Organic Molecules: adapted from the 600 PubChem molecules created from [2], evenly divided into validation and evaluation datasets.
Calculating Molecular Weights in Organic Molecules: adapted from the 600 PubChem… See the full description on the dataset page: https://huggingface.co/datasets/shangzhu/ChemQA.SpeechTextMatching_Tedlium2Train
Dataset Card for "SpeechTextMatching_TEDLIUM2Train"
More Information needed
SFT-Qwenshanghai_master_plan_beirThis is a copy of https://huggingface.co/datasets/jinaai/shanghai_master_plan reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/shanghai_master_plan_beir.expresso
Expresso — audio + text
A faithful re-publication of the official Expresso
dataset (Nguyen et al., Interspeech 2023) as a loadable HuggingFace audio dataset, sourced
directly from FAIR's official tar.
⚠️ License: CC-BY-NC-4.0 — non-commercial use only.
Configs
read — 11.6k mono read-speech utterances with human transcripts.
conversational — ~15.9k mono per-utterance turns derived from the stereo conversational dialogues, transcribed with Whisper Large V3 Turbo.… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/expresso.SWE-smith-pypokemon_data
Pokémon TCG Tournament Game Logs
Full game logs from an agent Pokémon TCG tournament ("Limited Card Battle"),
122,542 recorded games collected across 26 daily archives — an initial batch
of 15 undated archives (top/, medium_high/) plus dated days
2026-07-21 → 2026-07-31 (days_0721_0731/). Both batches use the same
per-day split rule and load together via the split config above.
Splits
Games in each daily archive were ranked by avg_score (per-agent average
rating… See the full description on the dataset page: https://huggingface.co/datasets/shantezhou/pokemon_data.shanghaitech-crowd-countingDACTYL
DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large language models Dataset
The DACTYL dataset is an AI-generated text detection dataset focusing primarily on one-shot or few-shot examples. We also include texts from continued pre-trained small language models.
For more information, refer to our paper.
Models Used
We used the following LLMs to generate texts.
OpenAI’s GPT-4o-mini and GPT-4o
Anthropic’s Claude Haiku and Sonnet 3.5
Mistral Small (24B)and… See the full description on the dataset page: https://huggingface.co/datasets/ShantanuT01/DACTYL.ViMedCSS-Cop
🩺 ViMedCSS: A Vietnamese Medical Code-Switching Speech Dataset (LREC 2026)
📖 Overview
ViMedCSS is a Vietnamese medical speech dataset for code-switching ASR, where each utterance contains at least one non-Vietnamese (mainly English) medical term embedded in Vietnamese speech.
📊 Dataset Statistics
Split Statistics (from ViMedCSS-Metadata)
Split
# Rows
Duration (hours)
Avg duration (s)
Total CS terms
train
11,832
24.30
7.39
12,314… See the full description on the dataset page: https://huggingface.co/datasets/shannonnonshan/ViMedCSS-Cop.Shanghai_Dialect_Conversational_Speech_Corpus
Corpus
This dataset is built from Magicdata ASR-CZDIACSC: A CHINESE SHANGHAI DIALECT CONVERSATIONAL SPEECH CORPUS
This corpus is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. Please refer to the license for further information.
Modifications: The audio is split in sentences based on the time span on the transcription file. Sentences that span less than 1 second is discarded. Topics of conversation is removed.
Usage… See the full description on the dataset page: https://huggingface.co/datasets/TingChen-ppmc/Shanghai_Dialect_Conversational_Speech_Corpus.exdarkaawaaz-transcript-cleanup-dataset
Aawaaz Transcript Cleanup Dataset
Training pairs for cleaning messy speech transcripts (ASR output, voice dictation) into well-formatted text while preserving the speaker's voice and meaning.
Dataset Description
Each example is a pair of:
input: A realistic messy transcript with filler words, false starts, self-corrections, grammar errors, and missing punctuation
output: The cleaned version with fillers removed, grammar fixed, punctuation added, and domain-appropriate… See the full description on the dataset page: https://huggingface.co/datasets/shantanugoel/aawaaz-transcript-cleanup-dataset.linux-commands
LinLM Dataset
A curated synthetic dataset for Linux command inference
Natural language description -> shell commands
Features:
Supports 10 languages
Arch Linux commands recognition
Fine-tune LLM for development, system administration, file operations, Git, Docker, and more
Usage
from datasets import load_dataset
dataset = load_dataset("missvector/linux-commands")
def format_for_training(example):
return {
"prompt": f"Convert to Linux command:… See the full description on the dataset page: https://huggingface.co/datasets/shanthropic/linux-commands.
