datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
raid
🚨 RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors 🚨
🌐 Website, 🖥️ Github, 📝 Paper
RAID is the largest & most comprehensive dataset for evaluating AI-generated text detectors.
It contains over 10 million documents spanning 11 LLMs, 11 genres, 4 decoding strategies, and 12 adversarial attacks.
It is designed to be the go-to location for trustworthy third-party evaluation of both open-source and closed-source generated text detectors.
Load… See the full description on the dataset page: https://huggingface.co/datasets/liamdugan/raid.RAIDThis dataset is for testing the adversarial robustness of AI-Generated Image Detectors, as described in the paper RAID: A Dataset for Testing the Adversarial Robustness of AI-Generated Image Detectors.
RadImageNet-VQA
RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
We introduce RadImageNet-VQA, a large-scale dataset designed for training and benchmarking radiologic VQA on CT and MRI exams. Built from the CT/MRI subset of RadImageNet and its expert-curated anatomical and pathological annotations, RadImageNet-VQA provides 750K images with 7.5M generated samples, including 750K medical captions for visual-text alignment and 6.75M… See the full description on the dataset page: https://huggingface.co/datasets/raidium/RadImageNet-VQA.rise-of-the-tomb-raider-gameplay-data
古墓丽影:崛起
This public dataset repository contains local gameplay data uploaded from F:\古墓丽影:崛起.
Contents
Files: 433
Total local size: 254.19 GB
Generated: 2026-06-08 19:29:07 UTC
File Types
.jsonl: 136
.json: 105
.png: 102
.mkv: 35
.txt: 33
.parquet: 22
Notes
This repository may contain gameplay video, images, Parquet files, JSON/JSONL metadata, and keyboard/mouse event logs.
The license is marked as other; review game footage… See the full description on the dataset page: https://huggingface.co/datasets/xiaoluo11/rise-of-the-tomb-raider-gameplay-data.rl-game-traces-rise-of-the-tomb-raider
古墓丽影:崛起
This public dataset repository contains gameplay trace data uploaded from F:\古墓丽影:崛起.
Contents
Files: 647
Total local size: 496.11 GB
Generated: 2026-06-06T01:03:39+00:00
File Types
.jsonl: 196
.json: 149
.png: 129
.parquet: 49
.mkv: 49
.txt: 49
.jpg: 15
.exe: 11
Notes
This repository may contain gameplay video, Parquet files, JSON/JSONL metadata, and input event logs.
The license is marked as other; review game footage… See the full description on the dataset page: https://huggingface.co/datasets/yinhuankuang/rl-game-traces-rise-of-the-tomb-raider.SID_Set
Dataset Card for SID_Set
Dataset Summary
We provide Social media Image Detection dataSet (SID-Set), which offers three key advantages:
Extensive volume: Featuring 300K AI-generated/tampered and authentic images with comprehensive annotations.
Broad diversity: Encompassing fully synthetic and tampered images across various classes.
Elevated realism: Including images that are predominantly indistinguishable from genuine ones through mere visual inspection.
Please… See the full description on the dataset page: https://huggingface.co/datasets/RAID-techjam/SID_Set.flaird-raid-pan261920-raider-waite-tarot-public-domain1920-raider-waite-tarot-public-domainRAID_none-encoded-gpt2Raiden-DeepSeek-R1Click here to support our open-source dataset and model releases!
Raiden-DeepSeek-R1 is a dataset containing creative-reasoning and analytic-reasoning responses, testing the limits of DeepSeek R1's reasoning skills!
This dataset contains:
63k 'creative_content' and 'analytical_reasoning' prompts from microsoft/orca-agentinstruct-1M-v1, with all responses generated by deepseek-ai/DeepSeek-R1.
Responses demonstrate the reasoning capabilities of DeepSeek's 685b parameter R1 reasoning model.… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Raiden-DeepSeek-R1.raiden_garments_folding_baseline_a2
raiden_garments_folding_baseline_a2
LeRobot v2.1 dataset of bimanual YAM (Raiden) garment folding, converted with per-subtask language labels (cell A2).
129 episodes / 95,762 frames / 30 fps
21 unique task strings
Cameras: observation.images.top, left_wrist, right_wrist (224×224, AV1)
State/action: 14-D joints (left 6+gripper, right 6+gripper)
Each LeRobot episode is one labeled subtask (not the high-level collect prompt). Typical sequence per garment:
single out the… See the full description on the dataset page: https://huggingface.co/datasets/Sshawnin/raiden_garments_folding_baseline_a2.executable-counterfactuals
Introduction
This repo contains all training and evaluation datasets used in "Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code". This work has been published in ICLR 2026.
Arxiv Paper
Github Repo (Work in Progress)
Counterfactual reasoning, a hallmark of intelligence, consists of three steps: inferring latent variables from observations (abduction), constructing alternative situations (interventions), and predicting the outcomes of the alternatives… See the full description on the dataset page: https://huggingface.co/datasets/Raidriar-Dai/executable-counterfactuals.Genshin_Impact_RaidenShogun_Voice_koreanRaid_split1920-raider-waite-tarot-public-domain-cleanedA cleaned up version of the multimodalart/1920-raider-waite-tarot-public-domain dataset, without the card borders and names
raidex-resultsraidex-requestsbackuplienchiangprobslistraid
🚨 RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors 🚨
🌐 Website, 🖥️ Github, 📝 Paper
RAID is the largest & most comprehensive dataset for evaluating AI-generated text detectors.
It contains over 10 million documents spanning 11 LLMs, 11 genres, 4 decoding strategies, and 12 adversarial attacks.
It is designed to be the go-to location for trustworthy third-party evaluation of both open-source and closed-source generated text detectors.
Load… See the full description on the dataset page: https://huggingface.co/datasets/Akjhtar/raid.raid-neologism-table-splits
RAID neologism table splits
Source-disjoint RAID train-derived paired splits for the two-token AI detector experiments.
Source dataset: liamdugan/raid, config raid, split train.
Seed: 20260501.
Base source partition: 10000 train source_ids, 3000 test source_ids, overlap 0.
Row format: one human text and one same-source_id AI text per row.
Protocols:
standard_train, standard_test: model, attack, decoding, repetition penalty, and domain sampled randomly.
model_<model>_train… See the full description on the dataset page: https://huggingface.co/datasets/danielfein/raid-neologism-table-splits.manga-colorization-masterMetricEval-BodyCT
MetricEval-BodyCT
This repository is a body CT benchmark for evaluating radiology report-generation metrics against
radiologists' judgment.
It covers 100 CT studies (50 chest and 50 abdomen/pelvis), with three candidate reports each. Every
candidate was independently annotated by multiple board-certified radiologists. The reference reports
are de-identified radiology reports from multiple US centers, provided by Segmed and redistributed
under the Data Use Agreement in LICENSE.… See the full description on the dataset page: https://huggingface.co/datasets/raidium/MetricEval-BodyCT.testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 12,
"total_frames": 6835,
"total_tasks":2,
"total_videos": 24,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:12"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/raidavid/test.Raiders-Of-The-Lost-Kek
Raiders Of The Lost Kek
The largest 4chan /pol/ dataset.
I extracted the post content, removed HTML nonesense, and 4chan specific things
like post number replies in text, etc.
There are a few sizes of datasets available
100kLines - first 100,000 lines of text from the dataset
300kLines - first 300,000 lines of text from the dataset
500kLines - first 500,000 lines of text from the dataset
maybe at some point once i have the compute ill upload the whole thing
link :… See the full description on the dataset page: https://huggingface.co/datasets/eddsterxyz/Raiders-Of-The-Lost-Kek.RAID-Plus
RAID+
🌐 Project Page, 🖥️ Code, 📊 Original RAID
RAID+ is an evaluation-only extension of the RAID benchmark(Dugan et al., 2024), regenerating RAID prompts using contemporary frontier models absent from the original dataset. It is intended for evaluating MGT detectors against LLMs released after RAID's publication.
This dataset was constructed as part of INSCONE: Unknown-Aware Detection of LLM-Generated Text via Informed Wild Data.
Models
Model
Samples… See the full description on the dataset page: https://huggingface.co/datasets/markstanl/RAID-Plus.Raiden-DeepSeek-R1-PREVIEWThis is a preview of the full Raiden-Deepseek-R1 creative and analytical reasoning dataset, containing the first ~6k rows. Get the full dataset here!
This dataset uses synthetic data generated by deepseek-ai/DeepSeek-R1.
The initial release of Raiden uses 'creative_content' and 'analytical_reasoning' prompts from microsoft/orca-agentinstruct-1M-v1.
Dataset has not been reviewed for format or accuracy. All responses are synthetic and provided without editing.
Use as you will.
atc-tts-voxtream
ATC TTS Voxtream Dataset
This dataset is prepared for training the Voxtream TTS model and follows the format of herimor/voxtream-train-9k.
Source:
Based on the Singaporean SG Eleven dataset: aether-raid/sg-aviation-el-combined
Description:
All features are provided as single batched .npy files:
mimi_codes_16cb.npy: Mimi codec tokens (16 codebooks)
phone_emb_indices.npy: Alignment of phoneme tokens to Mimi frames
phone_tokens.npy: Phoneme tokens
sem_label_shifts.npy: Monotonic… See the full description on the dataset page: https://huggingface.co/datasets/aether-raid/atc-tts-voxtream.details_Kquant03__Raiden-16x3.43B
Dataset Card for Evaluation run of Kquant03/Raiden-16x3.43B
Dataset automatically created during the evaluation run of model Kquant03/Raiden-16x3.43B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Raiden-16x3.43B.BraTS-2024-Complete
BraTS 2024 Complete Prepared Dataset
Brain Tumor Segmentation (Leave a like 💖 if this helped you)
Dataset Description
This is an organized and verified version of the BraTS 2024 challenge datasets, including three tumor types.
Included Datasets
Dataset
Type
Cases
Source
BraTS-GLI
Glioma
1,809
Synapse (Dec 2024)
BraTS-MEN-RT
Meningioma + RT
571
Synapse (Feb 2025)
BraTS-PED
Pediatric
348
Cancer Imaging Archive… See the full description on the dataset page: https://huggingface.co/datasets/RaidenShogunUltimate/BraTS-2024-Complete.
