CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01liamdugan /raid 🚨 RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors 🚨 🌐 Website, 🖥️ Github, 📝 Paper RAID is the largest & most comprehensive dataset for evaluating AI-generated text detectors. It contains over 10 million documents spanning 11 LLMs, 11 genres, 4 decoding strategies, and 12 adversarial attacks. It is designed to be the go-to location for trustworthy third-party evaluation of both open-source and closed-source generated text detectors. Load… See the full description on the dataset page: https://huggingface.co/datasets/liamdugan/raid.texttext-classification1M<n<10M27 likes5.5k downloads2y agoHugging Face02aimagelab /RAIDThis dataset is for testing the adversarial robustness of AI-Generated Image Detectors, as described in the paper RAID: A Dataset for Testing the Adversarial Robustness of AI-Generated Image Detectors. imageimage-classification2 likes4.5k downloads1y agoHugging Face03raidium /RadImageNet-VQAgated RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering We introduce RadImageNet-VQA, a large-scale dataset designed for training and benchmarking radiologic VQA on CT and MRI exams. Built from the CT/MRI subset of RadImageNet and its expert-curated anatomical and pathological annotations, RadImageNet-VQA provides 750K images with 7.5M generated samples, including 750K medical captions for visual-text alignment and 6.75M… See the full description on the dataset page: https://huggingface.co/datasets/raidium/RadImageNet-VQA.imagevisual-question-answering1M<n<10M92 likes983 downloads3mo agoHugging Face04xiaoluo11 /rise-of-the-tomb-raider-gameplay-data 古墓丽影:崛起 This public dataset repository contains local gameplay data uploaded from F:\古墓丽影:崛起. Contents Files: 433 Total local size: 254.19 GB Generated: 2026-06-08 19:29:07 UTC File Types .jsonl: 136 .json: 105 .png: 102 .mkv: 35 .txt: 33 .parquet: 22 Notes This repository may contain gameplay video, images, Parquet files, JSON/JSONL metadata, and keyboard/mouse event logs. The license is marked as other; review game footage… See the full description on the dataset page: https://huggingface.co/datasets/xiaoluo11/rise-of-the-tomb-raider-gameplay-data.videoreinforcement-learning0 likes953 downloads4mo agoHugging Face05yinhuankuang /rl-game-traces-rise-of-the-tomb-raider 古墓丽影:崛起 This public dataset repository contains gameplay trace data uploaded from F:\古墓丽影:崛起. Contents Files: 647 Total local size: 496.11 GB Generated: 2026-06-06T01:03:39+00:00 File Types .jsonl: 196 .json: 149 .png: 129 .parquet: 49 .mkv: 49 .txt: 49 .jpg: 15 .exe: 11 Notes This repository may contain gameplay video, Parquet files, JSON/JSONL metadata, and input event logs. The license is marked as other; review game footage… See the full description on the dataset page: https://huggingface.co/datasets/yinhuankuang/rl-game-traces-rise-of-the-tomb-raider.imagereinforcement-learning0 likes897 downloads4mo agoHugging Face06RAID-techjam /SID_Set Dataset Card for SID_Set Dataset Summary We provide Social media Image Detection dataSet (SID-Set), which offers three key advantages: Extensive volume: Featuring 300K AI-generated/tampered and authentic images with comprehensive annotations. Broad diversity: Encompassing fully synthetic and tampered images across various classes. Elevated realism: Including images that are predominantly indistinguishable from genuine ones through mere visual inspection. Please… See the full description on the dataset page: https://huggingface.co/datasets/RAID-techjam/SID_Set.imagetext-to-image100K<n<1M0 likes642 downloads23d agoHugging Face07MahmoodAnaam /flaird-raid-pan26text10M<n<100M0 likes438 downloads3d agoHugging Face08multimodalart /1920-raider-waite-tarot-public-domainimagen<1K58 likes294 downloads2y agoHugging Face09microssroads /1920-raider-waite-tarot-public-domainimagen<1K1 likes237 downloads3mo agoHugging Face10TheItCrOw /RAID_none-encoded-gpt2tabular100K<n<1M0 likes217 downloads1y agoHugging Face11sequelbox /Raiden-DeepSeek-R1Click here to support our open-source dataset and model releases! Raiden-DeepSeek-R1 is a dataset containing creative-reasoning and analytic-reasoning responses, testing the limits of DeepSeek R1's reasoning skills! This dataset contains: 63k 'creative_content' and 'analytical_reasoning' prompts from microsoft/orca-agentinstruct-1M-v1, with all responses generated by deepseek-ai/DeepSeek-R1. Responses demonstrate the reasoning capabilities of DeepSeek's 685b parameter R1 reasoning model.… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Raiden-DeepSeek-R1.texttext-generation10K<n<100K52 likes176 downloads2y agoHugging Face12Sshawnin /raiden_garments_folding_baseline_a2 raiden_garments_folding_baseline_a2 LeRobot v2.1 dataset of bimanual YAM (Raiden) garment folding, converted with per-subtask language labels (cell A2). 129 episodes / 95,762 frames / 30 fps 21 unique task strings Cameras: observation.images.top, left_wrist, right_wrist (224×224, AV1) State/action: 14-D joints (left 6+gripper, right 6+gripper) Each LeRobot episode is one labeled subtask (not the high-level collect prompt). Typical sequence per garment: single out the… See the full description on the dataset page: https://huggingface.co/datasets/Sshawnin/raiden_garments_folding_baseline_a2.tabularrobotics10K<n<100K0 likes155 downloads26d agoHugging Face13Raidriar-Dai /executable-counterfactuals Introduction This repo contains all training and evaluation datasets used in "Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code". This work has been published in ICLR 2026. Arxiv Paper Github Repo (Work in Progress) Counterfactual reasoning, a hallmark of intelligence, consists of three steps: inferring latent variables from observations (abduction), constructing alternative situations (interventions), and predicting the outcomes of the alternatives… See the full description on the dataset page: https://huggingface.co/datasets/Raidriar-Dai/executable-counterfactuals.tabular1K<n<10K0 likes135 downloads29d agoHugging Face14ricecake /Genshin_Impact_RaidenShogun_Voice_koreanaudion<1K2 likes134 downloads3y agoHugging Face15Shengkun /Raid_splittext100K<n<1M0 likes130 downloads1y agoHugging Face16multimodalart /1920-raider-waite-tarot-public-domain-cleanedA cleaned up version of the multimodalart/1920-raider-waite-tarot-public-domain dataset, without the card borders and names imagen<1K1 likes124 downloads1y agoHugging Face17cloudronin /raidex-results0 likes121 downloads2mo agoHugging Face18cloudronin /raidex-requests0 likes95 downloads3mo agoHugging Face19raidavid /backuplienchiangprobslistaudio10K<n<100K0 likes87 downloads6d agoHugging Face20Akjhtar /raid 🚨 RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors 🚨 🌐 Website, 🖥️ Github, 📝 Paper RAID is the largest & most comprehensive dataset for evaluating AI-generated text detectors. It contains over 10 million documents spanning 11 LLMs, 11 genres, 4 decoding strategies, and 12 adversarial attacks. It is designed to be the go-to location for trustworthy third-party evaluation of both open-source and closed-source generated text detectors. Load… See the full description on the dataset page: https://huggingface.co/datasets/Akjhtar/raid.texttext-classification1M<n<10M0 likes79 downloads1mo agoHugging Face21danielfein /raid-neologism-table-splits RAID neologism table splits Source-disjoint RAID train-derived paired splits for the two-token AI detector experiments. Source dataset: liamdugan/raid, config raid, split train. Seed: 20260501. Base source partition: 10000 train source_ids, 3000 test source_ids, overlap 0. Row format: one human text and one same-source_id AI text per row. Protocols: standard_train, standard_test: model, attack, decoding, repetition penalty, and domain sampled randomly. model_<model>_train… See the full description on the dataset page: https://huggingface.co/datasets/danielfein/raid-neologism-table-splits.0 likes69 downloads5mo agoHugging Face22Raid41 /manga-colorization-masterimagen<1K0 likes61 downloads3y agoHugging Face23raidium /MetricEval-BodyCTgated MetricEval-BodyCT This repository is a body CT benchmark for evaluating radiology report-generation metrics against radiologists' judgment. It covers 100 CT studies (50 chest and 50 abdomen/pelvis), with three candidate reports each. Every candidate was independently annotated by multiple board-certified radiologists. The reference reports are de-identified radiology reports from multiple US centers, provided by Segmed and redistributed under the Data Use Agreement in LICENSE.… See the full description on the dataset page: https://huggingface.co/datasets/raidium/MetricEval-BodyCT.tabulartext-rankingn<1K6 likes59 downloads6d agoHugging Face24raidavid /testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "so100", "total_episodes": 12, "total_frames": 6835, "total_tasks":2, "total_videos": 24, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:12" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/raidavid/test.tabularrobotics1K<n<10K0 likes58 downloads2y agoHugging Face25eddsterxyz /Raiders-Of-The-Lost-Kek Raiders Of The Lost Kek The largest 4chan /pol/ dataset. I extracted the post content, removed HTML nonesense, and 4chan specific things like post number replies in text, etc. There are a few sizes of datasets available 100kLines - first 100,000 lines of text from the dataset 300kLines - first 300,000 lines of text from the dataset 500kLines - first 500,000 lines of text from the dataset maybe at some point once i have the compute ill upload the whole thing link :… See the full description on the dataset page: https://huggingface.co/datasets/eddsterxyz/Raiders-Of-The-Lost-Kek.text1M<n<10M7 likes44 downloads3y agoHugging Face26markstanl /RAID-Plus RAID+ 🌐 Project Page, 🖥️ Code, 📊 Original RAID RAID+ is an evaluation-only extension of the RAID benchmark(Dugan et al., 2024), regenerating RAID prompts using contemporary frontier models absent from the original dataset. It is intended for evaluating MGT detectors against LLMs released after RAID's publication. This dataset was constructed as part of INSCONE: Unknown-Aware Detection of LLM-Generated Text via Informed Wild Data. Models Model Samples… See the full description on the dataset page: https://huggingface.co/datasets/markstanl/RAID-Plus.tabulartext-classification1K<n<10K1 likes43 downloads4mo agoHugging Face27sequelbox /Raiden-DeepSeek-R1-PREVIEWThis is a preview of the full Raiden-Deepseek-R1 creative and analytical reasoning dataset, containing the first ~6k rows. Get the full dataset here! This dataset uses synthetic data generated by deepseek-ai/DeepSeek-R1. The initial release of Raiden uses 'creative_content' and 'analytical_reasoning' prompts from microsoft/orca-agentinstruct-1M-v1. Dataset has not been reviewed for format or accuracy. All responses are synthetic and provided without editing. Use as you will. text1K<n<10K6 likes41 downloads2y agoHugging Face28aether-raid /atc-tts-voxtream ATC TTS Voxtream Dataset This dataset is prepared for training the Voxtream TTS model and follows the format of herimor/voxtream-train-9k. Source: Based on the Singaporean SG Eleven dataset: aether-raid/sg-aviation-el-combined Description: All features are provided as single batched .npy files: mimi_codes_16cb.npy: Mimi codec tokens (16 codebooks) phone_emb_indices.npy: Alignment of phoneme tokens to Mimi frames phone_tokens.npy: Phoneme tokens sem_label_shifts.npy: Monotonic… See the full description on the dataset page: https://huggingface.co/datasets/aether-raid/atc-tts-voxtream.0 likes40 downloads11mo agoHugging Face29open-llm-leaderboard-old /details_Kquant03__Raiden-16x3.43B Dataset Card for Evaluation run of Kquant03/Raiden-16x3.43B Dataset automatically created during the evaluation run of model Kquant03/Raiden-16x3.43B on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Kquant03__Raiden-16x3.43B.0 likes39 downloads3y agoHugging Face30RaidenShogunUltimate /BraTS-2024-Complete BraTS 2024 Complete Prepared Dataset Brain Tumor Segmentation (Leave a like 💖 if this helped you) Dataset Description This is an organized and verified version of the BraTS 2024 challenge datasets, including three tumor types. Included Datasets Dataset Type Cases Source BraTS-GLI Glioma 1,809 Synapse (Dec 2024) BraTS-MEN-RT Meningioma + RT 571 Synapse (Feb 2025) BraTS-PED Pediatric 348 Cancer Imaging Archive… See the full description on the dataset page: https://huggingface.co/datasets/RaidenShogunUltimate/BraTS-2024-Complete.imageimage-segmentation10K<n<100K0 likes39 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.