datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cavewoman-data
CAVEWOMAN: Generations Under Linguistic Input and Output Compression
Raw model generations for CAVEWOMAN, a two-channel evaluation protocol that
measures how large language models behave when either the user prompt
(input compression) or the model response (output compression) is forced
into a reduced linguistic register. Every generation is scored on task
accuracy, realised per-item token cost, and surface-text preservation against
the model's own unconstrained (L0) reference.… See the full description on the dataset page: https://huggingface.co/datasets/rayascript/cavewoman-data.cave
Description
This database contains a set multispectral images that were used to emulate the GAP camera. The images are of a wide variety of real-world materials and objects.
Image capture information
Camera
Cooled CCD camera (Apogee Alta U260)
Resolution
512 x 512 pixel
Filter
VariSpec liquid crystal tunable filter
Illuminant
CIE Standard Illuminant D65
Range of wevelength
400nm - 700nm
Steps
10nm
Number of band
31 band
Focal length
f/1.4… See the full description on the dataset page: https://huggingface.co/datasets/danaroth/cave.Echoes-Platos-CaveEchoes in Plato's Cave — Controlled Speech–Text Corpus
Controlled corpus of 14,400 synthetic English utterances in which the same 600 sentences are
rendered by 6 speakers × 4 emotions, so that speaker identity and prosody vary
while linguistic content is held fixed. It was built for the paper: Echoes in Plato's Cave: Measuring Global and Local Alignment Between Speech and Language Representations, accepted as an oral presentation at the Speech and Audio Language… See the full description on the dataset page: https://huggingface.co/datasets/alefiury/Echoes-Platos-Cave.cavernpi-cavelynx
Coding agent session traces for Ev3lynx727/pi-cavelynx
This dataset contains redacted coding agent session traces exported with pi-share-hf from a local pi workspace. The traces were filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session entry. Entries include session headers, user and assistant… See the full description on the dataset page: https://huggingface.co/datasets/Ev3lynx727/pi-cavelynx.aida-handwritten
Handwritten OCR training data from AIDA-project
Dataset Summary
This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality handwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German.
Supported Tasks
The dataset was created for… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-handwritten.cavernrebus
REBUS
REBUS: A Robust Evaluation Benchmark of Understanding Symbols
Paper | 🤗 Dataset | GitHub | Website
Introduction
Recent advances in large language models have led to the development of multimodal LLMs (MLLMs), which take both image data and text as an input. Virtually all of these models have been announced within the past year, leading to a significant need for benchmarks evaluating the abilities of these models to reason truthfully and accurately on a diverse… See the full description on the dataset page: https://huggingface.co/datasets/cavendishlabs/rebus.cave_bench
CAVE-Bench
You're Right, Let Me Fix It: How LLM Agents Damage Correct Work When Falsely Accused
Complete task pack: 365 Harbor tasks (172 inherited-resume, 193 self-built) with environments, verifiers, skills, adapters, and scoring.
After the work is already correct, a later message falsely accuses the agent.
Code and website: https://github.com/henrymao2004/agent-over-correction
Gallery: https://henrymao2004.github.io/agent-over-correction/gallery.html
arXiv: coming soon… See the full description on the dataset page: https://huggingface.co/datasets/sevens2004/cave_bench.CAVECAVE
Dataset Card for CAVE: Commonsense Anomalies in Visual Environments
🏠 Project Page📄 Paper (EMNLP 2025)💻 Code
Dataset Details
Dataset Description
CAVE is the first benchmark of real-world visual anomalies for evaluating Vision-Language Models (VLMs). It is curated from images captured in real-life settings (photographs and screenshots taken by individuals), sourced from Reddit.
The benchmark is grounded in cognitive science literature on how humans detect and… See the full description on the dataset page: https://huggingface.co/datasets/epfl-nlp/CAVE.chatgpt-clinic-caveats-fact-check
When ChatGPT Adds Caveats to Clinic Recommendations
This version 1.0 companion dataset fact-checks clinic-specific commercial and operational caveat families identified in a frozen corpus of 450 repeated ChatGPT answers from the parent study.
Author: Evgeniy Yudin, Founder and Strategy Lead
ORCID: https://orcid.org/0009-0007-8400-9561
Publisher: Rotgar Research
Published: 2026-09-03
Version DOI: https://doi.org/10.5281/zenodo.22304469
Zenodo record:… See the full description on the dataset page: https://huggingface.co/datasets/RotgarSett/chatgpt-clinic-caveats-fact-check.synthetic-caveman-thinkingcave-datasetaida-ship-info
Handwritten OCR training data from AIDA-project (Ship Registry)
Dataset Summary
This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality handwritten annotations from ship registry records — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German.
Supported… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-ship-info.aida-typewritten
typewritten OCR training data from AIDA-project
Dataset Summary
This dataset contains typewritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality typwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German.
Supported Tasks
The dataset was created for optical… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-typewritten.objectnav-sft-claude-cavemansimplified_soda_kr
SODA-KR (Simplified)
Korean translation of the SODA dataset (simplified version with 40722 samples).
Dataset Description
This is a Korean-translated version of the allenai/soda dataset.
Each sample contains speakers, narrative context, and dialogue translated to Korean.
Source Dataset
Original Dataset: allenai/soda
Original License: CC-BY-4.0
Citation: Please cite the original SODA paper
Translation Details
Translation Model:… See the full description on the dataset page: https://huggingface.co/datasets/CaveduckAI/simplified_soda_kr.steer-personality-rudeness-ko
📊 Dataset Overview
This dataset contains 1000 samples designed for extracting personality steering vectors using the Contrastive Activation Addition (CAA) method. Each sample presents a scenario with two response options: one exhibiting the target personality trait and one neutral.
Dataset Details
Property
Value
Personality Trait
매우 무례한
Generation Model
moonshotai/kimi-k2-0905
Source Dataset
CaveduckAI/simplified_soda_kr
Sample Count
1000
Generated… See the full description on the dataset page: https://huggingface.co/datasets/CaveduckAI/steer-personality-rudeness-ko.steer-personality-lewd-ko
📊 Dataset Overview
This dataset contains 15000 samples designed for extracting personality steering vectors using the Contrastive Activation Addition (CAA) method. Each sample presents a scenario with two response options: one exhibiting the target personality trait and one neutral.
Dataset Details
Property
Value
Personality Trait
Obscene, Erotic, Lewd, Sexy, Flirty, Playful, Seductive
Generation Model
moonshotai/kimi-k2-0905
Source Dataset… See the full description on the dataset page: https://huggingface.co/datasets/CaveduckAI/steer-personality-lewd-ko.cave-johnsonsteer-personality-extroversion-ko
📊 Dataset Overview
This dataset contains 100 samples designed for extracting personality steering vectors using the Contrastive Activation Addition (CAA) method. Each sample presents a scenario with two response options: one exhibiting the target personality trait and one neutral.
Dataset Details
Property
Value
Personality Trait
외향성 (Extroversion)
Generation Model
moonshotai/kimi-k2-0905
Source Dataset
CaveduckAI/simplified_soda_kr
Sample Count
100… See the full description on the dataset page: https://huggingface.co/datasets/CaveduckAI/steer-personality-extroversion-ko.swe-11k
SWE-11K
This dataset is generated from https://github.com/sdrobac/ijdar-2020.
Quick Start
from datasets import load_dataset
dataset = load_dataset("parquet", data_files="train.parquet", split="train")
Dataset Description
SWE-11K is a Swedish OCR dataset containing 11,094 image-text pairs extracted from historic Swedish newspapers and journals. Suitable for training and evaluating OCR models on Swedish historical text.
Supported Tasks
Optical… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/swe-11k.trl-mlt-2cave
Description
This database contains a set multispectral images that were used to emulate the GAP camera. The images are of a wide variety of real-world materials and objects.
Image capture information
Camera
Cooled CCD camera (Apogee Alta U260)
Resolution
512 x 512 pixel
Filter
VariSpec liquid crystal tunable filter
Illuminant
CIE Standard Illuminant D65
Range of wevelength
400nm - 700nm
Steps
10nm
Number of band
31 band
Focal length
f/1.4… See the full description on the dataset page: https://huggingface.co/datasets/fsd1970575668/cave.Finnish-OCR-evaluationAn evaluation dataset for Finnish OCR - Paddle OCR Format
trash-mult-dpocavejohnson_en-ljspeech
Cave Johnson — English (en)
LJSpeech dataset of Cave Johnson (en).
234 pairs
44.1kHz mono 16-bit PCM WAV (original wiki quality)
Piper TTS Training (High Quality on T4 GPU)
Preprocessing (downsample to 22.05kHz)
python3 -m piper_train.preprocess \
--language en-us \
--input-dir ./cavejohnson_en \
--output-dir ./train_cavejohnson_en \
--dataset-format ljspeech \
--single-speaker \
--sample-rate 22050
Training (Kaggle T4… See the full description on the dataset page: https://huggingface.co/datasets/RoxasYTB/cavejohnson_en-ljspeech.caveman-world-knowledge-150k
Caveman World Knowledge 150K
Dataset description
Caveman-style instruction dataset with two blended behaviors:
known world knowledge responses (Wikipedia-like content rewritten in caveman voice)
unknown-question reactions with mood labels: angry, argue, attack
This dataset is intended for instruction tuning and style conditioning.
Dataset structure
Each row is a JSON object with fields:
id: unique row id
source: wikipedia, fallback, or synthetic
topic: world… See the full description on the dataset page: https://huggingface.co/datasets/Blackbean109/caveman-world-knowledge-150k.theseus_ocr_tiny
Theseus Finnish OCR Dataset
Paragraph-level OCR dataset harvested from Theseus.fi,
the Finnish repository of university of applied sciences theses.
Each record is one paragraph crop extracted from a thesis PDF, paired with the
text extracted by pdfplumber.
Image Resolution
Paragraph crops are rendered at 300 DPI (dots per inch) with 2 px
padding on each side. At 300 DPI a standard A4 page is
2481 × 3507 pixels, giving high enough resolution for training
OCR and… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/theseus_ocr_tiny.
