datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Echoes-Platos-CaveEchoes in Plato's Cave — Controlled Speech–Text Corpus
Controlled corpus of 14,400 synthetic English utterances in which the same 600 sentences are
rendered by 6 speakers × 4 emotions, so that speaker identity and prosody vary
while linguistic content is held fixed. It was built for the paper: Echoes in Plato's Cave: Measuring Global and Local Alignment Between Speech and Language Representations, accepted as an oral presentation at the Speech and Audio Language… See the full description on the dataset page: https://huggingface.co/datasets/alefiury/Echoes-Platos-Cave.pi-cavelynx
Coding agent session traces for Ev3lynx727/pi-cavelynx
This dataset contains redacted coding agent session traces exported with pi-share-hf from a local pi workspace. The traces were filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session entry. Entries include session headers, user and assistant… See the full description on the dataset page: https://huggingface.co/datasets/Ev3lynx727/pi-cavelynx.aida-handwritten
Handwritten OCR training data from AIDA-project
Dataset Summary
This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality handwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German.
Supported Tasks
The dataset was created for… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-handwritten.rebus
REBUS
REBUS: A Robust Evaluation Benchmark of Understanding Symbols
Paper | 🤗 Dataset | GitHub | Website
Introduction
Recent advances in large language models have led to the development of multimodal LLMs (MLLMs), which take both image data and text as an input. Virtually all of these models have been announced within the past year, leading to a significant need for benchmarks evaluating the abilities of these models to reason truthfully and accurately on a diverse… See the full description on the dataset page: https://huggingface.co/datasets/cavendishlabs/rebus.chatgpt-clinic-caveats-fact-check
When ChatGPT Adds Caveats to Clinic Recommendations
This version 1.0 companion dataset fact-checks clinic-specific commercial and operational caveat families identified in a frozen corpus of 450 repeated ChatGPT answers from the parent study.
Author: Evgeniy Yudin, Founder and Strategy Lead
ORCID: https://orcid.org/0009-0007-8400-9561
Publisher: Rotgar Research
Published: 2026-09-03
Version DOI: https://doi.org/10.5281/zenodo.22304469
Zenodo record:… See the full description on the dataset page: https://huggingface.co/datasets/RotgarSett/chatgpt-clinic-caveats-fact-check.synthetic-caveman-thinkingcave-datasetaida-ship-info
Handwritten OCR training data from AIDA-project (Ship Registry)
Dataset Summary
This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality handwritten annotations from ship registry records — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German.
Supported… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-ship-info.aida-typewritten
typewritten OCR training data from AIDA-project
Dataset Summary
This dataset contains typewritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality typwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German.
Supported Tasks
The dataset was created for optical… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-typewritten.objectnav-sft-claude-cavemansimplified_soda_kr
SODA-KR (Simplified)
Korean translation of the SODA dataset (simplified version with 40722 samples).
Dataset Description
This is a Korean-translated version of the allenai/soda dataset.
Each sample contains speakers, narrative context, and dialogue translated to Korean.
Source Dataset
Original Dataset: allenai/soda
Original License: CC-BY-4.0
Citation: Please cite the original SODA paper
Translation Details
Translation Model:… See the full description on the dataset page: https://huggingface.co/datasets/CaveduckAI/simplified_soda_kr.steer-personality-rudeness-ko
📊 Dataset Overview
This dataset contains 1000 samples designed for extracting personality steering vectors using the Contrastive Activation Addition (CAA) method. Each sample presents a scenario with two response options: one exhibiting the target personality trait and one neutral.
Dataset Details
Property
Value
Personality Trait
매우 무례한
Generation Model
moonshotai/kimi-k2-0905
Source Dataset
CaveduckAI/simplified_soda_kr
Sample Count
1000
Generated… See the full description on the dataset page: https://huggingface.co/datasets/CaveduckAI/steer-personality-rudeness-ko.steer-personality-extroversion-ko
📊 Dataset Overview
This dataset contains 100 samples designed for extracting personality steering vectors using the Contrastive Activation Addition (CAA) method. Each sample presents a scenario with two response options: one exhibiting the target personality trait and one neutral.
Dataset Details
Property
Value
Personality Trait
외향성 (Extroversion)
Generation Model
moonshotai/kimi-k2-0905
Source Dataset
CaveduckAI/simplified_soda_kr
Sample Count
100… See the full description on the dataset page: https://huggingface.co/datasets/CaveduckAI/steer-personality-extroversion-ko.swe-11k
SWE-11K
This dataset is generated from https://github.com/sdrobac/ijdar-2020.
Quick Start
from datasets import load_dataset
dataset = load_dataset("parquet", data_files="train.parquet", split="train")
Dataset Description
SWE-11K is a Swedish OCR dataset containing 11,094 image-text pairs extracted from historic Swedish newspapers and journals. Suitable for training and evaluating OCR models on Swedish historical text.
Supported Tasks
Optical… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/swe-11k.trl-mlt-2Finnish-OCR-evaluationAn evaluation dataset for Finnish OCR - Paddle OCR Format
trash-mult-dpocaveman-world-knowledge-150k
Caveman World Knowledge 150K
Dataset description
Caveman-style instruction dataset with two blended behaviors:
known world knowledge responses (Wikipedia-like content rewritten in caveman voice)
unknown-question reactions with mood labels: angry, argue, attack
This dataset is intended for instruction tuning and style conditioning.
Dataset structure
Each row is a JSON object with fields:
id: unique row id
source: wikipedia, fallback, or synthetic
topic: world… See the full description on the dataset page: https://huggingface.co/datasets/Blackbean109/caveman-world-knowledge-150k.theseus_ocr_tiny
Theseus Finnish OCR Dataset
Paragraph-level OCR dataset harvested from Theseus.fi,
the Finnish repository of university of applied sciences theses.
Each record is one paragraph crop extracted from a thesis PDF, paired with the
text extracted by pdfplumber.
Image Resolution
Paragraph crops are rendered at 300 DPI (dots per inch) with 2 px
padding on each side. At 300 DPI a standard A4 page is
2481 × 3507 pixels, giving high enough resolution for training
OCR and… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/theseus_ocr_tiny.trash-mult-dpo-2fin-13k
FIN-13K
This dataset is generated from https://github.com/sdrobac/ijdar-2020.
Dataset Summary
FIN-13K is a Finnish OCR dataset containing 13,037 image-text pairs extracted from historic Finnish newspapers and journals. The dataset is derived from the FIN-BERT dataset and is suitable for training and evaluating OCR models on Finnish text.
Supported Tasks
Optical Character Recognition (OCR): Recognizing text from line-level images
OCR Post-correction: For… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/fin-13k.trl-mlt-1humanoid-caveman-dataCaveTrace-M2.7
KIMI-K2.5-1000000x
1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions)
Distribution:
Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl
Computer Science: 5%
Logical Questions: 5%
Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/FadeClip/CaveTrace-M2.7.ice-caves-context-qacavemen
