cave
Datasets
All datasets matching “cave”cavewoman-data
CAVEWOMAN: Generations Under Linguistic Input and Output Compression
Raw model generations for CAVEWOMAN, a two-channel evaluation protocol that
measures how large language models behave when either the user prompt
(input compression) or the model response (output compression) is forced
into a reduced linguistic register. Every generation is scored on task
accuracy, realised per-item token cost, and surface-text preservation against
the model's own unconstrained (L0) reference.… See the full description on the dataset page: https://huggingface.co/datasets/rayascript/cavewoman-data.cave
Description
This database contains a set multispectral images that were used to emulate the GAP camera. The images are of a wide variety of real-world materials and objects.
Image capture information
Camera
Cooled CCD camera (Apogee Alta U260)
Resolution
512 x 512 pixel
Filter
VariSpec liquid crystal tunable filter
Illuminant
CIE Standard Illuminant D65
Range of wevelength
400nm - 700nm
Steps
10nm
Number of band
31 band
Focal length
f/1.4… See the full description on the dataset page: https://huggingface.co/datasets/danaroth/cave.Echoes-Platos-CaveEchoes in Plato's Cave — Controlled Speech–Text Corpus
Controlled corpus of 14,400 synthetic English utterances in which the same 600 sentences are
rendered by 6 speakers × 4 emotions, so that speaker identity and prosody vary
while linguistic content is held fixed. It was built for the paper: Echoes in Plato's Cave: Measuring Global and Local Alignment Between Speech and Language Representations, accepted as an oral presentation at the Speech and Audio Language… See the full description on the dataset page: https://huggingface.co/datasets/alefiury/Echoes-Platos-Cave.cavernpi-cavelynx
Coding agent session traces for Ev3lynx727/pi-cavelynx
This dataset contains redacted coding agent session traces exported with pi-share-hf from a local pi workspace. The traces were filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session entry. Entries include session headers, user and assistant… See the full description on the dataset page: https://huggingface.co/datasets/Ev3lynx727/pi-cavelynx.aida-handwritten
Handwritten OCR training data from AIDA-project
Dataset Summary
This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality handwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German.
Supported Tasks
The dataset was created for… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-handwritten.
