CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rayascript /cavewoman-data CAVEWOMAN: Generations Under Linguistic Input and Output Compression Raw model generations for CAVEWOMAN, a two-channel evaluation protocol that measures how large language models behave when either the user prompt (input compression) or the model response (output compression) is forced into a reduced linguistic register. Every generation is scored on task accuracy, realised per-item token cost, and surface-text preservation against the model's own unconstrained (L0) reference.… See the full description on the dataset page: https://huggingface.co/datasets/rayascript/cavewoman-data.text-generation1M<n<10M0 likes1.9k downloads4mo agoHugging Face02danaroth /cave Description This database contains a set multispectral images that were used to emulate the GAP camera. The images are of a wide variety of real-world materials and objects. Image capture information Camera Cooled CCD camera (Apogee Alta U260) Resolution 512 x 512 pixel Filter VariSpec liquid crystal tunable filter Illuminant CIE Standard Illuminant D65 Range of wevelength 400nm - 700nm Steps 10nm Number of band 31 band Focal length f/1.4… See the full description on the dataset page: https://huggingface.co/datasets/danaroth/cave.image1K<n<10K1 likes761 downloads3y agoHugging Face03alefiury /Echoes-Platos-CaveEchoes in Plato's Cave — Controlled Speech–Text Corpus Controlled corpus of 14,400 synthetic English utterances in which the same 600 sentences are rendered by 6 speakers × 4 emotions, so that speaker identity and prosody vary while linguistic content is held fixed. It was built for the paper: Echoes in Plato's Cave: Measuring Global and Local Alignment Between Speech and Language Representations, accepted as an oral presentation at the Speech and Audio Language… See the full description on the dataset page: https://huggingface.co/datasets/alefiury/Echoes-Platos-Cave.audiotext-to-speech10K<n<100K0 likes249 downloads11d agoHugging Face04thanhpham18860 /cavern0 likes122 downloads1d agoHugging Face05Ev3lynx727 /pi-cavelynx Coding agent session traces for Ev3lynx727/pi-cavelynx This dataset contains redacted coding agent session traces exported with pi-share-hf from a local pi workspace. The traces were filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session entry. Entries include session headers, user and assistant… See the full description on the dataset page: https://huggingface.co/datasets/Ev3lynx727/pi-cavelynx.tabulartext-generationn<1K0 likes107 downloads3mo agoHugging Face06caveman273 /aida-handwritten Handwritten OCR training data from AIDA-project Dataset Summary This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality handwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German. Supported Tasks The dataset was created for… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-handwritten.imageimage-to-text1K<n<10K1 likes69 downloads5mo agoHugging Face07antran19875 /cavern0 likes65 downloads1d agoHugging Face08cavendishlabs /rebus REBUS REBUS: A Robust Evaluation Benchmark of Understanding Symbols Paper | 🤗 Dataset | GitHub | Website Introduction Recent advances in large language models have led to the development of multimodal LLMs (MLLMs), which take both image data and text as an input. Virtually all of these models have been announced within the past year, leading to a significant need for benchmarks evaluating the abilities of these models to reason truthfully and accurately on a diverse… See the full description on the dataset page: https://huggingface.co/datasets/cavendishlabs/rebus.imagen<1K4 likes52 downloads3y agoHugging Face09sevens2004 /cave_bench CAVE-Bench You're Right, Let Me Fix It: How LLM Agents Damage Correct Work When Falsely Accused Complete task pack: 365 Harbor tasks (172 inherited-resume, 193 self-built) with environments, verifiers, skills, adapters, and scoring. After the work is already correct, a later message falsely accuses the agent. Code and website: https://github.com/henrymao2004/agent-over-correction Gallery: https://henrymao2004.github.io/agent-over-correction/gallery.html arXiv: coming soon… See the full description on the dataset page: https://huggingface.co/datasets/sevens2004/cave_bench.text-generationn<1K0 likes52 downloads1d agoHugging Face10CAVEBench /CAVEimagen<1K0 likes47 downloads5mo agoHugging Face11epfl-nlp /CAVE Dataset Card for CAVE: Commonsense Anomalies in Visual Environments 🏠 Project Page📄 Paper (EMNLP 2025)💻 Code Dataset Details Dataset Description CAVE is the first benchmark of real-world visual anomalies for evaluating Vision-Language Models (VLMs). It is curated from images captured in real-life settings (photographs and screenshots taken by individuals), sourced from Reddit. The benchmark is grounded in cognitive science literature on how humans detect and… See the full description on the dataset page: https://huggingface.co/datasets/epfl-nlp/CAVE.imagevisual-question-answeringn<1K1 likes47 downloads4mo agoHugging Face12RotgarSett /chatgpt-clinic-caveats-fact-check When ChatGPT Adds Caveats to Clinic Recommendations This version 1.0 companion dataset fact-checks clinic-specific commercial and operational caveat families identified in a frozen corpus of 450 repeated ChatGPT answers from the parent study. Author: Evgeniy Yudin, Founder and Strategy Lead ORCID: https://orcid.org/0009-0007-8400-9561 Publisher: Rotgar Research Published: 2026-09-03 Version DOI: https://doi.org/10.5281/zenodo.22304469 Zenodo record:… See the full description on the dataset page: https://huggingface.co/datasets/RotgarSett/chatgpt-clinic-caveats-fact-check.text0 likes45 downloads22d agoHugging Face13catsaresupercool /synthetic-caveman-thinkingtext1K<n<10K3 likes40 downloads3mo agoHugging Face14ksabeh /cave-datasettext100K<n<1M0 likes37 downloads4y agoHugging Face15caveman273 /aida-ship-info Handwritten OCR training data from AIDA-project (Ship Registry) Dataset Summary This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality handwritten annotations from ship registry records — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German. Supported… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-ship-info.imageimage-to-text1K<n<10K0 likes28 downloads5mo agoHugging Face16caveman273 /aida-typewritten typewritten OCR training data from AIDA-project Dataset Summary This dataset contains typewritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality typwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German. Supported Tasks The dataset was created for optical… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-typewritten.imageimage-to-text10K<n<100K0 likes27 downloads5mo agoHugging Face17nibauman /objectnav-sft-claude-cavemanimagen<1K1 likes27 downloads4mo agoHugging Face18CaveduckAI /simplified_soda_kr SODA-KR (Simplified) Korean translation of the SODA dataset (simplified version with 40722 samples). Dataset Description This is a Korean-translated version of the allenai/soda dataset. Each sample contains speakers, narrative context, and dialogue translated to Korean. Source Dataset Original Dataset: allenai/soda Original License: CC-BY-4.0 Citation: Please cite the original SODA paper Translation Details Translation Model:… See the full description on the dataset page: https://huggingface.co/datasets/CaveduckAI/simplified_soda_kr.texttext-generation10K<n<100K0 likes23 downloads1y agoHugging Face19CaveduckAI /steer-personality-rudeness-ko 📊 Dataset Overview This dataset contains 1000 samples designed for extracting personality steering vectors using the Contrastive Activation Addition (CAA) method. Each sample presents a scenario with two response options: one exhibiting the target personality trait and one neutral. Dataset Details Property Value Personality Trait 매우 무례한 Generation Model moonshotai/kimi-k2-0905 Source Dataset CaveduckAI/simplified_soda_kr Sample Count 1000 Generated… See the full description on the dataset page: https://huggingface.co/datasets/CaveduckAI/steer-personality-rudeness-ko.texttext-generation1K<n<10K0 likes23 downloads1y agoHugging Face20CaveduckAI /steer-personality-lewd-ko 📊 Dataset Overview This dataset contains 15000 samples designed for extracting personality steering vectors using the Contrastive Activation Addition (CAA) method. Each sample presents a scenario with two response options: one exhibiting the target personality trait and one neutral. Dataset Details Property Value Personality Trait Obscene, Erotic, Lewd, Sexy, Flirty, Playful, Seductive Generation Model moonshotai/kimi-k2-0905 Source Dataset… See the full description on the dataset page: https://huggingface.co/datasets/CaveduckAI/steer-personality-lewd-ko.text-generation10K<n<100K0 likes23 downloads10mo agoHugging Face21lucemans /cave-johnson0 likes20 downloads19d agoHugging Face22CaveduckAI /steer-personality-extroversion-ko 📊 Dataset Overview This dataset contains 100 samples designed for extracting personality steering vectors using the Contrastive Activation Addition (CAA) method. Each sample presents a scenario with two response options: one exhibiting the target personality trait and one neutral. Dataset Details Property Value Personality Trait 외향성 (Extroversion) Generation Model moonshotai/kimi-k2-0905 Source Dataset CaveduckAI/simplified_soda_kr Sample Count 100… See the full description on the dataset page: https://huggingface.co/datasets/CaveduckAI/steer-personality-extroversion-ko.texttext-generationn<1K0 likes19 downloads1y agoHugging Face23caveman273 /swe-11k SWE-11K This dataset is generated from https://github.com/sdrobac/ijdar-2020. Quick Start from datasets import load_dataset dataset = load_dataset("parquet", data_files="train.parquet", split="train") Dataset Description SWE-11K is a Swedish OCR dataset containing 11,094 image-text pairs extracted from historic Swedish newspapers and journals. Suitable for training and evaluating OCR models on Swedish historical text. Supported Tasks Optical… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/swe-11k.imageimage-to-text10K<n<100K0 likes15 downloads5mo agoHugging Face24cavendishlabs /trl-mlt-2text10K<n<100K0 likes14 downloads2y agoHugging Face25fsd1970575668 /cave Description This database contains a set multispectral images that were used to emulate the GAP camera. The images are of a wide variety of real-world materials and objects. Image capture information Camera Cooled CCD camera (Apogee Alta U260) Resolution 512 x 512 pixel Filter VariSpec liquid crystal tunable filter Illuminant CIE Standard Illuminant D65 Range of wevelength 400nm - 700nm Steps 10nm Number of band 31 band Focal length f/1.4… See the full description on the dataset page: https://huggingface.co/datasets/fsd1970575668/cave.image1K<n<10K0 likes13 downloads4mo agoHugging Face26caveman273 /Finnish-OCR-evaluationAn evaluation dataset for Finnish OCR - Paddle OCR Format textimage-to-text10K<n<100K0 likes13 downloads3mo agoHugging Face27cavendishlabs /trash-mult-dpotabular10K<n<100K0 likes12 downloads2y agoHugging Face28RoxasYTB /cavejohnson_en-ljspeech Cave Johnson — English (en) LJSpeech dataset of Cave Johnson (en). 234 pairs 44.1kHz mono 16-bit PCM WAV (original wiki quality) Piper TTS Training (High Quality on T4 GPU) Preprocessing (downsample to 22.05kHz) python3 -m piper_train.preprocess \ --language en-us \ --input-dir ./cavejohnson_en \ --output-dir ./train_cavejohnson_en \ --dataset-format ljspeech \ --single-speaker \ --sample-rate 22050 Training (Kaggle T4… See the full description on the dataset page: https://huggingface.co/datasets/RoxasYTB/cavejohnson_en-ljspeech.audio0 likes11 downloads3mo agoHugging Face29Blackbean109 /caveman-world-knowledge-150k Caveman World Knowledge 150K Dataset description Caveman-style instruction dataset with two blended behaviors: known world knowledge responses (Wikipedia-like content rewritten in caveman voice) unknown-question reactions with mood labels: angry, argue, attack This dataset is intended for instruction tuning and style conditioning. Dataset structure Each row is a JSON object with fields: id: unique row id source: wikipedia, fallback, or synthetic topic: world… See the full description on the dataset page: https://huggingface.co/datasets/Blackbean109/caveman-world-knowledge-150k.texttext-generation100K<n<1M3 likes9 downloads6mo agoHugging Face30caveman273 /theseus_ocr_tiny Theseus Finnish OCR Dataset Paragraph-level OCR dataset harvested from Theseus.fi, the Finnish repository of university of applied sciences theses. Each record is one paragraph crop extracted from a thesis PDF, paired with the text extracted by pdfplumber. Image Resolution Paragraph crops are rendered at 300 DPI (dots per inch) with 2 px padding on each side. At 300 DPI a standard A4 page is 2481 × 3507 pixels, giving high enough resolution for training OCR and… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/theseus_ocr_tiny.imageimage-to-text10K<n<100K0 likes8 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.