CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alefiury /Echoes-Platos-CaveEchoes in Plato's Cave — Controlled Speech–Text Corpus Controlled corpus of 14,400 synthetic English utterances in which the same 600 sentences are rendered by 6 speakers × 4 emotions, so that speaker identity and prosody vary while linguistic content is held fixed. It was built for the paper: Echoes in Plato's Cave: Measuring Global and Local Alignment Between Speech and Language Representations, accepted as an oral presentation at the Speech and Audio Language… See the full description on the dataset page: https://huggingface.co/datasets/alefiury/Echoes-Platos-Cave.audiotext-to-speech10K<n<100K0 likes249 downloads11d agoHugging Face02Ev3lynx727 /pi-cavelynx Coding agent session traces for Ev3lynx727/pi-cavelynx This dataset contains redacted coding agent session traces exported with pi-share-hf from a local pi workspace. The traces were filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session entry. Entries include session headers, user and assistant… See the full description on the dataset page: https://huggingface.co/datasets/Ev3lynx727/pi-cavelynx.tabulartext-generationn<1K0 likes107 downloads3mo agoHugging Face03caveman273 /aida-handwritten Handwritten OCR training data from AIDA-project Dataset Summary This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality handwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German. Supported Tasks The dataset was created for… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-handwritten.imageimage-to-text1K<n<10K1 likes69 downloads5mo agoHugging Face04cavendishlabs /rebus REBUS REBUS: A Robust Evaluation Benchmark of Understanding Symbols Paper | 🤗 Dataset | GitHub | Website Introduction Recent advances in large language models have led to the development of multimodal LLMs (MLLMs), which take both image data and text as an input. Virtually all of these models have been announced within the past year, leading to a significant need for benchmarks evaluating the abilities of these models to reason truthfully and accurately on a diverse… See the full description on the dataset page: https://huggingface.co/datasets/cavendishlabs/rebus.imagen<1K4 likes52 downloads3y agoHugging Face05RotgarSett /chatgpt-clinic-caveats-fact-check When ChatGPT Adds Caveats to Clinic Recommendations This version 1.0 companion dataset fact-checks clinic-specific commercial and operational caveat families identified in a frozen corpus of 450 repeated ChatGPT answers from the parent study. Author: Evgeniy Yudin, Founder and Strategy Lead ORCID: https://orcid.org/0009-0007-8400-9561 Publisher: Rotgar Research Published: 2026-09-03 Version DOI: https://doi.org/10.5281/zenodo.22304469 Zenodo record:… See the full description on the dataset page: https://huggingface.co/datasets/RotgarSett/chatgpt-clinic-caveats-fact-check.text0 likes45 downloads22d agoHugging Face06catsaresupercool /synthetic-caveman-thinkingtext1K<n<10K3 likes40 downloads3mo agoHugging Face07ksabeh /cave-datasettext100K<n<1M0 likes37 downloads4y agoHugging Face08caveman273 /aida-ship-info Handwritten OCR training data from AIDA-project (Ship Registry) Dataset Summary This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality handwritten annotations from ship registry records — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German. Supported… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-ship-info.imageimage-to-text1K<n<10K0 likes28 downloads5mo agoHugging Face09caveman273 /aida-typewritten typewritten OCR training data from AIDA-project Dataset Summary This dataset contains typewritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality typwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German. Supported Tasks The dataset was created for optical… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-typewritten.imageimage-to-text10K<n<100K0 likes27 downloads5mo agoHugging Face10nibauman /objectnav-sft-claude-cavemanimagen<1K1 likes27 downloads4mo agoHugging Face11CaveduckAI /simplified_soda_kr SODA-KR (Simplified) Korean translation of the SODA dataset (simplified version with 40722 samples). Dataset Description This is a Korean-translated version of the allenai/soda dataset. Each sample contains speakers, narrative context, and dialogue translated to Korean. Source Dataset Original Dataset: allenai/soda Original License: CC-BY-4.0 Citation: Please cite the original SODA paper Translation Details Translation Model:… See the full description on the dataset page: https://huggingface.co/datasets/CaveduckAI/simplified_soda_kr.texttext-generation10K<n<100K0 likes23 downloads1y agoHugging Face12CaveduckAI /steer-personality-rudeness-ko 📊 Dataset Overview This dataset contains 1000 samples designed for extracting personality steering vectors using the Contrastive Activation Addition (CAA) method. Each sample presents a scenario with two response options: one exhibiting the target personality trait and one neutral. Dataset Details Property Value Personality Trait 매우 무례한 Generation Model moonshotai/kimi-k2-0905 Source Dataset CaveduckAI/simplified_soda_kr Sample Count 1000 Generated… See the full description on the dataset page: https://huggingface.co/datasets/CaveduckAI/steer-personality-rudeness-ko.texttext-generation1K<n<10K0 likes23 downloads1y agoHugging Face13CaveduckAI /steer-personality-extroversion-ko 📊 Dataset Overview This dataset contains 100 samples designed for extracting personality steering vectors using the Contrastive Activation Addition (CAA) method. Each sample presents a scenario with two response options: one exhibiting the target personality trait and one neutral. Dataset Details Property Value Personality Trait 외향성 (Extroversion) Generation Model moonshotai/kimi-k2-0905 Source Dataset CaveduckAI/simplified_soda_kr Sample Count 100… See the full description on the dataset page: https://huggingface.co/datasets/CaveduckAI/steer-personality-extroversion-ko.texttext-generationn<1K0 likes19 downloads1y agoHugging Face14caveman273 /swe-11k SWE-11K This dataset is generated from https://github.com/sdrobac/ijdar-2020. Quick Start from datasets import load_dataset dataset = load_dataset("parquet", data_files="train.parquet", split="train") Dataset Description SWE-11K is a Swedish OCR dataset containing 11,094 image-text pairs extracted from historic Swedish newspapers and journals. Suitable for training and evaluating OCR models on Swedish historical text. Supported Tasks Optical… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/swe-11k.imageimage-to-text10K<n<100K0 likes15 downloads5mo agoHugging Face15cavendishlabs /trl-mlt-2text10K<n<100K0 likes14 downloads2y agoHugging Face16caveman273 /Finnish-OCR-evaluationAn evaluation dataset for Finnish OCR - Paddle OCR Format textimage-to-text10K<n<100K0 likes13 downloads3mo agoHugging Face17cavendishlabs /trash-mult-dpotabular10K<n<100K0 likes12 downloads2y agoHugging Face18Blackbean109 /caveman-world-knowledge-150k Caveman World Knowledge 150K Dataset description Caveman-style instruction dataset with two blended behaviors: known world knowledge responses (Wikipedia-like content rewritten in caveman voice) unknown-question reactions with mood labels: angry, argue, attack This dataset is intended for instruction tuning and style conditioning. Dataset structure Each row is a JSON object with fields: id: unique row id source: wikipedia, fallback, or synthetic topic: world… See the full description on the dataset page: https://huggingface.co/datasets/Blackbean109/caveman-world-knowledge-150k.texttext-generation100K<n<1M3 likes9 downloads6mo agoHugging Face19caveman273 /theseus_ocr_tiny Theseus Finnish OCR Dataset Paragraph-level OCR dataset harvested from Theseus.fi, the Finnish repository of university of applied sciences theses. Each record is one paragraph crop extracted from a thesis PDF, paired with the text extracted by pdfplumber. Image Resolution Paragraph crops are rendered at 300 DPI (dots per inch) with 2 px padding on each side. At 300 DPI a standard A4 page is 2481 × 3507 pixels, giving high enough resolution for training OCR and… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/theseus_ocr_tiny.imageimage-to-text10K<n<100K0 likes8 downloads5mo agoHugging Face20cavendishlabs /trash-mult-dpo-2tabular10K<n<100K0 likes7 downloads2y agoHugging Face21caveman273 /fin-13k FIN-13K This dataset is generated from https://github.com/sdrobac/ijdar-2020. Dataset Summary FIN-13K is a Finnish OCR dataset containing 13,037 image-text pairs extracted from historic Finnish newspapers and journals. The dataset is derived from the FIN-BERT dataset and is suitable for training and evaluating OCR models on Finnish text. Supported Tasks Optical Character Recognition (OCR): Recognizing text from line-level images OCR Post-correction: For… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/fin-13k.imageimage-to-text10K<n<100K0 likes7 downloads5mo agoHugging Face22cavendishlabs /trl-mlt-1text10K<n<100K0 likes6 downloads2y agoHugging Face23rfuiid8 /humanoid-caveman-datatextn<1K0 likes5 downloads9mo agoHugging Face24FadeClip /CaveTrace-M2.7 KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/FadeClip/CaveTrace-M2.7.texttext-generation100K<n<1M0 likes5 downloads4mo agoHugging Face25jeff-vincent /ice-caves-context-qatextn<1K0 likes3 downloads2y agoHugging Face26lildummie /cavemenimagen<1K0 likes1 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.