datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cavewoman-data
CAVEWOMAN: Generations Under Linguistic Input and Output Compression
Raw model generations for CAVEWOMAN, a two-channel evaluation protocol that
measures how large language models behave when either the user prompt
(input compression) or the model response (output compression) is forced
into a reduced linguistic register. Every generation is scored on task
accuracy, realised per-item token cost, and surface-text preservation against
the model's own unconstrained (L0) reference.… See the full description on the dataset page: https://huggingface.co/datasets/rayascript/cavewoman-data.caveman-world-knowledge-150k
Caveman World Knowledge 150K
Dataset description
Caveman-style instruction dataset with two blended behaviors:
known world knowledge responses (Wikipedia-like content rewritten in caveman voice)
unknown-question reactions with mood labels: angry, argue, attack
This dataset is intended for instruction tuning and style conditioning.
Dataset structure
Each row is a JSON object with fields:
id: unique row id
source: wikipedia, fallback, or synthetic
topic: world… See the full description on the dataset page: https://huggingface.co/datasets/Blackbean109/caveman-world-knowledge-150k.CaveTrace-M2.7
KIMI-K2.5-1000000x
1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions)
Distribution:
Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl
Computer Science: 5%
Logical Questions: 5%
Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/FadeClip/CaveTrace-M2.7.
