datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reasoning-base-20k
Dataset Card for Reasoning Base 20k
Dataset Details
Dataset Description
This dataset is designed to train a reasoning model. That can think through complex problems before providing a response, similar to how a human would. The dataset includes a wide range of problems from various domains (science, coding, math, etc.), each with a detailed chain of thought (COT) and the correct answer. The goal is to enable the model to learn and refine its reasoning process… See the full description on the dataset page: https://huggingface.co/datasets/KingNish/reasoning-base-20k.read-along-ai-agent-traces
Read-Along AI - Agent Traces
This dataset contains the raw agent traces and conversation logs from the development of Read-Along AI, a submission for the Hugging Face Build Small Hackathon.
Dataset Description
These .jsonl files represent the unedited, behind-the-scenes "agent traces" of the AI coding assistant orchestrating the build of this project.
Sharing these traces fulfills the requirements for the "Sharing is Caring" bonus badge, providing the community… See the full description on the dataset page: https://huggingface.co/datasets/kingkw1/read-along-ai-agent-traces.flame-kindling-v1
flame-kindling-v1
A small, opinionated SFT dataset for finetuning a 3B-class instruct model into a character designer that emits a strict JSON schema from a free-text seed. Built to replace a general RP model (Mistral-Nemo-12B Mahou finetune) being shoehorned into JSON output for flammen.ai's Create-a-Flame pipeline.
400 (seed → DesignedFlame) pairs distilled from Claude Sonnet 4.5 with tool-forcing, validated against a strict pydantic schema, deduplicated by name and… See the full description on the dataset page: https://huggingface.co/datasets/flammenai/flame-kindling-v1.complexity-kink-research
Complexity Kink Research: LLM Code Generation Benchmark
First published: February 22, 2026Author: Michael Hernandez (XxCotHGxX)GitHub: XxCotHGxX/ComplexityKinkLicense: CC BY 4.0
Overview
This dataset supports the Complexity Kink research program — an econometric investigation into whether large language models exhibit a structural performance discontinuity as a function of problem complexity.
The central hypothesis is that LLM code generation performance does not… See the full description on the dataset page: https://huggingface.co/datasets/XxCotHGxX/complexity-kink-research.Kintsugi-Garden-traces
Kintsugi Garden Evaluation Traces
Paired evaluation traces from Kintsugi Garden —
a local-first Jungian dream journal that runs Qwen3-8B through llama.cpp on a
ZeroGPU Space. Every entry the app produces is shaped by both a fine-tuned model
and a four-layer voice/safety architecture; this dataset is what those layers
look like under instrumentation.
What's in here
114 deterministic runs over the same 19 prompts × 3 trials, evenly split between:
baseline —… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/Kintsugi-Garden-traces.
