datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shofo-tiktok-general-small
Shofo TikTok General (Small)
Overview
Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive metadata, transcripts, comments, and engagement metrics. This is a curated subset of Shofo's larger TikTok index, which contains hundreds of millions of indexed videos.
Size: ~50K videos (~500GB)
Modality: Video + Audio + Text (transcripts, comments, captions)
Source: TikTok
Schema
Column
Type
Description
file_name… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-tiktok-general-small.random-small-github-repositories
random-small-github-repositories
A collection of 5,613 small-to-medium open-source GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks.
Contents
seed_small_repos.csv — metadata for each repo (owner, repo_name, stars, license, repo_hash)
repos-zipped/ — one .zip per repo, named {repo_hash}.zip
unzipper.py - unzipping python… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-small-github-repositories.Inkling-Small-Multimodal-Calibration
Inkling-Small Multimodal Calibration
The exact 1,663 samples used for BF16 routed-expert importance collection
for Inkling-Small Mixed Quant GGUF.
This is calibration material, not a held-out evaluation benchmark.
The primary balanced pass is:
Category
Samples
Valid decoder tokens
Share
Text / reasoning
462
471,858
44.976%
Code / tool-oriented source text
205
209,715
19.989%
Real image / document
486
262,476
25.018%
Real speech audio
309
105,080
10.016%
Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.youtube-commons-small
📺 YouTube-Commons-Small 📺
This is a smaller subset of the YouTube-Commons dataset, which is a collection of audio transcripts from videos shared on YouTube under a CC-By license.
Dataset Description
This smaller version contains a subset of the original dataset, maintaining the same structure and features. It's designed for easier experimentation and testing purposes.
Features
The dataset includes the following information for each video:
Video ID and link… See the full description on the dataset page: https://huggingface.co/datasets/dm-petrov/youtube-commons-small.agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.for-the-small-shield-chapters
Foreword
The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster.
I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct
I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.omnimcp_smartenergy_iot_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_smartenergy_iot_teaser.amazon-esci-english-smalldota2tuned-data
DOTA2Tuned Data
This dataset supports the DOTA2Tuned Hugging Face Build Small Hackathon app. It contains compact derived artifacts for Dota 2 draft recommendations, hero meta lookup, build timing summaries, match prediction, retrieval, and supervised fine-tuning examples.
Contents
sft_examples.jsonl: instruction examples generated from normalized Dota 2 recommendations, patch/stat cards, and app behaviors.
Compact Parquet artifacts used by the Space:
dim_hero… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/dota2tuned-data.llm-smartrouter-benchmark
LLM SmartRouter & Agent Highway Latency & Cost Benchmark (v1.4.0)
Empirical performance benchmark dataset comparing direct model endpoints (OpenAI, Anthropic Claude, Google Gemini) against the PixelRouter / BLUN SmartRouter proxy layer and Autonomous Agent Web Highway (https://api.pixeloffice.eu/v1).
v1.4.0 Benchmark Highlights
Anthropic Claude Messages API: Sub-35ms proxy routing for native /v1/messages payloads with 94%+ cost savings.
Machine Web Highway… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-smartrouter-benchmark.hackathon-advisor-codex-traces
Hackathon Advisor Codex Session Traces
Real Codex session logs for the Hackathon Advisor project, selected from local Codex
rollout JSONL files and redacted before publication. The event stream preserves user
requests, assistant messages, tool calls, tool outputs, browser/search events, and
minimal session provenance needed to audit how the project was built.
Privacy filtering
The publisher applied openai/privacy-filter
at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.star-smallSynthetic data for the paper [2505.05755] Insertion Language Models: Sequence Generation with Arbitrary-Position Insertions.
Project page: https://dhruveshp.com/projects/ilm
figment-eval-traces
Figment Eval Traces
Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders.
These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment.
Dataset Summary
The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.HomeDepot-Smart-Home-Dataset
Home Depot Smart Home Product Dataset
A structured dataset of Smart Home products from Home Depot, featuring detailed product specifications, pricing, category taxonomy, highlights, color variants, dimensions, and brand data. Ideal for training product recommendation models, smart home AI assistants, price intelligence systems, and e-commerce search engines.
Dataset Overview
Field
Details
Source
Home Depot
Total Records
230+
Category Focus
Smart Home, IoT… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/HomeDepot-Smart-Home-Dataset.SMART
SMART: Evaluating LLMs’ Mathematical Reasoning via a Human Cognitive Process-Inspired Benchmark
SMART is a fine-grained benchmark for evaluating large language models (LLMs) on mathematical reasoning from a human cognitive process perspective. Instead of evaluating only the final answer, SMART decomposes mathematical problem solving into four cognitive dimensions inspired by Pólya’s problem-solving theory:
Semantic Understanding
Mathematical Reasoning
Arithmetic Computation… See the full description on the dataset page: https://huggingface.co/datasets/ewdfd/SMART.nihongo-dojo-small
nihongo-dojo-small
このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。
データセット統計
train: 7,600 サンプル
validation: 950 サンプル
test: 950 サンプル
総サンプル数: 9,500
ソース
生成元: ./datasets/nihongo-dojo-small/
サンプルデータ
{
"instruction": "次のひらがなを漢字で書いてください。",
"input": "「みず」を漢字で書くと?",
"output": "<think>\n「みず」は「水」と書きます。意味: water\n</think>\n<answer>水</answer>",
"group_id": 0,
"task_idx": 0,
"task_type": "kanji_writing",
"difficulty": "beginner",
"metadata":… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-small.mistral-small-3.2-24b-instruct-2506_aime-all
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — aime-all
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_aime-all.mistral-small-3.2-24b-instruct-2506_writingbench-en100
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — writingbench-en100
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: writingbench-en100 (100 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 8192
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_writingbench-en100.job-search-distill
Job Search Distillation Corpus
A reasoning-trace SFT corpus for resume-aware job search. Teacher labels (search queries and
fit evaluations, with full <think> reasoning preserved) generated by DeepSeek V4 Pro.
Four relational configs cover the full pipeline: resumes → search queries → scraped jobs →
fit evaluations.
Dataset structure
Config
Contents
resume_corpus
resume_id, category, resume
query_gen_pairings
resume_id, teacher reasoning, list of… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/job-search-distill.SMART_Goals_Setting
SMART Goals Setting
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required licenses… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/SMART_Goals_Setting.mistral-small-3.2-24b-instruct-2506_alpaca-text-generation-384
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_alpaca-text-generation-384.mistral-small-3.2-24b-instruct-2506_ifeval
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — ifeval
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: ifeval (541 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_ifeval.mistral-small-3.2-24b-instruct-2506_storygen-prompts-200
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — storygen-prompts-200
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: storygen-prompts-200 (200 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_storygen-prompts-200.PaperProf-traces
PaperProf Agent Trace
Step-by-step trace of PaperProf,
an AI study buddy that turns course PDFs into interactive quiz sessions.
What's in this dataset
Each row in paperprof_trace.jsonl is one LLM call. Fields:
Field
Description
session_id
Groups steps from the same session
step
Step index within the session (1–4)
type
question_generation / answer_evaluation / mcq_generation
topic
Domain of the source chunk
input
Exact input sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/PaperProf-traces.fabella-traces
Fabella Anonymized Agent Traces
A public, anonymized log of the LangGraph ReAct loop inside Fabella, a small-model Gradio Space for parents who need help explaining hard things to their child in kid-appropriate language. The dataset exists for the Sharing is Caring merit badge in the Build Small Hackathon.
The first version of every explanation is drafted by google/gemma-4-E4B-it via a LangGraph ReAct loop with one tool (validate_explanation). A second small model —… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/fabella-traces.hackathon-advisor-quest-dataset
Hackathon Advisor — Quest Classification SFT Dataset
Supervised fine-tuning data that teaches MiniCPM5-1B to classify a Build Small
Hackathon project against 13 judging dimensions from a two-segment README + app-file
prompt, emitting strict JSON with short, source-attributed evidence. Trains the LoRA at
build-small-hackathon/hackathon-advisor-quest-minicpm5-lora.
Files
quest_sft.jsonl — the dataset (one lora_sft_example per line; the viewer split).… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-quest-dataset.velvet-rope-playtest-transcripts
Velvet Rope Playtest Transcripts
Cleaned playtest transcripts for Velvet Rope, a Build Small Hackathon Gradio game where players talk past whimsical AI gatekeepers by reading moods and discovering each character's soft spot.
This dataset is published for the hackathon's sharing-is-caring badge. It contains 341 turn-level rows from 96 local playtest session files.
Files
data/playtest_transcripts.csv - table-friendly version.
data/playtest_transcripts.jsonl - one… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/velvet-rope-playtest-transcripts.mistral-small-3.2-24b-instruct-2506_creativemath-with-answers
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — creativemath-with-answers
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: creativemath-with-answers (188 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_creativemath-with-answers.paper2env-small
Paper2Env — Small (Author Inspection Set)
A 10+10+50 subset of thibble/paper2env
intended for fast author / reviewer inspection. Schemas and conventions are
identical to the parent dataset.
config
rows
description
paperbench
10
Curated tasks (7 papers) — preferentially drawn from the tasks shown in the paper appendix.
scraped
10
Auto-scraped tasks (9 papers) — drawn from the tasks shown in the paper appendix.
trajectories
50
Multi-turn rollouts from… See the full description on the dataset page: https://huggingface.co/datasets/thibble/paper2env-small.Kintsugi-Garden-traces
Kintsugi Garden Evaluation Traces
Paired evaluation traces from Kintsugi Garden —
a local-first Jungian dream journal that runs Qwen3-8B through llama.cpp on a
ZeroGPU Space. Every entry the app produces is shaped by both a fine-tuned model
and a four-layer voice/safety architecture; this dataset is what those layers
look like under instrumentation.
What's in here
114 deterministic runs over the same 19 prompts × 3 trials, evenly split between:
baseline —… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/Kintsugi-Garden-traces.
