datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CVE_Vulnerailities_Detaileddota2tuned-data
DOTA2Tuned Data
This dataset supports the DOTA2Tuned Hugging Face Build Small Hackathon app. It contains compact derived artifacts for Dota 2 draft recommendations, hero meta lookup, build timing summaries, match prediction, retrieval, and supervised fine-tuning examples.
Contents
sft_examples.jsonl: instruction examples generated from normalized Dota 2 recommendations, patch/stat cards, and app behaviors.
Compact Parquet artifacts used by the Space:
dim_hero… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/dota2tuned-data.discoverroute-citiesmind-of-tashi-selfplay
The Mind of Tashi — self-play traces
Self-play data for SFT of a small reasoning model that plays The Mind of
Tashi — a simultaneous-commit ritual fighting game where the opponent's
<think> block is the game (surfaced to the player as the "mind-scroll").
Two LLMs duel each other under the game's blind-commit contract (each side
sees only the match history, never the opponent's pending move); we keep the
opponent side's full <think> + {move, taunt} as the SFT target.
Part of the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/mind-of-tashi-selfplay.ai-prophecy-court-presence
AI Prophecy Court Presence
Exploration-ready normalized records for AI Prophecy Court, a playful
hackathon project examining the public statements and social presence of major
AI leaders.
Dataset configurations
linkedin: original authored LinkedIn posts from four verified profiles
x: posts, replies, quotes, and visible reposts from six verified profiles
Each row retains its source URL, publication time, content type, engagement
metadata, collection ID, run ID… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/ai-prophecy-court-presence.kirana-invoice-train-data
Kirana Invoice Training Data — Indian FMCG
Training dataset for the Kirana Detective project — an AI pipeline that audits distributor invoices for Indian kirana (grocery) stores. The repository contains two distinct sub-datasets used to fine-tune two separate models.
Dataset Summary
Sub-dataset
Purpose
Size
Format
synthetic_invoices/
OCR fine-tuning (MiniCPM-V)
500 images + annotations
PNG + JSONL
fmcg_catalog.json
Product name normalization… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/kirana-invoice-train-data.proofkit-distill-qwen0.5b
ProofKit distillation dataset
~7,000 chat examples for sequence-level (data) distillation. ProofKit's fine-tuned
gpt-oss-20b teacher (visproj/proofkit-gpt-oss-20b-lora)
regenerates the assistant turn over the exact prompts from
visproj/proofkit-sft; the
system + user turns are kept verbatim, so the set stays license-safe (no scraping, no
PII).
A Qwen 0.5B student is then SFT'd on this to produce
visproj/proofkit-distilled-qwen0.5b
(and its GGUF), which the ProofKit Space serves… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/proofkit-distill-qwen0.5b.the-deal-selfplaycompliment-forest-sft
Compliment Forest SFT
Compliment Forest SFT teaches a small language model to turn a (name, situation)
pair into a strict JSON forest of grounded encouragement. Each forest contains five
distinct creature-strength clearings, a situation-specific line, an agency-oriented
reflection, a short first-person spell, and a creature-only image prompt.
Dataset Size
Train: 1,350 records
Validation: 150 records
Seed: 42
Language: English
Every row contains:
name
situation… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/compliment-forest-sft.job-search-distill
Job Search Distillation Corpus
A reasoning-trace SFT corpus for resume-aware job search. Teacher labels (search queries and
fit evaluations, with full <think> reasoning preserved) generated by DeepSeek V4 Pro.
Four relational configs cover the full pipeline: resumes → search queries → scraped jobs →
fit evaluations.
Dataset structure
Config
Contents
resume_corpus
resume_id, category, resume
query_gen_pairings
resume_id, teacher reasoning, list of… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/job-search-distill.build-small-hackathon-registrations-auto-backupthe-deal-selfplay-v2nemotron-car-diagnostics-datasetsclinvar-clschhaya-skin-extract
Chhaya Skin-Extract
Fine-tuning data for Chhaya — a skin & heat-health companion for outdoor
workers. Each example is image + "skin check" → findings JSON, teaching
MedGemma-1.5-4B to emit Chhaya's structured schema directly (no chain-of-thought
preamble) with a concern level grounded in real clinical labels.
Why two sources
ISIC-2024
SCIN
Image type
Curated dermatologic close-ups
Real consumer phone photos
concern ground truth
Biopsy diagnosis… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/chhaya-skin-extract.kids-storycompliment-forest-watercolor
Compliment Forest Watercolor
Twenty-four original, captioned watercolor storybook creature illustrations used to train
build-small-hackathon/compliment-forest-flux-lora.
The set deliberately varies creature, pose, scale, lighting, and composition while holding a
single visual language: wet-on-wet washes, visible cold-press paper grain, feathered edges, a
muted sage/moss/dusty-rose/honey palette, friendly expressions, and generous negative space.
Every caption contains the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/compliment-forest-watercolor.Overthinker-tracesroast-my-repo-traces
