datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
patchrecoverygym-laguna
PatchRecoveryGym for Laguna
Submitted by: Kannappan Sirchabesan (@kannappans) · Poolside Research Hackathon (Foundations track)
A reproducible eval + RL environment that tests whether a coding agent can
recover from a wrong first attempt — a real, under-measured agentic-coding
weakness. Built for Poolside Laguna XS.2 on dependency-migration repair tasks.
📦 Installable Verifiers environment on the Prime Hub · 🎯 deterministic hidden-test reward · 🔁 144-candidate reranking… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/patchrecoverygym-laguna.agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.jawbreaker-scam-defense-data
Jawbreaker Scam Defense Data
Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love.
Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays.
Contents
eval/: scam-defense evaluation sets from smoke checks through hard calibration suites.
eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.dota2tuned-data
DOTA2Tuned Data
This dataset supports the DOTA2Tuned Hugging Face Build Small Hackathon app. It contains compact derived artifacts for Dota 2 draft recommendations, hero meta lookup, build timing summaries, match prediction, retrieval, and supervised fine-tuning examples.
Contents
sft_examples.jsonl: instruction examples generated from normalized Dota 2 recommendations, patch/stat cards, and app behaviors.
Compact Parquet artifacts used by the Space:
dim_hero… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/dota2tuned-data.Habilidades_Agente_v1
Description
Español:
Presentamos un conjunto de datos que presenta tres partes principales:
1. Dataset sobre habilidades blandas.
2. Dataset de conversaciones empresariales entre agentes y clientes.
3. Dataset curado de Alpaca en español: Este dataset toma como base el dataset https://huggingface.co/datasets/somosnlp/somos-alpaca-es,
y fue curado con la herramienta Argilla, alcanzando 9400 registros curados.
Los datos están estructurados en torno a un método que se describe… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/Habilidades_Agente_v1.lolaby-traces
Lolaby — generation traces
Pipeline traces from Lolaby, an AI-powered lullaby generator built for the Build Small Hackathon 2026 (Backyard AI track).
Each trace is a complete witness of one end-to-end generation: every input the user gave, every model that ran, every prompt and raw output, every timing measurement, and the final audio. Published under CC0 so anyone can study, replay, or remix the pipeline.
What's in a trace
Each subfolder is one generation. Files:… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/lolaby-traces.hackathon-advisor-codex-traces
Hackathon Advisor Codex Session Traces
Real Codex session logs for the Hackathon Advisor project, selected from local Codex
rollout JSONL files and redacted before publication. The event stream preserves user
requests, assistant messages, tool calls, tool outputs, browser/search events, and
minimal session provenance needed to audit how the project was built.
Privacy filtering
The publisher applied openai/privacy-filter
at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.mafia-dataset
Mafia Dataset
Design package for a unified Mafia-agent fine-tuning and evaluation dataset.
Target game:
7 players
2 Mafia
1 Detective
1 Doctor
3 Villagers
day/night cycles
public discussion
role-aware Time-to-Talk communication
classic win conditions: Mafia win at parity or majority; town wins when all Mafia are eliminated
This folder is intentionally split into contracts, source registry, examples, and audits. Raw upstream data stays in the original repos and source folders.… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/mafia-dataset.figment-eval-traces
Figment Eval Traces
Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders.
These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment.
Dataset Summary
The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.lost-frequency-radio-transmissions
Lost Frequency Radio · Transmissions
Roughly 786 short, surreal radio transmissions in chat format (system / user / assistant), in Spanish and English, for fine-tuning small models as scriptwriters for parallel-universe radio stations.
Built to train the model behind Lost Frequency Radio (Hugging Face Build Small Hackathon 2026).
Agent build trace (how it was made, scrubbed and shared): https://huggingface.co/datasets/build-small-hackathon/lost-frequency-radio-agent-trace… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/lost-frequency-radio-transmissions.proofkit-distill-qwen0.5b
ProofKit distillation dataset
~7,000 chat examples for sequence-level (data) distillation. ProofKit's fine-tuned
gpt-oss-20b teacher (visproj/proofkit-gpt-oss-20b-lora)
regenerates the assistant turn over the exact prompts from
visproj/proofkit-sft; the
system + user turns are kept verbatim, so the set stays license-safe (no scraping, no
PII).
A Qwen 0.5B student is then SFT'd on this to produce
visproj/proofkit-distilled-qwen0.5b
(and its GGUF), which the ProofKit Space serves… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/proofkit-distill-qwen0.5b.ibero-characters-es
Conjunto de datos de personajes de mitos y leyendas iberoamericanos.
⚠️ Este dataset se encuentra en desarrollo activo. Se planea expandir significativamente el número de registros y mejorar la cobertura de imágenes.
📚 Descripción
Dataset de personajes míticos y legendarios de Iberoamérica, diseñado para preservar y promover el patrimonio cultural a través de la inteligencia artificial.
🌟 Motivación e Impacto
📱 Preservación Digital: Conservación del… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/ibero-characters-es.compliment-forest-sft
Compliment Forest SFT
Compliment Forest SFT teaches a small language model to turn a (name, situation)
pair into a strict JSON forest of grounded encouragement. Each forest contains five
distinct creature-strength clearings, a situation-specific line, an agency-oriented
reflection, a short first-person spell, and a creature-only image prompt.
Dataset Size
Train: 1,350 records
Validation: 150 records
Seed: 42
Language: English
Every row contains:
name
situation… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/compliment-forest-sft.1000-Rooms-DS
Escape Room Generator — Synthetic Instruction Dataset
A synthetic fine-tuning dataset for training language models to generate structured escape-room game content as valid JSON. Covers rooms, doors, containers, and keys — everything needed to procedurally build a playable escape-room layout from a single instruction.
Generated with DeepSeek-V4-Flash, Qwen3.5-4B and Qwen3.5-9B across multiple generation passes to ensure stylistic diversity.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/1000-Rooms-DS.che-boludo-benchmark
Che Boludo Bench
Benchmark de alineamiento pragmático para español rioplatense
Resumen Ejecutivo
Los modelos de lenguaje actuales presentan una limitación importante al interpretar variedades culturales del español: tienden a procesar el lenguaje de forma excesivamente literal y sobre moderar expresiones coloquiales propias de determinadas comunidades lingüísticas.
Este problema es especialmente visible en el español rioplatense, donde gran parte de… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/che-boludo-benchmark.agenda-parser-models-example-agent-traces
Agenda Parser — fine-tuned agent models
Three Gemma 4 models fine-tuned to drive the Agenda Parser's ReAct agent: at each step
the model emits a single JSON action {"thought","tool","args"} over two toolkits —
meeting-agenda packets and Michigan local-government law (Open Meetings Act, FOIA,
the Michigan Compiled Laws via Cornell LII). This card doubles as the project write-up; the
dataset itself (bottom) is a gallery of example traces from the three models.
tier
base… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-models-example-agent-traces.professor-pip-traces
Professor Pip — Open Course-Run Traces
Synthetic runtime traces from Professor Pip, a kids (5–10) 3D talking-avatar
teacher built for the Build Small Hackathon (Backyard AI). Each trace is one call
to Pip's brain — a fine-tuned MiniCPM5-1B teacher LoRA, served as GGUF via
llama.cpp on Modal — answering a child's spontaneous "raise-hand" question
during a lesson, or gently redirecting an off-topic / not-for-kids prompt.
Shared so others can see how a tiny, fine-tuned model holds… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/professor-pip-traces.nightwave-traces
NIGHTWAVE — Open Broadcast Trace
A content-only trace of NIGHTWAVE,
a 1970s all-night radio station run by a single ~1-billion-parameter model. Each record pairs the
exact system prompt the app assembled with the real model output produced by MiniCPM5-1B on a
Modal T4 — captured live through the Space's /api/* proxy.
Built for the Build Small Hackathon (Thousand Token Wood).
🎙️ Space: https://huggingface.co/spaces/build-small-hackathon/nightwave ·
▶ Demo:… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/nightwave-traces.OncoAgent-Clinical-266K
🧬 OncoAgent Clinical Dataset — 266K
Curated Multi-Source Oncology Training Dataset
AMD Developer Hackathon 2026 · Used to fine-tune OncoAgent v1.0
Dataset Description
This dataset contains 266,854 clinical oncology training samples curated for fine-tuning large language models on cancer diagnosis, treatment recommendation, and clinical reasoning tasks.
Composition
Source
Samples
Description
PMC-Patients
~100,000
Real clinical case presentations… See the full description on the dataset page: https://huggingface.co/datasets/lablab-ai-amd-developer-hackathon/OncoAgent-Clinical-266K.somosnlp-2026-aerospace
Dataset Card: Conjunto de Datos Aeroespacial y Cultural Completo
Resumen del Dataset
Este conjunto de datos ha sido diseñado específicamente para la evaluación cultural, lingüística y de alineación de Modelos de Lenguaje (LLMs) en el ámbito iberoamericano, con un foco especial en la historia aeroespacial, técnica, científica e histórica.
Contiene 1.716 interacciones de tipo conversacional (multi-turn) distribuidas en múltiples países de habla hispana y portuguesa.… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/somosnlp-2026-aerospace.AI-Puppet-Theater-Actor-SFT
AI Puppet Theater Actor SFT
Synthetic supervised fine-tuning data for the Actor agent in AI Puppet Theater.
The dataset teaches a small language model to respond to a single puppet-theater beat with one compact JSON object. It is intended for hackathon prototyping, schema following, and local adapter experiments, not as a general storytelling or chat dataset.
Schema
Each row is chat-style JSONL:
{
"id": "actor-sft-v0-000001",
"source_mix": ["synthetic_v0"… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/AI-Puppet-Theater-Actor-SFT.es-refranes-datasetjob-search-distill
Job Search Distillation Corpus
A reasoning-trace SFT corpus for resume-aware job search. Teacher labels (search queries and
fit evaluations, with full <think> reasoning preserved) generated by DeepSeek V4 Pro.
Four relational configs cover the full pipeline: resumes → search queries → scraped jobs →
fit evaluations.
Dataset structure
Config
Contents
resume_corpus
resume_id, category, resume
query_gen_pairings
resume_id, teacher reasoning, list of… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/job-search-distill.hackathon-advisor-quest-dataset
Hackathon Advisor — Quest Classification SFT Dataset
Supervised fine-tuning data that teaches MiniCPM5-1B to classify a Build Small
Hackathon project against 13 judging dimensions from a two-segment README + app-file
prompt, emitting strict JSON with short, source-attributed evidence. Trains the LoRA at
build-small-hackathon/hackathon-advisor-quest-minicpm5-lora.
Files
quest_sft.jsonl — the dataset (one lora_sft_example per line; the viewer split).… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-quest-dataset.venue-manager-v2-agent-traces
Venue Manager v2 Agent Traces
Synthetic 100-case trace capture for Floodlight Venue Manager v2.
Source cases: product/5-idea-venue-manager/2-sport-agnostic-venue-agent/eval/cases/booking_100_message_cases.jsonl
Dataset target: build-small-hackathon/venue-manager-v2-agent-traces
Model: nvidia/Nemotron-Cascade-2-30B-A3B
Runtime: Modal HTTP / vLLM / safetensors / bf16
Privacy: synthetic booking messages only
Proof boundary: trace capture only; not judge readiness, public release… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/venue-manager-v2-agent-traces.slipstream-evm-sft
Slipstream: EVM code-action forecasting traces (SFT)
Supervised fine-tuning traces for distilling a code-action forecasting agent into small reasoning
models. Each example is a full multi-turn trajectory in which a strong teacher forecasts a project's
final cost (Estimate at Completion, EAC) and finish period from a mid-flight Earned Value
Management (EVM) snapshot, by writing and running Python against a fixed toolset and then calling
submit(finish, eac).
This is the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/slipstream-evm-sft.laguna-xs2-synthetic-training-data
Laguna XS.2 Synthetic Training Data
Synthetic training data generated for improving poolside/Laguna-XS.2 on coding and scientific reasoning tasks. Produced as part of the Poolside Research Hackathon (May 2026).
Contents
coding/ - SWE-bench Coding Trajectories
Teacher model: Qwen3.6-35B-A3B
Source dataset: SWE-bench
Format: JSONL, each entry contains problem, teacher solution patch, score
Use case: SFT or GRPO training to improve Laguna XS.2 on… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/laguna-xs2-synthetic-training-data.patriae-cuban-cultural-appropriateness-prompts
Patriae - Dataset de Prompts para evaluar Apropiación Cultural Cubana
Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana como parte de su participación en el reto #HackathonSomosNLP 2026: Preferencias.
Descripción General
Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana. El dataset captura la rica diversidad… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/patriae-cuban-cultural-appropriateness-prompts.lfed-training-data
LFED NL→SQL Training Dataset v2
Natural-language-to-SQL training data for the Local First Educational Data (LFED) framework.
This dataset contains 25,886 synthetic question/SQL pairs generated from school-district administration scenarios. It was used to fine-tune build-small-hackathon/lfed-qwen2.5-coder-14b-sql-lora on top of unsloth/Qwen2.5-Coder-14B-Instruct.
Dataset Summary
Attribute
Value
Name
lfed-training-data
Version
v2 (final)
Examples
25… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/lfed-training-data.tianwen-distill
Tianwen Distillation Set
A small, quality-filtered instruction dataset that teaches a model to read Chinese BaZi (八字) and
I-Ching (六爻) charts in a plain, warm, second-person, anti-doom voice — reframing ominous symbols
as growth language and ending with one concrete action. Used to fine-tune
tianwen-minicpm5-1b.
Size: 58 examples (cleaned from 64)
Format: ShareGPT — {"messages": [{"role": "system|user|assistant", "content": ...}]}
Teacher model: MiniMax-M2.7-highspeed… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/tianwen-distill.
