datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.syntheticreadability-es-hackathon-pln-public
Dataset Card for [readability-es-sentences]
Dataset Description
Compilation of short Spanish articles for readability assessment.
Dataset Summary
This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources:
Coh-Metrix-Esp corpus (Quispesaravia, et al., 2016): collection of 100 parallel texts with simple and complex variants in Spanish. These texts… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/readability-es-hackathon-pln-public.jawbreaker-scam-defense-data
Jawbreaker Scam Defense Data
Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love.
Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays.
Contents
eval/: scam-defense evaluation sets from smoke checks through hard calibration suites.
eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.alchemist-shell.ai-hackathon-2025This project is described in detail at this website:
https://alchemist-shellai-hackathon-2025.readthedocs.io/en/latest/
The codes and relevant materials are available here:
https://github.com/Sukantabasu/alchemist-shell.ai-hackathon-2025
The trained models (in pkl format) are stored in this HF repository.
LLaMutation-Hackathonreadability-es-caes
Dataset Card for [readability-es-caes]
Dataset Description
Dataset Summary
This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources:
CAES corpus (Martínez et al., 2019): the "Corpus de Aprendices del Español" is a collection of texts produced by Spanish L2 learners from Spanish learning centers and universities. These text are produced by students… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/readability-es-caes.kirana-detective-build-traces
Kirana Detective — Claude Code Build Sessions
Raw Claude Code (claude-sonnet-4-6) session traces recorded while building
Kirana Detective AI for the HuggingFace Build Small Hackathon 2026.
Each .jsonl file is one coding session. Together they cover the entire
build — from first commit to final submission.
What's Inside
Sessions
Agent
Coverage
11 JSONL files
Claude Code (Sonnet 4.6)
Full project build
Sessions include
Designing the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/kirana-detective-build-traces.hackathon-advisor-codex-traces
Hackathon Advisor Codex Session Traces
Real Codex session logs for the Hackathon Advisor project, selected from local Codex
rollout JSONL files and redacted before publication. The event stream preserves user
requests, assistant messages, tool calls, tool outputs, browser/search events, and
minimal session provenance needed to audit how the project was built.
Privacy filtering
The publisher applied openai/privacy-filter
at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.pit-wall-chaos-tracesCodex agent traces for Pit Wall Chaos, a Build Small Hackathon project.
Space link: https://huggingface.co/spaces/build-small-hackathon/pit-wall-chaos
job-search-assistant-agent-traceNeuroBait-Codex-Traces
Codex Session Traces
This folder contains Codex rollout JSONL traces related to the
NeuroBait Build Small Model project.
Included traces:
rollout-2026-06-08T17-03-23-019ea6af-db29-7801-ac01-46dfc88f90b0.jsonl
rollout-2026-06-09T07-10-21-019ea9b7-4610-7223-906e-2d0dba8bae7f.jsonl
rollout-2026-06-09T16-00-29-019eab9c-a18a-7de1-8967-ea63db425a4f.jsonl
The traces were selected because their session metadata contains the project
working directory:… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/NeuroBait-Codex-Traces.processrl-terminal-environments
ProcessRL Terminal Environments
ProcessRL is a collection of behavior-conditioned terminal environments for training and evaluating agent process control. The tasks are designed around failures that appear in interactive terminal work: stopping after a misleading successful command, repeating an unproductive action, failing to pivot after a dead end, losing track of migrated state, and leaving partial progress unfinished.
This release contains the first public train/heldout… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/processrl-terminal-environments.figment-eval-traces
Figment Eval Traces
Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders.
These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment.
Dataset Summary
The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.TinyNarrator-agent-tracesSpaces link: https://huggingface.co/spaces/build-small-hackathon/TinyNarrator
biomed_squad_es_v2
Dataset Card for biomed_squad_es_v2
This Dataset was created as part of the "Extractive QA Biomedicine" project developed during the 2022 Hackathon organized by SOMOS NLP.
Dataset Summary
This is a subset of the dev squad_es (v2) dataset (automatic translation of the Stanford Question Answering Dataset v2 into Spanish) containing questions related to the biomedical domain.
License, distribution and usage conditions of the original Squad_es Dataset apply.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/biomed_squad_es_v2.MatchWise-agent-tracein-your-own-worlds-tracesCodex agent traces for In your own wor(l)ds, a Build Small Hackathon project.
Space link: https://huggingface.co/spaces/build-small-hackathon/in-your-own-worlds
clue-vibes-tracesCodex agent traces for Clue Vibes, a Build Small Hackathon project.
Space link: https://huggingface.co/spaces/build-small-hackathon/clue-vibes
kicky-ai-codex-trace
Kicky AI - Codex agent trace (sanitized)
A redacted OpenAI Codex CLI session trace from building
Kicky AI for the Build Small
Hackathon - shared for the Sharing is Caring badge so others can see how the build went.
Format: Codex CLI JSONL session log (each record = {payload, timestamp, type}).
All secrets removed (HF / Modal / Roboflow tokens, shared secrets, emails) - verified 0 leaks.
Blog write-up: https://dcrey7.substack.com/p/world-fut-coach
bedtime-story-machine-trace
🌙 Bedtime Story Machine — Agent Build Trace
This dataset contains the build trace for the Bedtime Story Machine project, built for the Build Small Hackathon.
What's in this trace
Complete build steps from concept to deployment
Architecture decisions and model choices
Challenges encountered and solutions
Tools and infrastructure used
Project
→ Live App
→ GitHub
nvidia-hackathon-dataset
Crack & Pavement Segmentation — Curated Dataset Index
This repository is a curated reference index of public datasets relevant to
crack and pavement segmentation. It does not redistribute any third-party
images; it points to each dataset's canonical source so that researchers and
practitioners can download them from the original authors under the original
license terms.
For each dataset the index records: publisher, source link, brief domain
description, approximate size, mask… See the full description on the dataset page: https://huggingface.co/datasets/crackedcity/nvidia-hackathon-dataset.thousand-token-wood-traces
Thousand Token Wood -- Council Agent Traces
Open agent traces from Thousand Token Wood,
a tiny emergent economy where five woodland creatures trade goods for pebbles,
gossip, and react to reskinned market-history "Wood Legends". Each creature
thinks on a different lab's small model (distinct-engine budget 29.5B <= 32B):
Creature
Model
Lab
Oona (owl)
openai/gpt-oss-20b
OpenAI
Bramble (squirrel)
openbmb/MiniCPM3-4B
OpenBMB
Fenn (fox)
nvidia/Nemotron-Mini-4B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/thousand-token-wood-traces.sense-garden-tracesCodex agent traces for Sense Garden, a Build Small Hackathon project.
Space link: https://huggingface.co/spaces/build-small-hackathon/sense-garden
nli-esannotations_creators:
crowdsourced
other
language_creators:
other
crowdsourced
languages:
es
licenses:
cc-by-sa-4.0
multilinguality:
monolingual
pretty_name: ESnli
size_categories:
unknown
source_datasets:
extended|snli
extended|xnli
extended|multi_nli
task_categories:
text-classification
task_ids:
natural-language-inference
Dataset Card for nli-es
Dataset Summary
A Spanish Natural Language Inference dataset put together from the sources:
the Spanish slice of the XNLI… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/nli-es.winogrande_train_s_spanishThis is the Spanish version of Winogrande Small (640 instances) for training only.
The translation was done manually by a group of experts. The dataset will still be improved in the future.
we also acknowledge Somos-NLP for this achievement.
CV-Mistral-Hackathon_doom-mistral-final-latest-10-frames-ShareGPTOriginal set is one frame per sample.
This does the latest 10 samples, like so:
Sample 1: Frames 1, 2, 3, 4, 5, 6, 7, 8, 9, 10
Sample 2: Frames 2, 3, 4, 5, 6, 7, 8, 9, 10, 11
Sample 3: Frames 3, 4, 5, 6, 7, 8, 9, 10, 11, 12
Sample 4: Frames 4, 5, 6, 7, 8, 9, 10, 11, 12, 13
If it's trained this way maybe it will have a better understanding of the ASCII since it's not doing 0-shot every time, and gets some previous information.
exam-panic-rescue-build-trace
Exam Panic Rescue App Traces
Public-safe traces for the Build Small Hackathon project Exam Panic Rescue (Backyard AI track).
This dataset supports the Sharing is Caring path and documents the Backyard AI real-user evidence. It focuses on what the app does, not on private build strategy notes.
Configs / files
real_user_traces.jsonl (config: real_user) — a real student's anonymized session. A final-year university Machine Learning student (published as "R." with… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/exam-panic-rescue-build-trace.1000-Rooms-tracesThis dataset contains Codex traces from ferrariedhgs's 1000-Rooms submission (space, dataset and fine-tuned model)
cxr-draft-auditor-traces
CXR Draft Auditor - Open Audit Traces
RESEARCH / EDUCATIONAL QA ONLY. These traces come from a research tool that is NOT a medical device, NOT a diagnostic tool, and NOT a substitute for a qualified radiologist. Nothing here may be used for clinical decision-making, screening, or patient care. The model is frequently wrong.
This is a small open-trace dataset for my CXR Draft Auditor Space, captured for the Build Small Hackathon. Each record is one full audit run from the live… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/cxr-draft-auditor-traces.
