datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
huggingface-krew-hackathon2023mozilla_commonvoice_hackathon_preprocessed_train_batch_3
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_3"
More Information needed
chat_databrain-hackathon-2023-embed-datamozilla_commonvoice_hackathon_preprocessed_train_batch_2
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_2"
More Information needed
informes_discriminacion_gitana
Resumen del dataset
Se trata de un dataset en español, extraído del centro de documentación de la Fundación Secretariado Gitano, en el que se presentan distintas situaciones discriminatorias acontecidas por el pueblo gitano. Puesto que el objetivo del modelo es crear un sistema de generación de actuaciones que permita minimizar el impacto de una situación discriminatoria, se hizo un scrappeo y se extrajeron todos los PDFs que contuvieron casos de discriminación con el formato… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/informes_discriminacion_gitana.protein-ligand-design
🧪 Protein-Ligand Design Gym — Team JAMMY
poolside Laguna Hackathon submission. A tool-use reinforcement-learning
environment that teaches an LLM to reason like a bench computational chemist /
protein engineer — by measuring, not guessing.
The problem
Proteins are the molecular machines inside living cells, each built from a long
string of amino-acid "letters". Ligands are the small molecules — most drugs
among them — that bind to a protein to switch it on or… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/protein-ligand-design.spanish-to-quechua
Spanish to Quechua
Dataset Description
This dataset is a recopilation of webs and others datasets that shows in dataset creation section. This contains translations from spanish (es) to Qechua of Ayacucho (qu).
Dataset Structure
Data Fields
es: The sentence in Spanish.
qu: The sentence in Quechua of Ayacucho.
Data Splits
train: To train the model (102 747 sentences).
Validation: To validate the model during training (12 844 sentences).… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/spanish-to-quechua.syntheticmozilla_commonvoice_hackathon_preprocessed_train_batch_5
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_5"
More Information needed
jawbreaker-scam-defense-data
Jawbreaker Scam Defense Data
Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love.
Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays.
Contents
eval/: scam-defense evaluation sets from smoke checks through hard calibration suites.
eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.alchemist-shell.ai-hackathon-2025This project is described in detail at this website:
https://alchemist-shellai-hackathon-2025.readthedocs.io/en/latest/
The codes and relevant materials are available here:
https://github.com/Sukantabasu/alchemist-shell.ai-hackathon-2025
The trained models (in pkl format) are stored in this HF repository.
mozilla_commonvoice_hackathon_preprocessed_train_batch_1
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_1"
More Information needed
mozilla_commonvoice_hackathon_preprocessed_train_batch_4
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_4"
More Information needed
mozilla_commonvoice_hackathon_preprocessed_train_batch_6
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_6"
More Information needed
hotel_datasetsreadability-es-hackathon-pln-public
Dataset Card for [readability-es-sentences]
Dataset Description
Compilation of short Spanish articles for readability assessment.
Dataset Summary
This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources:
Coh-Metrix-Esp corpus (Quispesaravia, et al., 2016): collection of 100 parallel texts with simple and complex variants in Spanish. These texts… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/readability-es-hackathon-pln-public.open-pulse-hackathon-data-analysis
LauzHack Projects Dataset
Dataset Summary
This dataset contains comprehensive information about projects submitted to
LauzHack (EPFL's student-run hackathon) from 2023 to 2025. Each project
includes details about the project title, description, team members, awards, and
categories.
LauzHack is an annual 24-hour hackathon hosted at EPFL (École Polytechnique
Fédérale de Lausanne) in Lausanne, Switzerland, bringing together students and
hackers to create innovative solutions… See the full description on the dataset page: https://huggingface.co/datasets/SDSC/open-pulse-hackathon-data-analysis.kirana-detective-build-traces
Kirana Detective — Claude Code Build Sessions
Raw Claude Code (claude-sonnet-4-6) session traces recorded while building
Kirana Detective AI for the HuggingFace Build Small Hackathon 2026.
Each .jsonl file is one coding session. Together they cover the entire
build — from first commit to final submission.
What's Inside
Sessions
Agent
Coverage
11 JSONL files
Claude Code (Sonnet 4.6)
Full project build
Sessions include
Designing the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/kirana-detective-build-traces.biopharma-hackathon
Biopharma hackathon data
Two independent datasets share this repo. They have different sources and different licenses,
and nothing joins them:
GenomeScreen (relational) — 5 tables, the
DrugCLIP genome-wide virtual screen parsed into parquet.
Parkinson's disease subgraph — 5 tables, a
pathway-centric neighbourhood extracted from PrimeKG, as a graph and as a
disease→pathway→protein→drug tree, plus an environmental-toxin overlay on the same
pathways.
1. GenomeScreen… See the full description on the dataset page: https://huggingface.co/datasets/conradry/biopharma-hackathon.Axolotl-Spanish-Nahuatl
Axolotl-Spanish-Nahuatl : Parallel corpus for Spanish-Nahuatl machine translation
Dataset Collection
In order to get a good translator, we collected and cleaned two of the most complete Nahuatl-Spanish parallel corpora available. Those are Axolotl collected by an expert team at UNAM and Bible UEDIN Nahuatl Spanish crawled by Christos Christodoulopoulos and Mark Steedman from Bible Gateway site.
After this, we ended with 12,207 samples from Axolotl due to misalignments and… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/Axolotl-Spanish-Nahuatl.pinecone_hackathon
Dataset Card for "pinecone_hackathon"
More Information needed
figment-eval-traces
Figment Eval Traces
Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders.
These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment.
Dataset Summary
The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.CVE_Vulnerailities_Detaileddota2tuned-data
DOTA2Tuned Data
This dataset supports the DOTA2Tuned Hugging Face Build Small Hackathon app. It contains compact derived artifacts for Dota 2 draft recommendations, hero meta lookup, build timing summaries, match prediction, retrieval, and supervised fine-tuning examples.
Contents
sft_examples.jsonl: instruction examples generated from normalized Dota 2 recommendations, patch/stat cards, and app behaviors.
Compact Parquet artifacts used by the Space:
dim_hero… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/dota2tuned-data.tat_hackathon_asr
Hackathon Tatar ASR
Dataset Summary
Hackathon Tatar ASR is a speech dataset distributed during the "Татар.Бу Хакатон" (Tatar.Bu Hackathon) held in Tatarstan in May 2024. This dataset likely consists of newly collected crowdsourced recordings created after the last release of TatSC (Tatar Speech Corpus), although some intersections with TatSC might be present. While TatSC contains 269.1 hours of transcribed speech with 271,914 utterances, this hackathon dataset comprises… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tat_hackathon_asr.neutral-es
Spanish Gender Neutralization
Spanish is a beautiful language and it has many ways of referring to people, neutralizing the genders and using some of the resources inside the language. One would say Todas las personas asistentes instead of Todos los asistentes and it would end in a more inclusive way for talking about people. This dataset collects a set of manually anotated examples of gendered-to-neutral spanish transformations.
The intended use of this dataset is to train a… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/neutral-es.hackathon-advisor-codex-traces
Hackathon Advisor Codex Session Traces
Real Codex session logs for the Hackathon Advisor project, selected from local Codex
rollout JSONL files and redacted before publication. The event stream preserves user
requests, assistant messages, tool calls, tool outputs, browser/search events, and
minimal session provenance needed to audit how the project was built.
Privacy filtering
The publisher applied openai/privacy-filter
at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.lolaby-traces
Lolaby — generation traces
Pipeline traces from Lolaby, an AI-powered lullaby generator built for the Build Small Hackathon 2026 (Backyard AI track).
Each trace is a complete witness of one end-to-end generation: every input the user gave, every model that ran, every prompt and raw output, every timing measurement, and the final audio. Published under CC0 so anyone can study, replay, or remix the pipeline.
What's in a trace
Each subfolder is one generation. Files:… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/lolaby-traces.
