datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mozilla_commonvoice_hackathon_preprocessed_train_batch_3
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_3"
More Information needed
brain-hackathon-2023-embed-datamozilla_commonvoice_hackathon_preprocessed_train_batch_2
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_2"
More Information needed
informes_discriminacion_gitana
Resumen del dataset
Se trata de un dataset en español, extraído del centro de documentación de la Fundación Secretariado Gitano, en el que se presentan distintas situaciones discriminatorias acontecidas por el pueblo gitano. Puesto que el objetivo del modelo es crear un sistema de generación de actuaciones que permita minimizar el impacto de una situación discriminatoria, se hizo un scrappeo y se extrajeron todos los PDFs que contuvieron casos de discriminación con el formato… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/informes_discriminacion_gitana.protein-ligand-design
🧪 Protein-Ligand Design Gym — Team JAMMY
poolside Laguna Hackathon submission. A tool-use reinforcement-learning
environment that teaches an LLM to reason like a bench computational chemist /
protein engineer — by measuring, not guessing.
The problem
Proteins are the molecular machines inside living cells, each built from a long
string of amino-acid "letters". Ligands are the small molecules — most drugs
among them — that bind to a protein to switch it on or… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/protein-ligand-design.spanish-to-quechua
Spanish to Quechua
Dataset Description
This dataset is a recopilation of webs and others datasets that shows in dataset creation section. This contains translations from spanish (es) to Qechua of Ayacucho (qu).
Dataset Structure
Data Fields
es: The sentence in Spanish.
qu: The sentence in Quechua of Ayacucho.
Data Splits
train: To train the model (102 747 sentences).
Validation: To validate the model during training (12 844 sentences).… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/spanish-to-quechua.mozilla_commonvoice_hackathon_preprocessed_train_batch_5
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_5"
More Information needed
mozilla_commonvoice_hackathon_preprocessed_train_batch_1
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_1"
More Information needed
mozilla_commonvoice_hackathon_preprocessed_train_batch_4
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_4"
More Information needed
mozilla_commonvoice_hackathon_preprocessed_train_batch_6
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_6"
More Information needed
open-pulse-hackathon-data-analysis
LauzHack Projects Dataset
Dataset Summary
This dataset contains comprehensive information about projects submitted to
LauzHack (EPFL's student-run hackathon) from 2023 to 2025. Each project
includes details about the project title, description, team members, awards, and
categories.
LauzHack is an annual 24-hour hackathon hosted at EPFL (École Polytechnique
Fédérale de Lausanne) in Lausanne, Switzerland, bringing together students and
hackers to create innovative solutions… See the full description on the dataset page: https://huggingface.co/datasets/SDSC/open-pulse-hackathon-data-analysis.biopharma-hackathon
Biopharma hackathon data
Two independent datasets share this repo. They have different sources and different licenses,
and nothing joins them:
GenomeScreen (relational) — 5 tables, the
DrugCLIP genome-wide virtual screen parsed into parquet.
Parkinson's disease subgraph — 5 tables, a
pathway-centric neighbourhood extracted from PrimeKG, as a graph and as a
disease→pathway→protein→drug tree, plus an environmental-toxin overlay on the same
pathways.
1. GenomeScreen… See the full description on the dataset page: https://huggingface.co/datasets/conradry/biopharma-hackathon.pinecone_hackathon
Dataset Card for "pinecone_hackathon"
More Information needed
CVE_Vulnerailities_Detaileddota2tuned-data
DOTA2Tuned Data
This dataset supports the DOTA2Tuned Hugging Face Build Small Hackathon app. It contains compact derived artifacts for Dota 2 draft recommendations, hero meta lookup, build timing summaries, match prediction, retrieval, and supervised fine-tuning examples.
Contents
sft_examples.jsonl: instruction examples generated from normalized Dota 2 recommendations, patch/stat cards, and app behaviors.
Compact Parquet artifacts used by the Space:
dim_hero… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/dota2tuned-data.tat_hackathon_asr
Hackathon Tatar ASR
Dataset Summary
Hackathon Tatar ASR is a speech dataset distributed during the "Татар.Бу Хакатон" (Tatar.Bu Hackathon) held in Tatarstan in May 2024. This dataset likely consists of newly collected crowdsourced recordings created after the last release of TatSC (Tatar Speech Corpus), although some intersections with TatSC might be present. While TatSC contains 269.1 hours of transcribed speech with 271,914 utterances, this hackathon dataset comprises… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tat_hackathon_asr.neutral-es
Spanish Gender Neutralization
Spanish is a beautiful language and it has many ways of referring to people, neutralizing the genders and using some of the resources inside the language. One would say Todas las personas asistentes instead of Todos los asistentes and it would end in a more inclusive way for talking about people. This dataset collects a set of manually anotated examples of gendered-to-neutral spanish transformations.
The intended use of this dataset is to train a… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/neutral-es.devpost-hackathon-projects
Hackathon Projects
Summary
This dataset contains
200k+ hackathon project descriptions
from 6700+ hackathons
last updated Jan 2025
Data Description
combined_hackathons.parquet: This contains all the projects
hackathon_id
project_link: Link to the hackathon project
full_desc: Full description of the project
title: Title of project
brief_desc: Summary of project
team_members: Each member of the team in a list
prize: Prizes won in a list
tags… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/devpost-hackathon-projects.DiagTrast
Dataset Card for "DiagTrast"
Table of Content
Table of Contents
Dataset Description
Dataset Summary
Supported Tasks and Leaderboards
Languages
Dataset Structure
Data Instances
Data Fields
Data Splits
Dataset Creation
Curation Rationale
Source Data
Annotations
Considerations for Using the Data
Social Impact of Dataset
Discussion of Biases
Other Known Limitations
Team members
Dataset Description
Dataset Summary
For the creation of this… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/DiagTrast.MESD
Dataset Card for MESD
Dataset Summary
Contiene los datos de la base MESD procesados para hacer 'finetuning' de un modelo 'Wav2Vec' en el Hackaton organizado por 'Somos NLP'.
Ejemplo de referencia:
https://colab.research.google.com/github/huggingface/notebooks/blob/master/examples/audio_classification.ipynb
Hemos accedido a la base MESD para obtener ejemplos.
Breve descripción de los autores de la base MESD:
"La Base de Datos del Discurso Emocional Mexicano (MESD en… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/MESD.discoverroute-citiesgradio-agents-mcp-hackathon-certificates
Dataset Card for "gradio-agents-mcp-hackathon-certificates"
More Information needed
sanskrit-sandhi-split-hackathon
Dataset Card for "sanskrit-sandhi-split-hackathon"
More Information needed
gastronomia-hispana-dpo
Gastronomía Hispana DPO
Descripción del Dataset
Este dataset contiene pares de preferencias para el entrenamiento de modelos de lenguaje especializados en gastronomía hispana utilizando la técnica DPO (Direct Preference Optimization). Los datos incluyen conversaciones sobre cocina internacional con un enfoque particular en recetas, ingredientes, técnicas culinarias y tradiciones gastronómicas del mundo hispano.
Estructura del Dataset
El dataset contiene las… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/gastronomia-hispana-dpo.podcasts-ner-es
Dataset Card for "podcasts-ner-es"
Dataset Summary
This dataset comprises of small text snippets extracted from the "Deforme Semanal" podcast,
accompanied by annotations that identify the presence of a predetermined set of entities.
The purpose of this dataset is to facilitate Named Entity Recognition (NER) tasks.
The dataset was created to aid in the identification of entities such as famous people, books, or films in podcasts.
The transcription of the audio was first… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/podcasts-ner-es.mind-of-tashi-selfplay
The Mind of Tashi — self-play traces
Self-play data for SFT of a small reasoning model that plays The Mind of
Tashi — a simultaneous-commit ritual fighting game where the opponent's
<think> block is the game (surfaced to the player as the "mind-scroll").
Two LLMs duel each other under the game's blind-commit contract (each side
sees only the match history, never the opponent's pending move); we keep the
opponent side's full <think> + {move, taunt} as the SFT target.
Part of the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/mind-of-tashi-selfplay.ai-prophecy-court-presence
AI Prophecy Court Presence
Exploration-ready normalized records for AI Prophecy Court, a playful
hackathon project examining the public statements and social presence of major
AI leaders.
Dataset configurations
linkedin: original authored LinkedIn posts from four verified profiles
x: posts, replies, quotes, and visible reposts from six verified profiles
Each row retains its source URL, publication time, content type, engagement
metadata, collection ID, run ID… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/ai-prophecy-court-presence.ask2democracy-cfqa-salud-pension
About Ask2Democracy-cfqa-salud-pension
Ask2Democracy-cfqa-salud-pension is an instructional, context-based generative dataset created using the text reforms of Colombian health and pension systems in Spanish(March 23).
The text was pre-processed and augmented using the chat-gpt-turbo API.
Creado por Jorge Henao 🇨🇴 Twitter LinkedIn Linktree
Con el apoyo de David Torres 🇨🇴 Twitter LinkedIn
Different prompt engineering experiments were conducted to obtain high-quality… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/ask2democracy-cfqa-salud-pension.Agents_MCP_Hackathon_Tools_ListAlong with the MCP Explorer submitted to the great Agents-MCP-Hackathon of june 2025, I added a loop in my code to generate a list of the MCP Tools submitted at the event, feel free to explorer, use it as a tool selector for your agent or as a dataset enlighting the trends in the MCP tools ecosystem.
mcp-birthday-hackathon-certificates
