CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pseudolab /huggingface-krew-hackathon2023textn<1K2 likes1.3k downloads3y agoHugging Face02Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_3 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_3" More Information needed audio10K<n<100K0 likes525 downloads3y agoHugging Face03hackathon-ai-politician /chat_datatext1K<n<10K0 likes475 downloads2y agoHugging Face04nattasunit /brain-hackathon-2023-embed-datatabular100K<n<1M0 likes356 downloads3y agoHugging Face05Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_2 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_2" More Information needed audio10K<n<100K0 likes336 downloads3y agoHugging Face06somosnlp-hackathon-2023 /informes_discriminacion_gitana Resumen del dataset Se trata de un dataset en español, extraído del centro de documentación de la Fundación Secretariado Gitano, en el que se presentan distintas situaciones discriminatorias acontecidas por el pueblo gitano. Puesto que el objetivo del modelo es crear un sistema de generación de actuaciones que permita minimizar el impacto de una situación discriminatoria, se hizo un scrappeo y se extrajeron todos los PDFs que contuvieron casos de discriminación con el formato… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/informes_discriminacion_gitana.imagetext-classification1K<n<10K8 likes236 downloads3y agoHugging Face07poolside-laguna-hackathon /protein-ligand-design 🧪 Protein-Ligand Design Gym — Team JAMMY poolside Laguna Hackathon submission. A tool-use reinforcement-learning environment that teaches an LLM to reason like a bench computational chemist / protein engineer — by measuring, not guessing. The problem Proteins are the molecular machines inside living cells, each built from a long string of amino-acid "letters". Ligands are the small molecules — most drugs among them — that bind to a protein to switch it on or… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/protein-ligand-design.textquestion-answering1K<n<10K1 likes207 downloads3mo agoHugging Face08somosnlp-hackathon-2022 /spanish-to-quechua Spanish to Quechua Dataset Description This dataset is a recopilation of webs and others datasets that shows in dataset creation section. This contains translations from spanish (es) to Qechua of Ayacucho (qu). Dataset Structure Data Fields es: The sentence in Spanish. qu: The sentence in Quechua of Ayacucho. Data Splits train: To train the model (102 747 sentences). Validation: To validate the model during training (12 844 sentences).… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/spanish-to-quechua.texttranslation100K<n<1M16 likes204 downloads4y agoHugging Face09legal-hackathon-2024 /synthetictabular100K<n<1M0 likes202 downloads2y agoHugging Face10Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_5 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_5" More Information needed audio10K<n<100K0 likes193 downloads3y agoHugging Face11build-small-hackathon /jawbreaker-scam-defense-data Jawbreaker Scam Defense Data Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love. Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays. Contents eval/: scam-defense evaluation sets from smoke checks through hard calibration suites. eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.texttext-classification10K<n<100K6 likes180 downloads4mo agoHugging Face12build-small-hackathon /agenda-parser-tool-traces Agenda Parser — tool-calling reasoning traces ReAct tool-calling traces for the Agenda Parser agents: each row is one agent step — a {system, user, assistant} chat example where the assistant emits a single JSON action {"thought", "tool", "args"}. Two agents are covered (tagged by meta.domain): agenda — the uploaded-packet research agent, over real public-meeting agenda packets (tools: list/read items, semantic + exact search, summarize, report). Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.documenttext-generation1K<n<10K0 likes179 downloads4mo agoHugging Face13sukantabasu /alchemist-shell.ai-hackathon-2025This project is described in detail at this website: https://alchemist-shellai-hackathon-2025.readthedocs.io/en/latest/ The codes and relevant materials are available here: https://github.com/Sukantabasu/alchemist-shell.ai-hackathon-2025 The trained models (in pkl format) are stored in this HF repository. textn<1K0 likes175 downloads1y agoHugging Face14Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_1 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_1" More Information needed audio10K<n<100K0 likes172 downloads3y agoHugging Face15Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_4 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_4" More Information needed audio10K<n<100K0 likes162 downloads3y agoHugging Face16Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_6 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_6" More Information needed audio10K<n<100K0 likes159 downloads3y agoHugging Face17traversaal-ai-hackathon /hotel_datasetsimage1K<n<10K3 likes149 downloads3y agoHugging Face18somosnlp-hackathon-2022 /readability-es-hackathon-pln-public Dataset Card for [readability-es-sentences] Dataset Description Compilation of short Spanish articles for readability assessment. Dataset Summary This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources: Coh-Metrix-Esp corpus (Quispesaravia, et al., 2016): collection of 100 parallel texts with simple and complex variants in Spanish. These texts… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/readability-es-hackathon-pln-public.texttext-classification1K<n<10K3 likes133 downloads3y agoHugging Face19SDSC /open-pulse-hackathon-data-analysis LauzHack Projects Dataset Dataset Summary This dataset contains comprehensive information about projects submitted to LauzHack (EPFL's student-run hackathon) from 2023 to 2025. Each project includes details about the project title, description, team members, awards, and categories. LauzHack is an annual 24-hour hackathon hosted at EPFL (École Polytechnique Fédérale de Lausanne) in Lausanne, Switzerland, bringing together students and hackers to create innovative solutions… See the full description on the dataset page: https://huggingface.co/datasets/SDSC/open-pulse-hackathon-data-analysis.tabularn<1K0 likes129 downloads5mo agoHugging Face20build-small-hackathon /kirana-detective-build-traces Kirana Detective — Claude Code Build Sessions Raw Claude Code (claude-sonnet-4-6) session traces recorded while building Kirana Detective AI for the HuggingFace Build Small Hackathon 2026. Each .jsonl file is one coding session. Together they cover the entire build — from first commit to final submission. What's Inside Sessions Agent Coverage 11 JSONL files Claude Code (Sonnet 4.6) Full project build Sessions include Designing the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/kirana-detective-build-traces.tabularn<1K0 likes121 downloads4mo agoHugging Face21conradry /biopharma-hackathon Biopharma hackathon data Two independent datasets share this repo. They have different sources and different licenses, and nothing joins them: GenomeScreen (relational) — 5 tables, the DrugCLIP genome-wide virtual screen parsed into parquet. Parkinson's disease subgraph — 5 tables, a pathway-centric neighbourhood extracted from PrimeKG, as a graph and as a disease→pathway→protein→drug tree, plus an environmental-toxin overlay on the same pathways. 1. GenomeScreen… See the full description on the dataset page: https://huggingface.co/datasets/conradry/biopharma-hackathon.tabular1M<n<10M0 likes110 downloads1mo agoHugging Face22somosnlp-hackathon-2022 /Axolotl-Spanish-Nahuatl Axolotl-Spanish-Nahuatl : Parallel corpus for Spanish-Nahuatl machine translation Dataset Collection In order to get a good translator, we collected and cleaned two of the most complete Nahuatl-Spanish parallel corpora available. Those are Axolotl collected by an expert team at UNAM and Bible UEDIN Nahuatl Spanish crawled by Christos Christodoulopoulos and Mark Steedman from Bible Gateway site. After this, we ended with 12,207 samples from Axolotl due to misalignments and… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/Axolotl-Spanish-Nahuatl.tabulartranslation10K<n<100K16 likes88 downloads3y agoHugging Face23factored /pinecone_hackathon Dataset Card for "pinecone_hackathon" More Information needed text100K<n<1M0 likes88 downloads3y agoHugging Face24build-small-hackathon /figment-eval-traces Figment Eval Traces Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders. These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment. Dataset Summary The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.tabulartext-generation100K<n<1M0 likes86 downloads3mo agoHugging Face25build-small-hackathon /CVE_Vulnerailities_Detailedtext10K<n<100K17 likes86 downloads3mo agoHugging Face26build-small-hackathon /dota2tuned-data DOTA2Tuned Data This dataset supports the DOTA2Tuned Hugging Face Build Small Hackathon app. It contains compact derived artifacts for Dota 2 draft recommendations, hero meta lookup, build timing summaries, match prediction, retrieval, and supervised fine-tuning examples. Contents sft_examples.jsonl: instruction examples generated from normalized Dota 2 recommendations, patch/stat cards, and app behaviors. Compact Parquet artifacts used by the Space: dim_hero… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/dota2tuned-data.tabulartext-generation1M<n<10M0 likes82 downloads3mo agoHugging Face27yasalma /tat_hackathon_asr Hackathon Tatar ASR Dataset Summary Hackathon Tatar ASR is a speech dataset distributed during the "Татар.Бу Хакатон" (Tatar.Bu Hackathon) held in Tatarstan in May 2024. This dataset likely consists of newly collected crowdsourced recordings created after the last release of TatSC (Tatar Speech Corpus), although some intersections with TatSC might be present. While TatSC contains 269.1 hours of transcribed speech with 271,914 utterances, this hackathon dataset comprises… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tat_hackathon_asr.audioautomatic-speech-recognition10K<n<100K0 likes80 downloads1y agoHugging Face28somosnlp-hackathon-2022 /neutral-es Spanish Gender Neutralization Spanish is a beautiful language and it has many ways of referring to people, neutralizing the genders and using some of the resources inside the language. One would say Todas las personas asistentes instead of Todos los asistentes and it would end in a more inclusive way for talking about people. This dataset collects a set of manually anotated examples of gendered-to-neutral spanish transformations. The intended use of this dataset is to train a… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/neutral-es.texttranslation1K<n<10K7 likes79 downloads4y agoHugging Face29build-small-hackathon /hackathon-advisor-codex-traces Hackathon Advisor Codex Session Traces Real Codex session logs for the Hackathon Advisor project, selected from local Codex rollout JSONL files and redacted before publication. The event stream preserves user requests, assistant messages, tool calls, tool outputs, browser/search events, and minimal session provenance needed to audit how the project was built. Privacy filtering The publisher applied openai/privacy-filter at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.tabulartext-generationn<1K0 likes77 downloads4mo agoHugging Face30build-small-hackathon /lolaby-traces Lolaby — generation traces Pipeline traces from Lolaby, an AI-powered lullaby generator built for the Build Small Hackathon 2026 (Backyard AI track). Each trace is a complete witness of one end-to-end generation: every input the user gave, every model that ran, every prompt and raw output, every timing measurement, and the final audio. Published under CC0 so anyone can study, replay, or remix the pipeline. What's in a trace Each subfolder is one generation. Files:… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/lolaby-traces.audiotext-to-audion<1K0 likes74 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.