CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pseudolab /huggingface-krew-hackathon2023textn<1K2 likes1.5k downloads3y agoHugging Face02hackathon-ai-politician /chat_datatext1K<n<10K0 likes512 downloads2y agoHugging Face03Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_3 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_3" More Information needed audio10K<n<100K0 likes490 downloads3y agoHugging Face04nattasunit /brain-hackathon-2023-embed-datatabular100K<n<1M0 likes349 downloads3y agoHugging Face05Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_2 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_2" More Information needed audio10K<n<100K0 likes287 downloads3y agoHugging Face06somosnlp-hackathon-2023 /informes_discriminacion_gitana Resumen del dataset Se trata de un dataset en español, extraído del centro de documentación de la Fundación Secretariado Gitano, en el que se presentan distintas situaciones discriminatorias acontecidas por el pueblo gitano. Puesto que el objetivo del modelo es crear un sistema de generación de actuaciones que permita minimizar el impacto de una situación discriminatoria, se hizo un scrappeo y se extrajeron todos los PDFs que contuvieron casos de discriminación con el formato… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/informes_discriminacion_gitana.imagetext-classification1K<n<10K8 likes267 downloads3y agoHugging Face07build-small-hackathon /agenda-parser-tool-traces Agenda Parser — tool-calling reasoning traces ReAct tool-calling traces for the Agenda Parser agents: each row is one agent step — a {system, user, assistant} chat example where the assistant emits a single JSON action {"thought", "tool", "args"}. Two agents are covered (tagged by meta.domain): agenda — the uploaded-packet research agent, over real public-meeting agenda packets (tools: list/read items, semantic + exact search, summarize, report). Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.documenttext-generation1K<n<10K0 likes206 downloads3mo agoHugging Face08poolside-laguna-hackathon /protein-ligand-design 🧪 Protein-Ligand Design Gym — Team JAMMY poolside Laguna Hackathon submission. A tool-use reinforcement-learning environment that teaches an LLM to reason like a bench computational chemist / protein engineer — by measuring, not guessing. The problem Proteins are the molecular machines inside living cells, each built from a long string of amino-acid "letters". Ligands are the small molecules — most drugs among them — that bind to a protein to switch it on or… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/protein-ligand-design.textquestion-answering1K<n<10K1 likes205 downloads3mo agoHugging Face09somosnlp-hackathon-2022 /spanish-to-quechua Spanish to Quechua Dataset Description This dataset is a recopilation of webs and others datasets that shows in dataset creation section. This contains translations from spanish (es) to Qechua of Ayacucho (qu). Dataset Structure Data Fields es: The sentence in Spanish. qu: The sentence in Quechua of Ayacucho. Data Splits train: To train the model (102 747 sentences). Validation: To validate the model during training (12 844 sentences).… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/spanish-to-quechua.texttranslation100K<n<1M16 likes192 downloads4y agoHugging Face10Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_5 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_5" More Information needed audio10K<n<100K0 likes180 downloads3y agoHugging Face11legal-hackathon-2024 /synthetictabular100K<n<1M0 likes175 downloads2y agoHugging Face12Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_1 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_1" More Information needed audio10K<n<100K0 likes159 downloads3y agoHugging Face13Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_6 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_6" More Information needed audio10K<n<100K0 likes152 downloads3y agoHugging Face14traversaal-ai-hackathon /hotel_datasetsimage1K<n<10K3 likes151 downloads3y agoHugging Face15Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_4 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_4" More Information needed audio10K<n<100K0 likes146 downloads3y agoHugging Face16somosnlp-hackathon-2022 /readability-es-hackathon-pln-public Dataset Card for [readability-es-sentences] Dataset Description Compilation of short Spanish articles for readability assessment. Dataset Summary This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources: Coh-Metrix-Esp corpus (Quispesaravia, et al., 2016): collection of 100 parallel texts with simple and complex variants in Spanish. These texts… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/readability-es-hackathon-pln-public.texttext-classification1K<n<10K3 likes128 downloads3y agoHugging Face17SDSC /open-pulse-hackathon-data-analysis LauzHack Projects Dataset Dataset Summary This dataset contains comprehensive information about projects submitted to LauzHack (EPFL's student-run hackathon) from 2023 to 2025. Each project includes details about the project title, description, team members, awards, and categories. LauzHack is an annual 24-hour hackathon hosted at EPFL (École Polytechnique Fédérale de Lausanne) in Lausanne, Switzerland, bringing together students and hackers to create innovative solutions… See the full description on the dataset page: https://huggingface.co/datasets/SDSC/open-pulse-hackathon-data-analysis.tabularn<1K0 likes127 downloads5mo agoHugging Face18build-small-hackathon /jawbreaker-scam-defense-data Jawbreaker Scam Defense Data Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love. Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays. Contents eval/: scam-defense evaluation sets from smoke checks through hard calibration suites. eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.texttext-classification10K<n<100K6 likes125 downloads3mo agoHugging Face19conradry /biopharma-hackathon Biopharma hackathon data Two independent datasets share this repo. They have different sources and different licenses, and nothing joins them: GenomeScreen (relational) — 5 tables, the DrugCLIP genome-wide virtual screen parsed into parquet. Parkinson's disease subgraph — 5 tables, a pathway-centric neighbourhood extracted from PrimeKG, as a graph and as a disease→pathway→protein→drug tree, plus an environmental-toxin overlay on the same pathways. 1. GenomeScreen… See the full description on the dataset page: https://huggingface.co/datasets/conradry/biopharma-hackathon.tabular1M<n<10M0 likes96 downloads1mo agoHugging Face20build-small-hackathon /CVE_Vulnerailities_Detailedtext10K<n<100K17 likes94 downloads3mo agoHugging Face21sukantabasu /alchemist-shell.ai-hackathon-2025This project is described in detail at this website: https://alchemist-shellai-hackathon-2025.readthedocs.io/en/latest/ The codes and relevant materials are available here: https://github.com/Sukantabasu/alchemist-shell.ai-hackathon-2025 The trained models (in pkl format) are stored in this HF repository. textn<1K0 likes90 downloads1y agoHugging Face22somosnlp-hackathon-2022 /Axolotl-Spanish-Nahuatl Axolotl-Spanish-Nahuatl : Parallel corpus for Spanish-Nahuatl machine translation Dataset Collection In order to get a good translator, we collected and cleaned two of the most complete Nahuatl-Spanish parallel corpora available. Those are Axolotl collected by an expert team at UNAM and Bible UEDIN Nahuatl Spanish crawled by Christos Christodoulopoulos and Mark Steedman from Bible Gateway site. After this, we ended with 12,207 samples from Axolotl due to misalignments and… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/Axolotl-Spanish-Nahuatl.tabulartranslation10K<n<100K16 likes89 downloads3y agoHugging Face23yasalma /tat_hackathon_asr Hackathon Tatar ASR Dataset Summary Hackathon Tatar ASR is a speech dataset distributed during the "Татар.Бу Хакатон" (Tatar.Bu Hackathon) held in Tatarstan in May 2024. This dataset likely consists of newly collected crowdsourced recordings created after the last release of TatSC (Tatar Speech Corpus), although some intersections with TatSC might be present. While TatSC contains 269.1 hours of transcribed speech with 271,914 utterances, this hackathon dataset comprises… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tat_hackathon_asr.audioautomatic-speech-recognition10K<n<100K0 likes79 downloads1y agoHugging Face24build-small-hackathon /dota2tuned-data DOTA2Tuned Data This dataset supports the DOTA2Tuned Hugging Face Build Small Hackathon app. It contains compact derived artifacts for Dota 2 draft recommendations, hero meta lookup, build timing summaries, match prediction, retrieval, and supervised fine-tuning examples. Contents sft_examples.jsonl: instruction examples generated from normalized Dota 2 recommendations, patch/stat cards, and app behaviors. Compact Parquet artifacts used by the Space: dim_hero… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/dota2tuned-data.tabulartext-generation1M<n<10M0 likes79 downloads3mo agoHugging Face25factored /pinecone_hackathon Dataset Card for "pinecone_hackathon" More Information needed text100K<n<1M0 likes78 downloads3y agoHugging Face26somosnlp-hackathon-2023 /Habilidades_Agente_v1 Description Español: Presentamos un conjunto de datos que presenta tres partes principales: 1. Dataset sobre habilidades blandas. 2. Dataset de conversaciones empresariales entre agentes y clientes. 3. Dataset curado de Alpaca en español: Este dataset toma como base el dataset https://huggingface.co/datasets/somosnlp/somos-alpaca-es, y fue curado con la herramienta Argilla, alcanzando 9400 registros curados. Los datos están estructurados en torno a un método que se describe… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/Habilidades_Agente_v1.texttext-generation10K<n<100K22 likes69 downloads3y agoHugging Face27Kearm /LLaMutation-Hackathontext10K<n<100K0 likes67 downloads2y agoHugging Face28taylor-geospatial /neural-earth-fields-hackathon Neural Earth Fields Hackathon Land cover CGLC-MODIS-LCZ-100m.cog.tif is a Cloud Optimized GeoTIFF version of A hybrid 100-m global land cover dataset with Local Climate Zones for WRF by Matthias Demuzere, Cenlin He, Alberto Martilli, and Andrea Zonato (2023). The source dataset is licensed under CC BY 4.0. The file contains one uint8 class ID per pixel on the source's nominal 100 m EPSG:4326 grid. It uses ZSTD compression, 512 × 512 tiles, and mode-resampled… See the full description on the dataset page: https://huggingface.co/datasets/taylor-geospatial/neural-earth-fields-hackathon.imagen<1K0 likes66 downloads12d agoHugging Face29alvanlii /devpost-hackathon-projects Hackathon Projects Summary This dataset contains 200k+ hackathon project descriptions from 6700+ hackathons last updated Jan 2025 Data Description combined_hackathons.parquet: This contains all the projects hackathon_id project_link: Link to the hackathon project full_desc: Full description of the project title: Title of project brief_desc: Summary of project team_members: Each member of the team in a list prize: Prizes won in a list tags… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/devpost-hackathon-projects.tabular100K<n<1M3 likes61 downloads2y agoHugging Face30somosnlp-hackathon-2022 /readability-es-caes Dataset Card for [readability-es-caes] Dataset Description Dataset Summary This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources: CAES corpus (Martínez et al., 2019): the "Corpus de Aprendices del Español" is a collection of texts produced by Spanish L2 learners from Spanish learning centers and universities. These text are produced by students… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/readability-es-caes.texttext-classification10K<n<100K3 likes59 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.