CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hackercupai /hackercup Data Preview The data available in this preview contains a 10 row dataset: Sample Dataset ("sample"): This is a subset of the full dataset, containing data from 2023. To view full dataset, download output_dataset.parquet. This contains data from 2011 to 2023. Fields The dataset include the following fields: name (string) year (string) round (string) statement (string) input (string) solution (string) code (string) sample_input (string) sample_output (string) images… See the full description on the dataset page: https://huggingface.co/datasets/hackercupai/hackercup.imagen<1K24 likes2.2k downloads2y agoHugging Face02open-index /hacker-news-rss Hacker News RSS Feed Directory TL;DR — We visited every unique domain ever posted to Hacker News, found which ones publish RSS/Atom feeds, and packaged the results as monthly parquet snapshots with rich metadata. 623,957 feeds discovered across 1,755,955 hosts, spanning 232 months from 2006-10 to 2026-03. Last updated: 2026-04-05T09:21:39Z Why this exists RSS is not dead — it's just hard to discover. The <link rel="alternate"> tag that points to a site's feed is… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news-rss.imagetext-classification100K<n<1M2 likes1.5k downloads6mo agoHugging Face03julien040 /hacker-news-posts Hacker News Stories Dataset This is a dataset containing approximately 4 million stories from Hacker News, exported to a Parquet file. The dataset includes the following fields: id (int64): The unique identifier of the story. title (string): The title of the story. url (string): The URL of the story. score (int64): The score of the story. time (int64): The time the story was posted, in Unix time. comments (int64): The number of comments on the story. author (string): The… See the full description on the dataset page: https://huggingface.co/datasets/julien040/hacker-news-posts.tabular100K<n<1M8 likes1.3k downloads2mo agoHugging Face04hackaprompt /hackaprompt-datasetgated Dataset Card for HackAPrompt 💻🔍 This dataset contains submissions from a prompt hacking competition. An in-depth analysis of the dataset has been accepted at the EMNLP 2023 conference. 📊👾 Submissions were sourced from two environments: a playground for experimentation and an official submissions platform. The playground itself can be accessed here 🎮 More details about the competition itself here 🏆 Dataset Details 📋 Dataset Description 📄 We… See the full description on the dataset page: https://huggingface.co/datasets/hackaprompt/hackaprompt-dataset.tabular100K<n<1M104 likes948 downloads3y agoHugging Face05DexopT /NOSK-Hackingtext100K<n<1M0 likes538 downloads5mo agoHugging Face06Hack90 /ncbi_part_0_v1tabular10M<n<100M0 likes529 downloads3y agoHugging Face07Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_3 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_3" More Information needed audio10K<n<100K0 likes525 downloads3y agoHugging Face08Hack90 /ref_seq_bacteria_part_2tabular100K<n<1M0 likes514 downloads3y agoHugging Face09OpenPipe /hacker-news Hacker News posts and comments This is a dataset of all HN posts and comments, current as of November 1, 2023. tabular10M<n<100M28 likes502 downloads2y agoHugging Face10Hack90 /ref_seq_vertebrate_non_mammal_part_1tabular100K<n<1M0 likes458 downloads3y agoHugging Face11mistral-hackaton-2026 /zebra-cot-mistral-small-3.2-24b-preprocessed Zebra-CoT Preprocessed — Mistral Hackathon 2026 Preprocessed version of the Zebra-CoT dataset for fine-tuning Mistral-Small-3.2-24B-Instruct. Format text: formatted as [INST] question [/INST] <think> reasoning </think> answer image: PIL JPEG image for the corresponding visual task Usage Fine-tuning Mistral-Small-3.2-24B on chain-of-thought visual reasoning. Hackathon Created for Mistral Hackaton 2026 — Fine-tuning track with W&B. imagevisual-question-answering100K<n<1M0 likes383 downloads7mo agoHugging Face12hackelle /BigEarthNetV2-Lithuania-Summer-LMDB TU Berlin RSiM DIMA BigEarth BIFOLD reBEN — Lithuania Summer Subset (pre-converted to LMDB) ⚠️ Unofficial mirror. This is an unofficial, community-provided pre-conversion of a subset of the BigEarthNet v2.0 (reBEN) dataset into LMDB format. It is provided as a convenience for researchers who wish to get started quickly without running the full conversion pipeline. In case of any discrepancy, the original publication and the original files always take… See the full description on the dataset page: https://huggingface.co/datasets/hackelle/BigEarthNetV2-Lithuania-Summer-LMDB.textimage-classification1K<n<10K0 likes371 downloads6mo agoHugging Face13nattasunit /brain-hackathon-2023-embed-datatabular100K<n<1M0 likes356 downloads3y agoHugging Face14Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_2 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_2" More Information needed audio10K<n<100K0 likes336 downloads3y agoHugging Face15Manusagents /NOSK-Hackingtext100K<n<1M0 likes330 downloads2mo agoHugging Face16Hack90 /ref_seq_vertebrate_non_mammal_part_2tabularn<1K0 likes318 downloads3y agoHugging Face17Hack90 /virus_dna_dataset[Needs More Information] Dataset Card for virus_dna_dataset Dataset Summary A collection of full virus genome dna, the dataset was built from NCBI data Supported Tasks and Leaderboards [Needs More Information] Languages DNA Dataset Structure Data Instances { 'Description' : 'NC_030848.1 Haloarcula californiae icosahedral...', 'dna_sequence' : 'TCATCTC TCTCTCT CTCTCTT GTTCCCG CGCCCGC CCGCCC...', 'sequence_length':'35787'… See the full description on the dataset page: https://huggingface.co/datasets/Hack90/virus_dna_dataset.tabular1M<n<10M8 likes287 downloads3y agoHugging Face18TechyCode /bon-pim-hacking-experimentstext10K<n<100K0 likes267 downloads2mo agoHugging Face19ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.tabulartext-generation10K<n<100K0 likes250 downloads3mo agoHugging Face20labofsahil /hackernews-vector-search-datasetThe Hacker News dataset contains 28.74 million postings and their vector embeddings. The embeddings were generated using SentenceTransformers model all-MiniLM-L6-v2. The dimension of each embedding vector is 384. Created by clickhouse more info: https://clickhouse.com/docs/getting-started/example-datasets/hackernews-vector-search-dataset tabular10M<n<100M2 likes247 downloads10mo agoHugging Face21somosnlp-hackathon-2023 /informes_discriminacion_gitana Resumen del dataset Se trata de un dataset en español, extraído del centro de documentación de la Fundación Secretariado Gitano, en el que se presentan distintas situaciones discriminatorias acontecidas por el pueblo gitano. Puesto que el objetivo del modelo es crear un sistema de generación de actuaciones que permita minimizar el impacto de una situación discriminatoria, se hizo un scrappeo y se extrajeron todos los PDFs que contuvieron casos de discriminación con el formato… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/informes_discriminacion_gitana.imagetext-classification1K<n<10K8 likes236 downloads3y agoHugging Face22ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.tabulartext-generation10K<n<100K0 likes226 downloads3mo agoHugging Face23poolside-laguna-hackathon /protein-ligand-design 🧪 Protein-Ligand Design Gym — Team JAMMY poolside Laguna Hackathon submission. A tool-use reinforcement-learning environment that teaches an LLM to reason like a bench computational chemist / protein engineer — by measuring, not guessing. The problem Proteins are the molecular machines inside living cells, each built from a long string of amino-acid "letters". Ligands are the small molecules — most drugs among them — that bind to a protein to switch it on or… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/protein-ligand-design.textquestion-answering1K<n<10K1 likes207 downloads3mo agoHugging Face24somosnlp-hackathon-2022 /spanish-to-quechua Spanish to Quechua Dataset Description This dataset is a recopilation of webs and others datasets that shows in dataset creation section. This contains translations from spanish (es) to Qechua of Ayacucho (qu). Dataset Structure Data Fields es: The sentence in Spanish. qu: The sentence in Quechua of Ayacucho. Data Splits train: To train the model (102 747 sentences). Validation: To validate the model during training (12 844 sentences).… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/spanish-to-quechua.texttranslation100K<n<1M16 likes204 downloads4y agoHugging Face25Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_5 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_5" More Information needed audio10K<n<100K0 likes193 downloads3y agoHugging Face26Nobody05 /NOSK-Hackingtext100K<n<1M0 likes188 downloads2mo agoHugging Face27Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_1 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_1" More Information needed audio10K<n<100K0 likes172 downloads3y agoHugging Face28Hack90 /ncbi_genbank_part_41 Dataset Card for "ncbi_genbank_part_41" More Information needed tabular100K<n<1M0 likes171 downloads3y agoHugging Face29Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_4 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_4" More Information needed audio10K<n<100K0 likes162 downloads3y agoHugging Face30Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_6 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_6" More Information needed audio10K<n<100K0 likes159 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.