CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abhika-m /fava-flagged-demo Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/abhika-m/fava-flagged-demo.textn<1K0 likes1.4k downloads2y agoHugging Face02flax-community /conceptual-12m-mbart-50-multilingualimage10M<n<100M2 likes626 downloads5y agoHugging Face03flax-community /conceptual-captions-12This file contains English captions from Conceptual 12M dataset by Google. Since we don't own the images, we have provided the link to images, name of downloaded file, and caption for that image in the TSV file. We would like to thank Luke Melas for helping us get the cleaned CC-12M data on our TPU-VMs. image10M<n<100M5 likes361 downloads3y agoHugging Face04ksumit /food17-flaggedimagen<1K0 likes251 downloads2y agoHugging Face05flax-community /conceptual-12m-multilingual-marianThis dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models. Data distribution is following: train_file_marian_final.tsv: 10010625 captions (2502656 captions of English, German, Spanish, French each) val_file_marian_final.tsv:… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian.text10M<n<100M1 likes219 downloads3y agoHugging Face06flax-community /conceptual-12m-multilingual-marian-128This dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models (with sequence length 128). Data distribution is following: train_file_marian_final.tsv: 10002432 captions (2500608 captions of English, German, Spanish, French each)… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian-128.text10M<n<100M0 likes192 downloads3y agoHugging Face07huanmit /flan-t5-boosting-mmlu_cottext1K<n<10K1 likes162 downloads3y agoHugging Face08nasa-ibm-ai4science /surya-bench-flare-forecasting Full-disk Solar Flare Forecasting Dataset Dataset Summary This dataset provides labels for solar flare forecasting derived from NOAA GOES flare events from May 2010 to December 2024. Labels are constructed using a 24h rolling prediction window sampled at an hourly cadence. Each window is annotated with both max GOES class (based on peak X-ray flux) and cumulative flare index. Two derived binary labels are included for forecasting tasks: label_max: 1 if the maximum… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/surya-bench-flare-forecasting.tabular100K<n<1M1 likes139 downloads9mo agoHugging Face09FlameF0X /evalstabularn<1K1 likes99 downloads18d agoHugging Face10flax-community /conceptual-12m-multilingual-marian-esimage1M<n<10M0 likes89 downloads3y agoHugging Face11electricsheepafrica /africa-synth-energy-oilgas-gas-flaring-nigeria Africa Synth Energy Oilgas Gas Flaring Nigeria | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-oilgas-gas-flaring-nigeria.tabulartabular-classification1K<n<10K0 likes87 downloads1mo agoHugging Face12rockdrigoma /spanish-nahuatl-flaggingtextn<1K2 likes84 downloads4y agoHugging Face13flax-community /multilingual-vqatext1M<n<10M0 likes77 downloads3y agoHugging Face14AreejAlotaibi12 /skin-cancer-flagged-dataset Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/AreejAlotaibi12/skin-cancer-flagged-dataset.imagen<1K0 likes76 downloads2y agoHugging Face15Flaglab /CIL-Dataset CIL-Dataset: Colombian Indigenous Languages Parallel and monolingual corpora for machine translation between Spanish and four Colombian Indigenous languages: Wayuunaiki, Inga, Kamëntsá and Nasa Yuwe. Compiled for the thesis Data-Centric Strategies for Low-Resource Neural Machine Translation in Colombian Indigenous Languages (Universidad de los Andes, FLAG Lab). Languages Language Column ISO 639-3 Spanish esp spa Wayuunaiki way guc Inga ing inb… See the full description on the dataset page: https://huggingface.co/datasets/Flaglab/CIL-Dataset.tabulartranslation100K<n<1M1 likes71 downloads3mo agoHugging Face16flamiinngo /agronomy-qa-agriculture Agronomy QA Pairs — Agriculture Instruction Dataset A concise, English-language question-and-answer dataset covering practical agriculture: crop management and planting, soil health and fertility, irrigation, and pest and disease control, with a smaller amount of livestock content. Curated for instruction fine-tuning of language models in the agriculture domain. Rows 1,914 question/answer pairs Unique questions 1,914 — no duplicate questions Unique answers 1,870… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/agronomy-qa-agriculture.textquestion-answering1K<n<10K1 likes67 downloads2mo agoHugging Face17retarfi /flare-fomctext1K<n<10K1 likes56 downloads2y agoHugging Face18frostglade /flat-birthday-cf2c09 flat-birthday-cf2c09 Synthetic products test data: 41 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/frostglade/flat-birthday-cf2c09.tabularn<1K0 likes50 downloads15d agoHugging Face19flamiinngo /market-news-qa Market News QA — Market-Analysis Instruction Dataset Concise question-and-answer pairs for market analysis and financial news interpretation: classifying news by market area, reading sentiment, identifying who a story matters to, and answering forward-looking questions from earnings calls. Built for the Adaption Labs AutoScientist Challenge (Market-Analysis & News category). Rows 10,011 Distinct answers 8,469 (85%) Duplicate questions none Nulls none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/market-news-qa.textquestion-answering10K<n<100K1 likes49 downloads2mo agoHugging Face20huanmit /flan-t5-boosting-bbh_cottext1K<n<10K3 likes44 downloads3y agoHugging Face21cureskin /indian-skincare-ingredient-flags Indian Skincare Ingredient Flags 1,448 cosmetic (INCI) ingredients found in skincare sold in India, with dermatology-referenced flags: Malassezia (fungal-acne) trigger status, comedogenicity rating, fragrance/EU-allergen, alcohol type and pregnancy-caution. Published by CureSkin (Cure and Care Wellness Pvt. Ltd.) — the data layer behind Indian Skincare, Decoded, medically reviewed by Dr. Charu Sharma, Head of Dermatology at CureSkin. Blank cells mean "not assessed", never "safe"… See the full description on the dataset page: https://huggingface.co/datasets/cureskin/indian-skincare-ingredient-flags.tabular1K<n<10K2 likes43 downloads3mo agoHugging Face22flax-sentence-embeddings /Gender_Bias_Evaluation_SetThis dataset has been created as part of the Flax/JAX community week for testing the flax-sentence-embeddings Sentence Similarity models for Gender Bias but can be used for other use-cases as well related to evaluating Gender Bias. The Following Dataset has been created for Evaluating Gender Bias for different models, based on various stereotypical occupations. The Structure of the dataset is of the following type: Base Sentence Occupation Steretypical_Gender Male Sentence Female… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/Gender_Bias_Evaluation_Set.text1K<n<10K4 likes39 downloads2mo agoHugging Face23wren-ilands /flagged-antibody-catalogs Flagged antibody catalogs A one-row-per-product view of the 2026 antibody validation-image audit: which catalog numbers have at least one problematic validation image on the vendor's own product page. Source and credit Derived from Richardson, R. & David, S. (2026), "Problematic images in vendor antibody verification data", Zenodo, version 260825, published 25 August 2026, DOI 10.5281/zenodo.22090940, licensed CC BY 4.0. The source audit documents 18,943… See the full description on the dataset page: https://huggingface.co/datasets/wren-ilands/flagged-antibody-catalogs.text10K<n<100K0 likes38 downloads5d agoHugging Face24ovi054 /mail-flag-data Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/ovi054/mail-flag-data.tabularn<1K0 likes35 downloads2y agoHugging Face25flamiinngo /math-code-qa Math & Code QA — Instruction Dataset Worked mathematical solutions and short code answers, built for the Adaption Labs AutoScientist Challenge (Math & Code category). Rows 5,200 Math 3,600 Code 1,600 Distinct answers 5,199 (100%) Duplicate questions none Nulls none Question length median 27 words Answer length median 58 words (max 89) License CC-BY-4.0 What makes the math rows unusual Every math answer is short worked reasoning… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa.textquestion-answering1K<n<10K1 likes34 downloads2mo agoHugging Face26huanmit /flan-t5-boosting-bbh_directtext1K<n<10K1 likes32 downloads3y agoHugging Face27relai-ai /flask-standardSamples in this benchmark were generated by RELAI using the following data source(s): Data Source Name: Flask Documentation Data Source Link: https://flask.palletsprojects.com/en/stable/ Data Source License: https://flask.palletsprojects.com/en/stable/license/ Data Source Authors: Pallets Project AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai textquestion-answering1K<n<10K0 likes32 downloads1y agoHugging Face28flaviawallen /MNLP_M3_rag_documentstext1K<n<10K0 likes30 downloads1y agoHugging Face29kellydoesstuff /RotBot_Flags Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/kellydoesstuff/RotBot_Flags.textn<1K0 likes29 downloads5mo agoHugging Face30flamiinngo /math-code-qa-v2 Math & Code QA v2 — Instruction Dataset Worked mathematical solutions and short code answers, spanning arithmetic word problems through to algebra, geometry and combinatorics. Built for the Adaption Labs AutoScientist Challenge (Math & Code category). The model trained on this beats Llama-3.3-70B-Instruct 72 to 28 on the held-out Math category evaluation. Rows 5,297 (4,197 math, 1,100 code) Distinct answers 5,297 (100%) Duplicate questions none Nulls none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa-v2.textquestion-answering1K<n<10K0 likes28 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.