CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yatsbm /montreal_firetabular1M<n<10M0 likes2.1k downloads2y agoHugging Face02colettemb /firmstabular100K<n<1M0 likes169 downloads2y agoHugging Face03Cleanlab /fire-financial-ner-extractionThis dataset can be used for benchmarking LLM Structured Outputs via the code here: https://github.com/cleanlab/structured-output-benchmark/ text1K<n<10K0 likes152 downloads10mo agoHugging Face04uzw /us_ssa_gender_neutral_first_namesThis is the official dataset for Beyond Binary Gender Labels: Revealing Gender Bias in LLMs through Gender-Neutral Name Predictions Name-based gender prediction has traditionally categorized individuals as either female or male based on their names, using a binary classification system. That binary approach can be problematic in the cases of gender-neutral names that do not align with any one gender, among other reasons. Relying solely on binary gender categories without recognizing… See the full description on the dataset page: https://huggingface.co/datasets/uzw/us_ssa_gender_neutral_first_names.tabulartext-classification10K<n<100K0 likes145 downloads1y agoHugging Face05Firmansyah-Ibrahim /idt5-v4-results-final-lora-s123-20260912T013040606815Z final-lora-s123-20260912T013040606815Z Run artifacts and per-item predictions. Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables. See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity. Metrics { "n": 267, "rule_version": "structural-proxy-v0.4-grounding-separated", "parse_success_pct": 94.7565543071161, "bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-lora-s123-20260912T013040606815Z.texttext-generationn<1K0 likes110 downloads12d agoHugging Face06fireapproved /plants-evidence FireApproved.com plant flammability evidence directory Structured plant-flammability evidence from FireApproved.com, a citation directory for wildfire-resistant construction and plants. The site issues no ratings of its own. Every status is computed from cited evidence. 3,801 plant records / 27,007 evidence rows, plus the 130-source registry those rows are attributed to. Each row keeps the source's own wording, the edition, and the date the document was read. Classifications… See the full description on the dataset page: https://huggingface.co/datasets/fireapproved/plants-evidence.text10K<n<100K0 likes107 downloads25d agoHugging Face07FirstBML1 /afrofinchain-multilingual-web3 AfroFinChain — Multilingual Web3 & Blockchain Dataset Multilingual Web3 & blockchain dataset in Yoruba, Hausa, Igbo, and Nigerian Pidgin with 1,451 terminology entries and 1,451 conversational Q&A pairs. Designed for LLM fine-tuning, financial literacy, and conversational AI in low-resource African languages. Uses culturally grounded analogies (e.g., ajo, adashi, isusu) to make DeFi concepts actually understandable. Built with Adaptive Data by Adaption as part of the Adaption… See the full description on the dataset page: https://huggingface.co/datasets/FirstBML1/afrofinchain-multilingual-web3.texttext-generation1K<n<10K0 likes104 downloads5mo agoHugging Face08ccosme /FiReCS Dataset Card for Filipino-English Reviews with Code-Switching (FiReCS) Dataset Summary We introduce FiReCS, the first sentiment-annotated corpus of product and service reviews involving Filipino-English code-switching. The data set is composed of 10,487 reviews with a fairly balanced number per sentiment class. Inter-annotator agreement is high with a Kripendorffs’s α for ordinal metric of 0.83. Three human annotators were tasked to manually label reviews according to… See the full description on the dataset page: https://huggingface.co/datasets/ccosme/FiReCS.texttext-classification10K<n<100K2 likes96 downloads2y agoHugging Face095dollarfootballapi /premier-league-first-goal-impact-2025-26 What Is the First Goal Worth? — 2025/26 Premier League Match-level data behind a 5DollarFootballAPI study of how the first confirmed goal changed Bet365's normalized in-play win probabilities during the 2025/26 Premier League season. Across 347 usable matches, the median within-match increase in the scoring team's normalized win probability was 23.1 percentage points (bootstrap 95% CI: 22.4–24.2). The median first goal after minute 75 moved the probability by 62.2 points… See the full description on the dataset page: https://huggingface.co/datasets/5dollarfootballapi/premier-league-first-goal-impact-2025-26.tabularn<1K0 likes89 downloads29d agoHugging Face10letrinhan /vn-provinces-mean-age-first-marriage Vietnam provinces mean age at first marriage Provincial and regional mean age at first marriage (years). Coverage 2010 and 2013-2024. Year 2024 is preliminary. Tables cover provinces, regions and national total. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Comparison Color key Files provinces (819 rows) data/provinces.csv data/provinces.dta… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-mean-age-first-marriage.tabularn<1K0 likes89 downloads3d agoHugging Face11Firmansyah-Ibrahim /idt5-v4-results-final-lora-s2026-20260912T034640190015Z final-lora-s2026-20260912T034640190015Z Run artifacts and per-item predictions. Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables. See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity. Metrics { "n": 267, "rule_version": "structural-proxy-v0.4-grounding-separated", "parse_success_pct": 95.88014981273409, "bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-lora-s2026-20260912T034640190015Z.texttext-generationn<1K0 likes87 downloads12d agoHugging Face12Firmansyah-Ibrahim /indo-bloom-corpus 🇮🇩 Indo-Bloom-AQG: A Unified Framework for Controllable Indonesian AQG ⚠️ RESEARCH ARTIFACT STATUS: SILVER VERSION (Work in Progress) This dataset serves as the preliminary corpus (Silver Standard) for the ongoing Doctoral Dissertation at Universitas Negeri Malang (UM). Current State: Unannotated / Pre-validation with Heuristic Bloom Labels Target Final State: Gold Standard (Expert Validated with Bloom's Taxonomy Labels) 🔒 FROZEN — v0.1 Silver This version is permanently… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/indo-bloom-corpus.tabulartext-generation1K<n<10K0 likes81 downloads7mo agoHugging Face13letrinhan /vn-provinces-criminal-cases-first-instance Vietnam criminal cases first-instance trial Vietnam criminal cases first-instance trial. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Comparison Color key Files provinces (189 rows) data/provinces.csv data/provinces.dta data/provinces.xlsx regions (18 rows) data/regions.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-criminal-cases-first-instance.tabularn<1K0 likes80 downloads3d agoHugging Face14firecrawl /scrape-content-dataset-v1 Scrape Content Dataset v1 A human-curated benchmark dataset for evaluating web scraping engines on content quality. Overview This dataset contains 1,000 web pages with human-annotated ground truth for evaluating how well web scraping engines capture core content while avoiding noise (navigation, ads, footers, etc.). The dataset was created in 2025-10-21 and may become outdated over time. Dataset Structure CSV format with columns: id: Sequential identifier url:… See the full description on the dataset page: https://huggingface.co/datasets/firecrawl/scrape-content-dataset-v1.text1K<n<10K0 likes78 downloads11mo agoHugging Face15letrinhan /vn-provinces-fires-explosions Vietnam fires and explosions Vietnam fires and explosions. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Comparison Color key Files provinces (189 rows) data/provinces.csv data/provinces.dta data/provinces.xlsx regions (18 rows) data/regions.csv data/regions.dta… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-fires-explosions.tabularn<1K0 likes75 downloads3d agoHugging Face16Firmansyah-Ibrahim /idt5-v4-results-final-fft-s2026-20260911T063507183162Z final-fft-s2026-20260911T063507183162Z Run artifacts and per-item predictions. Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables. See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity. Metrics { "n": 267, "rule_version": "structural-proxy-v0.4-grounding-separated", "parse_success_pct": 88.01498127340824, "bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-fft-s2026-20260911T063507183162Z.texttext-generationn<1K0 likes64 downloads12d agoHugging Face17Firmansyah-Ibrahim /idt5-v4-results-final-lora-s42-20260912T063343815032Z final-lora-s42-20260912T063343815032Z Run artifacts and per-item predictions. Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables. See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity. Metrics { "n": 267, "rule_version": "structural-proxy-v0.4-grounding-separated", "parse_success_pct": 92.88389513108615, "bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-lora-s42-20260912T063343815032Z.texttext-generationn<1K0 likes61 downloads12d agoHugging Face18fireapproved /osfm-bml-censuses FireApproved.com CAL FIRE OSFM Building Materials Listing — WUI category censuses Complete public censuses of the CAL FIRE Office of the State Fire Marshal Building Materials Listing, published by FireApproved.com. One census per in-scope category code, from a frozen snapshot. Every listing in a covered category is present — a census, not a selection. 276 listings across 16 category codes (decking, siding, doors, windows, vents, eaves, roofing, ignition-resistant and… See the full description on the dataset page: https://huggingface.co/datasets/fireapproved/osfm-bml-censuses.textn<1K0 likes60 downloads25d agoHugging Face19eltorio /french_first_names_insee_2024 French First Names from Death Records (1970-2024) This dataset contains French first names extracted from death records provided by INSEE (French National Institute of Statistics and Economic Studies) covering the period from 1970 to September 2024. Dataset Description Data Source The data is sourced from INSEE's death records database. It includes first names of deceased individuals in France, providing valuable insights into naming patterns across different… See the full description on the dataset page: https://huggingface.co/datasets/eltorio/french_first_names_insee_2024.tabular10K<n<100K4 likes53 downloads2y agoHugging Face20CooperBench /qwen35-9b-plan-first-coop What this is Cooperative two-agent coding dataset: 211 task pairs across 18 repos, generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a plan-first prompt variant — agents are prompted to produce an explicit implementation plan before writing code, then coordinate to reconcile plans before proceeding. Patches are auto-merged after both submit. At a glance Field Value Model Qwen/Qwen3.5-9B Agent mini_swe_agent (plan-first prompt)… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-plan-first-coop.tabularn<1K0 likes46 downloads3mo agoHugging Face21Firmansyah-Ibrahim /idt5-v4-results-final-fft-s42-20260910T135823740810Z final-fft-s42-20260910T135823740810Z Run artifacts and per-item predictions. Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables. See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity. Metrics { "n": 267, "rule_version": "structural-proxy-v0.4-grounding-separated", "parse_success_pct": 93.63295880149813, "bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-fft-s42-20260910T135823740810Z.texttext-generationn<1K0 likes43 downloads12d agoHugging Face22Firmansyah-Ibrahim /idt5-v4-results-final-fft-s123-20260910T215641650583Z final-fft-s123-20260910T215641650583Z Run artifacts and per-item predictions. Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables. See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity. Metrics { "n": 267, "rule_version": "structural-proxy-v0.4-grounding-separated", "parse_success_pct": 93.25842696629213, "bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-fft-s123-20260910T215641650583Z.texttext-generationn<1K0 likes41 downloads12d agoHugging Face23smartnanotubes /snt-fire-enose-samplegated SNT Fire/Smoke E-Nose - Sample Data Subset A small, curated sample of raw recordings from SmartNanotubes' 4×16-channel carbon-nanotube (CNT) electronic-nose arrays, released so that researchers and partners can see the signal quality and experiment with the data. This is a teaser subset, not the full training corpus. It accompanies the model card at smartnanotubes/snt-fire-enose-5class and the interactive demo at smartnanotubes/snt-fire-enose-demo. What's in it… See the full description on the dataset page: https://huggingface.co/datasets/smartnanotubes/snt-fire-enose-sample.tabulartabular-classification10K<n<100K2 likes36 downloads14d agoHugging Face24uzw /canada_ssa_gender_neutral_first_namesThis is the official dataset for Beyond Binary Gender Labels: Revealing Gender Bias in LLMs through Gender-Neutral Name Predictions Name-based gender prediction has traditionally categorized individuals as either female or male based on their names, using a binary classification system. That binary approach can be problematic in the cases of gender-neutral names that do not align with any one gender, among other reasons. Relying solely on binary gender categories without recognizing… See the full description on the dataset page: https://huggingface.co/datasets/uzw/canada_ssa_gender_neutral_first_names.tabulartext-classification1K<n<10K0 likes34 downloads1y agoHugging Face25hasibul1ah /article_2019_first_500tabularn<1K0 likes31 downloads3y agoHugging Face26CarPeAs /first_dataset_iabdExtraído de https://github.com/anthony-wang/BestPractices/tree/master/data. Campos: Formula (string) T (float64): Temperatura (K) CP (float64): Capacidad calorífica (J/mol K) tabular1K<n<10K0 likes31 downloads2y agoHugging Face27mark-miller /firm-shoe-3dd73b firm-shoe-3dd73b Synthetic weather test data: 40 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/mark-miller/firm-shoe-3dd73b.tabularn<1K0 likes31 downloads13d agoHugging Face28hafsteinn /ice_and_fire Ice and Fire Comment Dataset Description The Ice and Fire Dataset is a collection of comments from the Icelandic blog platform, blog.is, that have been annotated in several tasks. Dataset Structure Data Fields annotator_id: An integer identifier for the annotator who labeled the comment. label: The label assigned to the comment. task_type: The type of task the comment was annotated for (see paper). show_blog_post: A boolean indicating whether the… See the full description on the dataset page: https://huggingface.co/datasets/hafsteinn/ice_and_fire.texttext-classification1K<n<10K1 likes30 downloads2y agoHugging Face29CooperBench /qwen35-9b-question-first-coop-random-50 What this is Cooperative two-agent coding dataset: 49 task pairs across 15 repos (random-50 subset), generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a question-first prompt variant — agents begin by asking each other clarifying questions about their respective features before starting implementation, aiming to surface integration concerns early. All 49 pairs were successfully evaluated. At a glance Field Value Model… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-question-first-coop-random-50.tabularn<1K0 likes30 downloads3mo agoHugging Face30fsvf /prop-firm-industry-analysis-2026 Proprietary Trading Firm Industry Dataset (2026) Dataset Description This dataset provides a structured, multi-table overview of the proprietary (prop) trading firm industry as of early 2026. It covers 31 firms across four complementary CSV files, capturing firm characteristics, evaluation program structures, payout policies, and aggregate industry-level metrics. The dataset is intended for researchers, analysts, and practitioners studying the retail-facing… See the full description on the dataset page: https://huggingface.co/datasets/fsvf/prop-firm-industry-analysis-2026.tabulartabular-classificationn<1K0 likes30 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.