CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gtak1 /panlex-meanings Dataset Card for panlex-meanings This is a dataset of words in several thousand languages, extracted from https://panlex.org. Dataset Details Dataset Description This dataset has been extracted from https://panlex.org (the 20240301 database dump) and rearranged on the per-language basis. Each language subset consists of expressions (words and phrases). Each expression is associated with some meanings (if there is more than one meaning, they are in separate… See the full description on the dataset page: https://huggingface.co/datasets/gtak1/panlex-meanings.tabulartranslation10M<n<100M0 likes6.8k downloads8mo agoHugging Face02cointegrated /panlex-meanings Dataset Card for panlex-meanings This is a dataset of words in several thousand languages, extracted from https://panlex.org. Dataset Details Dataset Description This dataset has been extracted from https://panlex.org (the 20240301 database dump) and rearranged on the per-language basis. Each language subset consists of expressions (words and phrases). Each expression is associated with some meanings (if there is more than one meaning, they are in separate… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/panlex-meanings.tabulartranslation100M<n<1B14 likes5.9k downloads1y agoHugging Face03Fllamber /TCGA-PANCAN-HiSeq-2770x20530gene expression cancer RNA-Seq - Check the original submission: - https://www.synapse.org/Synapse:syn2812925 - is maintained by the cancer genome atlas pan-cancer analysis project. - TCGA-PANCAN-HiSeq-2770x20530 Files combined: unc.edu_BRCA_IlluminaHiSeq_RNASeqV2.geneExp (20530, 957) BRCA unc.edu_KIRC_IlluminaHiSeq_RNASeqV2.geneExp (20530, 552) KIRC unc.edu_LUAD_IlluminaHiSeq_RNASeqV2.geneExp (20530, 413) LUAD unc.edu_THCA_IlluminaHiSeq_RNASeqV2.geneExp (20530, 471) THCA… See the full description on the dataset page: https://huggingface.co/datasets/Fllamber/TCGA-PANCAN-HiSeq-2770x20530.tabular1K<n<10K0 likes2.5k downloads2y agoHugging Face04Beijing-AISI /panda-bench PandaBench PandaBench is a comprehensive benchmark for evaluating Large Language Model (LLM) safety, focusing on jailbreak attacks, defense mechanisms, and evaluation methodologies. The PandaGuard framework architecture illustrating the end-to-end pipeline for LLM safety evaluation. The system connects three key components: Attackers, Defenders, and Judges. Dataset Description This repository contains the benchmark results from extensive evaluations of various… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/panda-bench.tabulartext-generation100K<n<1M0 likes1.5k downloads1y agoHugging Face05ogutsevda /graph-pannuke Graph-PanNuke: A Cell-Graph Dataset for Nucleus Classification from PanNuke Graph-PanNuke is a node-level classification dataset derived from the PanNuke pan-cancer histology dataset. We use all slides at 40× magnification. Each tissue patch is converted into a cell-graph where nodes represent detected cell nuclei and edges encode spatial proximity. The task is predicting the cell type of each nucleus across 5 classes. Note that node features describe cell morphology, texture… See the full description on the dataset page: https://huggingface.co/datasets/ogutsevda/graph-pannuke.textgraph-ml1K<n<10K10 likes1.1k downloads7mo agoHugging Face06dszohib /graph-pannuke Graph-PanNuke: A Cell-Graph Dataset for Nucleus Classification from PanNuke Graph-PanNuke is a node-level classification dataset derived from the PanNuke pan-cancer histology dataset. We use all slides at 40× magnification. Each tissue patch is converted into a cell-graph where nodes represent detected cell nuclei and edges encode spatial proximity. The task is predicting the cell type of each nucleus across 5 classes. Note that node features describe cell morphology, texture… See the full description on the dataset page: https://huggingface.co/datasets/dszohib/graph-pannuke.textgraph-ml1K<n<10K0 likes280 downloads7mo agoHugging Face07multimodalart /panda-70m Panda 70M dataset by Snap Inc 70M video-caption pairs Code for downloading: https://github.com/snap-research/Panda-70M/dataset_dataloading textimage-to-text100K<n<1M15 likes216 downloads2y agoHugging Face08fsyfb /AI_Hype_Index_Panel_Data Multi-Agent AI Washing Index Panel Data for Chinese A-share Listed Firms, 2015-2024 Dataset Description This dataset provides firm-year panel measurements of AI washing among Chinese A-share listed companies from 2015 to 2024. It contains structured scores, qualitative classifications, adversarial multi-agent evaluation records, and verification evidence extracted from annual reports and firm-level AI capability indicators. The dataset is designed for academic… See the full description on the dataset page: https://huggingface.co/datasets/fsyfb/AI_Hype_Index_Panel_Data.imagetext-classification10K<n<100K0 likes213 downloads24d agoHugging Face09pankajbiswas6 /prism-hinglish-hate-speech PRISM - Code-Mixed Hinglish Hate-Speech Dataset Binary hate-speech dataset of code-mixed Hindi-English (Hinglish) text, used in the project Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text (RSET, The Assam Royal Global University). Source: combined_hate_speech_dataset on Kaggle. Companion model repository: Hinglish Hate-Speech Classification - BiLSTM / LSTM track Summary Attribute Value Total samples (raw) 29,550… See the full description on the dataset page: https://huggingface.co/datasets/pankajbiswas6/prism-hinglish-hate-speech.texttext-classification10K<n<100K0 likes137 downloads3mo agoHugging Face10letrinhan /imf-weo-fiscal-panel IMF WEO general-government fiscal panel (country × year) Country-year panel of IMF World Economic Outlook general-government fiscal indicators: revenue, expenditure, fiscal and primary balances, gross/net debt, and structural balance (% of GDP). Wide table is the primary product for panel regressions. a long table is included for extension. Includes weo_vintage, is_forecast, and actual_cutoff from WEO metadata. Figures Hero Comparison Files… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/imf-weo-fiscal-panel.tabulartabular-regression10K<n<100K0 likes136 downloads6d agoHugging Face11lbourdois /panlex PanLex January 1, 2024 version of PanLex Language Vocabulary with 24,650,274 rows covering 6,152 languages. Columns vocab: contains the text entry. 639-3: contains the ISO 639-3 languages tags to allow users to filter on the language(s) of their choice. 639-3_english_name: the English language name associated to the code ISO 639-3. var_code: contains a code to differentiate language variants. In practice, this is the code 639-3 + a number. If 000, it corresponds to… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/panlex.text10M<n<100M8 likes125 downloads3y agoHugging Face12chakra-labs /pango-sample Pango Sample: Real-World Computer Use Agent Training Data Pango represents Productivity Applications with Natural GUI Observations and trajectories. Dataset Description This dataset contains authentic computer interaction data collected from users performing real work tasks in productivity applications. The data was collected through Pango, a crowdsourced platform where users are compensated for contributing their natural computer interactions during actual work sessions.… See the full description on the dataset page: https://huggingface.co/datasets/chakra-labs/pango-sample.textn<1K4 likes115 downloads1y agoHugging Face13WABC /Panda-70Mtext1M<n<10M0 likes75 downloads4mo agoHugging Face14Panos21 /Tripadvisortext10K<n<100K0 likes73 downloads2y agoHugging Face15nameissakthi /tn-water-panels Tamil Nadu Water Panels Cleaned, analysis-ready hydrological series for the Cauvery basin and Tamil Nadu's major reservoirs, assembled from Indian government open data. Why this exists. The underlying data is public but not usable as published. The national water portal's CWC daily reservoir dataset covers only Odisha and Madhya Pradesh, and enumerating its Tamil Nadu resources returned no reservoir file. The archived reservoir bulletins are weekly PDFs across two incompatible… See the full description on the dataset page: https://huggingface.co/datasets/nameissakthi/tn-water-panels.tabular10K<n<100K0 likes68 downloads20d agoHugging Face16davecook1985 /solar-panel-yield-2026 Solar Panel Cleaning Yield Recovery — Datasets Open data companion to the Solar Panel Cleaning Yield Recovery working paper and reference calculator. Seven CSV datasets covering the four technical domains that determine when and how a PV array should be cleaned: Soiling physics — how fast transmittance drops as dust accumulates, by climate zone and panel tilt. Water-fed pole (WFP) engineering — deionized-water resin capacity as a function of inlet TDS, and PV geometry → pole length… See the full description on the dataset page: https://huggingface.co/datasets/davecook1985/solar-panel-yield-2026.tabulartabular-regressionn<1K0 likes67 downloads6mo agoHugging Face17YOUSIKI /PanoWantexttext-to-video10K<n<100K4 likes66 downloads10mo agoHugging Face18AkikJana /scramble-control-panels Scramble-control panels for cofolding confidence metrics Per-fold confidence scores for peptide–protein complexes, folded under Boltz-1, Boltz-2, Chai-1 and a few-step-distilled model, with each cognate peptide scored against permutations of itself as well as against unrelated decoys. 2,456 folds across 16 inference arms and 75 receptors. A permutation — a scramble — preserves amino-acid composition and length exactly and destroys only sequence order. Decoy comparisons cannot… See the full description on the dataset page: https://huggingface.co/datasets/AkikJana/scramble-control-panels.tabular1K<n<10K0 likes59 downloads1mo agoHugging Face19kartoun /Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0Acknowledgment: The dataset was created by Dr. Uri Kartoun. Description: The dataset was designed for the classification of text descriptions into seven stages of pancreatic cancer. It comprises two sets: a training set and a held-out set. Each set contains 700 blobs of text, with each blob representing a specific stage of pancreatic cancer. There are 100 text blobs for each of the seven defined stages in both files. Data Collection and Preparation: The text blobs were generated using… See the full description on the dataset page: https://huggingface.co/datasets/kartoun/Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0.text1K<n<10K0 likes53 downloads1y agoHugging Face20myothiha /pandas_table_qa_ft_v3text1K<n<10K1 likes52 downloads1y agoHugging Face21reddest-panda /arm-asmtexttext-generation1M<n<10M0 likes40 downloads2y agoHugging Face22future-edge-group /ipulse-ai-batch5-advisor-forecast-panel iPulse AI Batch 5 Advisor Forecast Panel This dataset exposes a compact, anonymized panel of production forecasts from iPulse AI, Future Edge Group's Open Agentic Investment Research Platform. It is designed for research on forecast combination, disagreement, correlated errors, regime dependence, and the effective number of independent forecasters. The release contains seven showcase assets, twelve advisor configurations per asset, quarterly forecast paths extending five years… See the full description on the dataset page: https://huggingface.co/datasets/future-edge-group/ipulse-ai-batch5-advisor-forecast-panel.tabular1K<n<10K0 likes39 downloads1mo agoHugging Face23PedroCuisinier2025 /OBD2_panel_opel_2012 📘 Dataset: OBD-II Telemetry – Opel Corsa 1.2 (2012) Real-world automotive telemetry recorded from a 2012 Opel Corsa (A12XER, 84 hp), collected using an ELM327 OBD-II adapter and python-OBD. 📊 Overview 394,406 rows 28 columns Time-ordered samples from 2025-04-30 → 2025-12-02 Sampling frequency: 3–12 Hz depending on PID latency Real OBD-II sensor readings + derived fields (fuel usage, torque, power, gear estimate) Each row corresponds to a single OBD-II polling cycle… See the full description on the dataset page: https://huggingface.co/datasets/PedroCuisinier2025/OBD2_panel_opel_2012.tabular100K<n<1M0 likes37 downloads10mo agoHugging Face24myothiha /pandas_table_qa_ft_v1tabularn<1K0 likes32 downloads1y agoHugging Face25costinflation /pantyliner-prices-raw-dataset-2026 25,466 raw U.S. pantyliner price observations across 12 ZIP markets and 29 days. Pantyliner Prices Raw Dataset (2026) Analyze 25,466 unaggregated product-level listed retail prices for disposable pantyliners across 12 U.S. ZIP markets from July 13 through August 10, 2026. The single analysis-ready CSV preserves titles, dates, geography, package quantities, listed prices, and a source-neutral comparable-price field. What “raw” means here: unaggregated product-level… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/pantyliner-prices-raw-dataset-2026.tabulartabular-regression10K<n<100K1 likes29 downloads1mo agoHugging Face26EstellaCheng42 /panel Panel: A Human Pairwise-Preference Benchmark for Open-Ended Dialogue Panel is a 1,800-pair human pairwise-preference benchmark for evaluating LLM-as-a-Judge systems in open-ended dialogue, introduced in the EMNLP 2026 paper: Multi-Expert Conformal Risk Control for Pairwise LLM Judging in Open-Ended Dialogue Ming Cheng, Yusheng Dai, Qiuhong Ke, Zhaolin Chen, Lizhen Qu (Monash University) All candidate responses are generated by open-weight LLMs, so judge logits are fully… See the full description on the dataset page: https://huggingface.co/datasets/EstellaCheng42/panel.texttext-classification1K<n<10K0 likes29 downloads1mo agoHugging Face27reddest-panda /arm-asm-xsmalltexttext-generation100K<n<1M0 likes24 downloads2y agoHugging Face28Pankaj8922 /nli-high-quality NLI High-Quality Balanced Dataset A combined, filtered, and class-balanced natural language inference (NLI) dataset built from MNLI, SNLI, FEVER-NLI, and ANLI, intended for fine-tuning NLI models for use in zero-shot text classification via the entailment trick (hypothesis = "This example is about {label}."). The goal of this dataset was quality and generalization over raw volume: rather than concatenating the four source datasets as-is, several filtering stages were applied to… See the full description on the dataset page: https://huggingface.co/datasets/Pankaj8922/nli-high-quality.texttext-classification100K<n<1M0 likes23 downloads2mo agoHugging Face29PanGD /lotus-QnAtextn<1K0 likes20 downloads2y agoHugging Face30panda04 /smart-home-datasettabular10K<n<100K0 likes20 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.