datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
panlex-meanings
Dataset Card for panlex-meanings
This is a dataset of words in several thousand languages, extracted from https://panlex.org.
Dataset Details
Dataset Description
This dataset has been extracted from https://panlex.org (the 20240301 database dump) and rearranged on the per-language basis.
Each language subset consists of expressions (words and phrases).
Each expression is associated with some meanings (if there is more than one meaning, they are in separate… See the full description on the dataset page: https://huggingface.co/datasets/gtak1/panlex-meanings.panlex-meanings
Dataset Card for panlex-meanings
This is a dataset of words in several thousand languages, extracted from https://panlex.org.
Dataset Details
Dataset Description
This dataset has been extracted from https://panlex.org (the 20240301 database dump) and rearranged on the per-language basis.
Each language subset consists of expressions (words and phrases).
Each expression is associated with some meanings (if there is more than one meaning, they are in separate… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/panlex-meanings.TCGA-PANCAN-HiSeq-2770x20530gene expression cancer RNA-Seq - Check the original submission: - https://www.synapse.org/Synapse:syn2812925 - is maintained by the cancer genome atlas pan-cancer analysis project. - TCGA-PANCAN-HiSeq-2770x20530
Files combined:
unc.edu_BRCA_IlluminaHiSeq_RNASeqV2.geneExp (20530, 957) BRCA
unc.edu_KIRC_IlluminaHiSeq_RNASeqV2.geneExp (20530, 552) KIRC
unc.edu_LUAD_IlluminaHiSeq_RNASeqV2.geneExp (20530, 413) LUAD
unc.edu_THCA_IlluminaHiSeq_RNASeqV2.geneExp (20530, 471) THCA… See the full description on the dataset page: https://huggingface.co/datasets/Fllamber/TCGA-PANCAN-HiSeq-2770x20530.panda-bench
PandaBench
PandaBench is a comprehensive benchmark for evaluating Large Language Model (LLM) safety, focusing on jailbreak attacks, defense mechanisms, and evaluation methodologies.
The PandaGuard framework architecture illustrating the end-to-end pipeline for LLM safety evaluation. The system connects three key components: Attackers, Defenders, and Judges.
Dataset Description
This repository contains the benchmark results from extensive evaluations of various… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/panda-bench.graph-pannuke
Graph-PanNuke: A Cell-Graph Dataset for Nucleus Classification from PanNuke
Graph-PanNuke is a node-level classification dataset derived from the PanNuke pan-cancer histology dataset. We use all slides at 40× magnification. Each tissue patch is converted into a cell-graph where nodes represent detected cell nuclei and edges encode spatial proximity. The task is predicting the cell type of each nucleus across 5 classes. Note that node features describe cell morphology, texture… See the full description on the dataset page: https://huggingface.co/datasets/ogutsevda/graph-pannuke.graph-pannuke
Graph-PanNuke: A Cell-Graph Dataset for Nucleus Classification from PanNuke
Graph-PanNuke is a node-level classification dataset derived from the PanNuke pan-cancer histology dataset. We use all slides at 40× magnification. Each tissue patch is converted into a cell-graph where nodes represent detected cell nuclei and edges encode spatial proximity. The task is predicting the cell type of each nucleus across 5 classes. Note that node features describe cell morphology, texture… See the full description on the dataset page: https://huggingface.co/datasets/dszohib/graph-pannuke.panda-70m
Panda 70M dataset by Snap Inc
70M video-caption pairs
Code for downloading: https://github.com/snap-research/Panda-70M/dataset_dataloading
AI_Hype_Index_Panel_Data
Multi-Agent AI Washing Index Panel Data for Chinese A-share Listed Firms, 2015-2024
Dataset Description
This dataset provides firm-year panel measurements of AI washing among Chinese A-share listed companies from 2015 to 2024. It contains structured scores, qualitative classifications, adversarial multi-agent evaluation records, and verification evidence extracted from annual reports and firm-level AI capability indicators.
The dataset is designed for academic… See the full description on the dataset page: https://huggingface.co/datasets/fsyfb/AI_Hype_Index_Panel_Data.prism-hinglish-hate-speech
PRISM - Code-Mixed Hinglish Hate-Speech Dataset
Binary hate-speech dataset of code-mixed Hindi-English (Hinglish) text, used in the project
Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text
(RSET, The Assam Royal Global University). Source: combined_hate_speech_dataset on Kaggle.
Companion model repository: Hinglish Hate-Speech Classification - BiLSTM / LSTM track
Summary
Attribute
Value
Total samples (raw)
29,550… See the full description on the dataset page: https://huggingface.co/datasets/pankajbiswas6/prism-hinglish-hate-speech.imf-weo-fiscal-panel
IMF WEO general-government fiscal panel (country × year)
Country-year panel of IMF World Economic Outlook general-government fiscal indicators: revenue, expenditure, fiscal and primary balances, gross/net debt, and structural balance (% of GDP). Wide table is the primary product for panel regressions. a long table is included for extension. Includes weo_vintage, is_forecast, and actual_cutoff from WEO metadata.
Figures
Hero
Comparison
Files… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/imf-weo-fiscal-panel.panlex
PanLex
January 1, 2024 version of PanLex Language Vocabulary with 24,650,274 rows covering 6,152 languages.
Columns
vocab: contains the text entry.
639-3: contains the ISO 639-3 languages tags to allow users to filter on the language(s) of their choice.
639-3_english_name: the English language name associated to the code ISO 639-3.
var_code: contains a code to differentiate language variants. In practice, this is the code 639-3 + a number. If 000, it corresponds to… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/panlex.pango-sample
Pango Sample: Real-World Computer Use Agent Training Data
Pango represents Productivity Applications with Natural GUI Observations and trajectories.
Dataset Description
This dataset contains authentic computer interaction data collected from users performing real work tasks in productivity applications. The data was collected through Pango, a crowdsourced platform where users are compensated for contributing their natural computer interactions during actual work sessions.… See the full description on the dataset page: https://huggingface.co/datasets/chakra-labs/pango-sample.Panda-70MTripadvisortn-water-panels
Tamil Nadu Water Panels
Cleaned, analysis-ready hydrological series for the Cauvery basin and Tamil Nadu's
major reservoirs, assembled from Indian government open data.
Why this exists. The underlying data is public but not usable as published. The
national water portal's CWC daily reservoir dataset covers only Odisha and Madhya
Pradesh, and enumerating its Tamil Nadu resources returned no reservoir file. The
archived reservoir bulletins are weekly PDFs across two incompatible… See the full description on the dataset page: https://huggingface.co/datasets/nameissakthi/tn-water-panels.solar-panel-yield-2026
Solar Panel Cleaning Yield Recovery — Datasets
Open data companion to the Solar Panel Cleaning Yield Recovery working paper and reference calculator. Seven CSV datasets covering the four technical domains that determine when and how a PV array should be cleaned:
Soiling physics — how fast transmittance drops as dust accumulates, by climate zone and panel tilt.
Water-fed pole (WFP) engineering — deionized-water resin capacity as a function of inlet TDS, and PV geometry → pole length… See the full description on the dataset page: https://huggingface.co/datasets/davecook1985/solar-panel-yield-2026.PanoWanscramble-control-panels
Scramble-control panels for cofolding confidence metrics
Per-fold confidence scores for peptide–protein complexes, folded under Boltz-1,
Boltz-2, Chai-1 and a few-step-distilled model, with each cognate peptide scored
against permutations of itself as well as against unrelated decoys.
2,456 folds across 16 inference arms and 75 receptors.
A permutation — a scramble — preserves amino-acid composition and length
exactly and destroys only sequence order. Decoy comparisons cannot… See the full description on the dataset page: https://huggingface.co/datasets/AkikJana/scramble-control-panels.Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0Acknowledgment: The dataset was created by Dr. Uri Kartoun.
Description: The dataset was designed for the classification of text descriptions into seven stages of pancreatic cancer. It comprises two sets: a training set and a held-out set. Each set contains 700 blobs of text, with each blob representing a specific stage of pancreatic cancer. There are 100 text blobs for each of the seven defined stages in both files.
Data Collection and Preparation: The text blobs were generated using… See the full description on the dataset page: https://huggingface.co/datasets/kartoun/Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0.pandas_table_qa_ft_v3arm-asmipulse-ai-batch5-advisor-forecast-panel
iPulse AI Batch 5 Advisor Forecast Panel
This dataset exposes a compact, anonymized panel of production forecasts from iPulse AI, Future Edge Group's Open Agentic Investment Research Platform. It is designed for research on forecast combination, disagreement, correlated errors, regime dependence, and the effective number of independent forecasters.
The release contains seven showcase assets, twelve advisor configurations per asset, quarterly forecast paths extending five years… See the full description on the dataset page: https://huggingface.co/datasets/future-edge-group/ipulse-ai-batch5-advisor-forecast-panel.OBD2_panel_opel_2012
📘 Dataset: OBD-II Telemetry – Opel Corsa 1.2 (2012)
Real-world automotive telemetry recorded from a 2012 Opel Corsa (A12XER, 84 hp), collected using an ELM327 OBD-II adapter and python-OBD.
📊 Overview
394,406 rows
28 columns
Time-ordered samples from 2025-04-30 → 2025-12-02
Sampling frequency: 3–12 Hz depending on PID latency
Real OBD-II sensor readings + derived fields (fuel usage, torque, power, gear estimate)
Each row corresponds to a single OBD-II polling cycle… See the full description on the dataset page: https://huggingface.co/datasets/PedroCuisinier2025/OBD2_panel_opel_2012.pandas_table_qa_ft_v1pantyliner-prices-raw-dataset-2026
25,466 raw U.S. pantyliner price observations across 12 ZIP markets and 29 days.
Pantyliner Prices Raw Dataset (2026)
Analyze 25,466 unaggregated product-level listed retail prices for disposable pantyliners across 12 U.S. ZIP markets from July 13 through August 10, 2026. The single analysis-ready CSV preserves titles, dates, geography, package quantities, listed prices, and a source-neutral comparable-price field.
What “raw” means here: unaggregated product-level… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/pantyliner-prices-raw-dataset-2026.panel
Panel: A Human Pairwise-Preference Benchmark for Open-Ended Dialogue
Panel is a 1,800-pair human pairwise-preference benchmark for evaluating LLM-as-a-Judge systems in open-ended dialogue, introduced in the EMNLP 2026 paper:
Multi-Expert Conformal Risk Control for Pairwise LLM Judging in Open-Ended Dialogue
Ming Cheng, Yusheng Dai, Qiuhong Ke, Zhaolin Chen, Lizhen Qu (Monash University)
All candidate responses are generated by open-weight LLMs, so judge logits are fully… See the full description on the dataset page: https://huggingface.co/datasets/EstellaCheng42/panel.arm-asm-xsmallnli-high-quality
NLI High-Quality Balanced Dataset
A combined, filtered, and class-balanced natural language inference (NLI)
dataset built from MNLI, SNLI, FEVER-NLI, and ANLI, intended for fine-tuning
NLI models for use in zero-shot text classification via the entailment trick
(hypothesis = "This example is about {label}.").
The goal of this dataset was quality and generalization over raw volume:
rather than concatenating the four source datasets as-is, several filtering
stages were applied to… See the full description on the dataset page: https://huggingface.co/datasets/Pankaj8922/nli-high-quality.lotus-QnAsmart-home-dataset
