datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
panlex-meanings
Dataset Card for panlex-meanings
This is a dataset of words in several thousand languages, extracted from https://panlex.org.
Dataset Details
Dataset Description
This dataset has been extracted from https://panlex.org (the 20240301 database dump) and rearranged on the per-language basis.
Each language subset consists of expressions (words and phrases).
Each expression is associated with some meanings (if there is more than one meaning, they are in separate… See the full description on the dataset page: https://huggingface.co/datasets/gtak1/panlex-meanings.panlex-meanings
Dataset Card for panlex-meanings
This is a dataset of words in several thousand languages, extracted from https://panlex.org.
Dataset Details
Dataset Description
This dataset has been extracted from https://panlex.org (the 20240301 database dump) and rearranged on the per-language basis.
Each language subset consists of expressions (words and phrases).
Each expression is associated with some meanings (if there is more than one meaning, they are in separate… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/panlex-meanings.TCGA-PANCAN-HiSeq-2770x20530gene expression cancer RNA-Seq - Check the original submission: - https://www.synapse.org/Synapse:syn2812925 - is maintained by the cancer genome atlas pan-cancer analysis project. - TCGA-PANCAN-HiSeq-2770x20530
Files combined:
unc.edu_BRCA_IlluminaHiSeq_RNASeqV2.geneExp (20530, 957) BRCA
unc.edu_KIRC_IlluminaHiSeq_RNASeqV2.geneExp (20530, 552) KIRC
unc.edu_LUAD_IlluminaHiSeq_RNASeqV2.geneExp (20530, 413) LUAD
unc.edu_THCA_IlluminaHiSeq_RNASeqV2.geneExp (20530, 471) THCA… See the full description on the dataset page: https://huggingface.co/datasets/Fllamber/TCGA-PANCAN-HiSeq-2770x20530.panda-bench
PandaBench
PandaBench is a comprehensive benchmark for evaluating Large Language Model (LLM) safety, focusing on jailbreak attacks, defense mechanisms, and evaluation methodologies.
The PandaGuard framework architecture illustrating the end-to-end pipeline for LLM safety evaluation. The system connects three key components: Attackers, Defenders, and Judges.
Dataset Description
This repository contains the benchmark results from extensive evaluations of various… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/panda-bench.AI_Hype_Index_Panel_Data
Multi-Agent AI Washing Index Panel Data for Chinese A-share Listed Firms, 2015-2024
Dataset Description
This dataset provides firm-year panel measurements of AI washing among Chinese A-share listed companies from 2015 to 2024. It contains structured scores, qualitative classifications, adversarial multi-agent evaluation records, and verification evidence extracted from annual reports and firm-level AI capability indicators.
The dataset is designed for academic… See the full description on the dataset page: https://huggingface.co/datasets/fsyfb/AI_Hype_Index_Panel_Data.imf-weo-fiscal-panel
IMF WEO general-government fiscal panel (country × year)
Country-year panel of IMF World Economic Outlook general-government fiscal indicators: revenue, expenditure, fiscal and primary balances, gross/net debt, and structural balance (% of GDP). Wide table is the primary product for panel regressions. a long table is included for extension. Includes weo_vintage, is_forecast, and actual_cutoff from WEO metadata.
Figures
Hero
Comparison
Files… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/imf-weo-fiscal-panel.tn-water-panels
Tamil Nadu Water Panels
Cleaned, analysis-ready hydrological series for the Cauvery basin and Tamil Nadu's
major reservoirs, assembled from Indian government open data.
Why this exists. The underlying data is public but not usable as published. The
national water portal's CWC daily reservoir dataset covers only Odisha and Madhya
Pradesh, and enumerating its Tamil Nadu resources returned no reservoir file. The
archived reservoir bulletins are weekly PDFs across two incompatible… See the full description on the dataset page: https://huggingface.co/datasets/nameissakthi/tn-water-panels.solar-panel-yield-2026
Solar Panel Cleaning Yield Recovery — Datasets
Open data companion to the Solar Panel Cleaning Yield Recovery working paper and reference calculator. Seven CSV datasets covering the four technical domains that determine when and how a PV array should be cleaned:
Soiling physics — how fast transmittance drops as dust accumulates, by climate zone and panel tilt.
Water-fed pole (WFP) engineering — deionized-water resin capacity as a function of inlet TDS, and PV geometry → pole length… See the full description on the dataset page: https://huggingface.co/datasets/davecook1985/solar-panel-yield-2026.scramble-control-panels
Scramble-control panels for cofolding confidence metrics
Per-fold confidence scores for peptide–protein complexes, folded under Boltz-1,
Boltz-2, Chai-1 and a few-step-distilled model, with each cognate peptide scored
against permutations of itself as well as against unrelated decoys.
2,456 folds across 16 inference arms and 75 receptors.
A permutation — a scramble — preserves amino-acid composition and length
exactly and destroys only sequence order. Decoy comparisons cannot… See the full description on the dataset page: https://huggingface.co/datasets/AkikJana/scramble-control-panels.ipulse-ai-batch5-advisor-forecast-panel
iPulse AI Batch 5 Advisor Forecast Panel
This dataset exposes a compact, anonymized panel of production forecasts from iPulse AI, Future Edge Group's Open Agentic Investment Research Platform. It is designed for research on forecast combination, disagreement, correlated errors, regime dependence, and the effective number of independent forecasters.
The release contains seven showcase assets, twelve advisor configurations per asset, quarterly forecast paths extending five years… See the full description on the dataset page: https://huggingface.co/datasets/future-edge-group/ipulse-ai-batch5-advisor-forecast-panel.OBD2_panel_opel_2012
📘 Dataset: OBD-II Telemetry – Opel Corsa 1.2 (2012)
Real-world automotive telemetry recorded from a 2012 Opel Corsa (A12XER, 84 hp), collected using an ELM327 OBD-II adapter and python-OBD.
📊 Overview
394,406 rows
28 columns
Time-ordered samples from 2025-04-30 → 2025-12-02
Sampling frequency: 3–12 Hz depending on PID latency
Real OBD-II sensor readings + derived fields (fuel usage, torque, power, gear estimate)
Each row corresponds to a single OBD-II polling cycle… See the full description on the dataset page: https://huggingface.co/datasets/PedroCuisinier2025/OBD2_panel_opel_2012.pandas_table_qa_ft_v1pantyliner-prices-raw-dataset-2026
25,466 raw U.S. pantyliner price observations across 12 ZIP markets and 29 days.
Pantyliner Prices Raw Dataset (2026)
Analyze 25,466 unaggregated product-level listed retail prices for disposable pantyliners across 12 U.S. ZIP markets from July 13 through August 10, 2026. The single analysis-ready CSV preserves titles, dates, geography, package quantities, listed prices, and a source-neutral comparable-price field.
What “raw” means here: unaggregated product-level… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/pantyliner-prices-raw-dataset-2026.smart-home-datasetUSA_100_Panos
Dataset
Panitikan-34
Panitikan-34 Corpus
Panitikan-34 contains Filipino literary texts from 34 Filipino authors from the late 19th to early 20th century. The data was gathered from the Tagalog books category of Project Gutenberg using the scrapy library. Various pre-processing techniques were also applied to the dataset which can also be adopted in other languages as discussed in the paper. Dictionaries, thesauruses, and works translated from other languages were excluded to solely focus on literary… See the full description on the dataset page: https://huggingface.co/datasets/Galahallt/Panitikan-34.Panitikan-10
Panitikan-10 Corpus
Panitikan-10 contains Filipino literary texts from 10 Filipino authors from the late 19th to early 20th century. The data was gathered from the Tagalog books category of Project Gutenberg using the scrapy library. Various pre-processing techniques were also applied to the dataset which can also be adopted in other languages as discussed in the paper. Dictionaries, thesauruses, and works translated from other languages were excluded to solely focus on literary… See the full description on the dataset page: https://huggingface.co/datasets/Galahallt/Panitikan-10.PANCANPantip_QA_200000_20220220demolegal-judicial-panel-coherence-convergence-v0.1What this dataset is
You receive
majority frame
concurrence frame
dissent frame
fracture signals
review signals
You decide
Does the panel converge on the same core legal frame
Answer
coherent
or
incoherent
Why this matters
Fractured panels predict
en banc risk
higher court review risk
weak precedent
unstable doctrine
RAG12000-LLaMA3.1-8B-gguf_AR-RAG_v2datasettesthfragas_evaluationV1Dataset-responsesv2with-evaluationpanaderias.csvDatos del ejemplo ficticio de una panadería, tomada del libro "The Manga Guide to Regression Analysis" de Shin Takahashi, Iroha Inoue
movies1213
