datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PHI-SPIKE-C172x-Community-Dataset-v1.0
PHI-SPIKE C172X Community Dataset v1.0
Dataset Summary
PHI-SPIKE C172X Community Dataset v1.0 is a simulation-based aerospace Prognostics and Health Management (PHM) dataset and training-artifact release developed from the PHI-SPIKE C172X research campaign.
The release provides:
JSBSim C172X reference telemetry;
benchmark metadata;
training histories;
trained PyTorch model checkpoints;
per-run evaluation metrics; and
five-seed campaign summaries.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/PHI-SPIKE-C172x-Community-Dataset-v1.0.spider-ko
Dataset Card for spider-ko: 한국어 Text-to-SQL 데이터셋
데이터셋 요약
Spider-KO는 Yale University의 Spider 데이터셋을 한국어로 번역한 텍스트-SQL 변환 데이터셋입니다. 원본 Spider 데이터셋의 자연어 질문을 한국어로 번역하여 구성하였습니다. 이 데이터셋은 다양한 도메인의 데이터베이스에 대한 질의와 해당 SQL 쿼리를 포함하고 있으며, 한국어 Text-to-SQL 모델 개발 및 평가에 활용될 수 있습니다.
지원 태스크 및 리더보드
text-to-sql: 한국어 자연어 질문을 SQL 쿼리로 변환하는 태스크에 사용됩니다.
언어
데이터셋의 질문은 한국어(ko)로 번역되었으며, SQL 쿼리는 영어 기반으로 유지되었습니다. 원본 영어 질문도 함께 제공됩니다.
데이터셋 구조
데이터 필드
db_id… See the full description on the dataset page: https://huggingface.co/datasets/huggingface-KREW/spider-ko.calcium-spike-inference-gcamp6f
Calcium spike inference, GCaMP6f
Two-photon and one-photon calcium-imaging recordings with ground-truth spikes, GCaMP6f only,
prepared for a spike-inference task. Eighty-one sessions across six behavioural paradigms, four
quality levels (typical, low SNR, high SNR, high motion) and two modalities, with per-session frame
rate (6.1 to 30.9 Hz), duration (50 to 403 s) and cell count (41 to 140) all varying.
heldout/heldout-000.npz ... heldout-026.npz 27 sessions: traces only, no… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/calcium-spike-inference-gcamp6f.whiskydb-fine-spirits-sample
🥃 WhiskyDB — Fine Spirits & Whisky Dataset (Free Sample)
Full dataset: whiskydb.dataengineered.io · $49 one-time (or $49 / month with the monthly refresh) → Buy once · Subscribe · the same sample on Kaggle
A free sample of WhiskyDB: a structured, relational dataset of whiskies and fine spirits built entirely from open, legally accessible public sources — government label registries (US TTB COLA), corporate registries (UK Companies House), the EU eAmbrosia GI register, Open… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/whiskydb-fine-spirits-sample.spiced
Dataset Card for SPICED
Dataset Summary
The Scientific Paraphrase and Information ChangE Dataset (SPICED) is a dataset of paired scientific findings from scientific papers, news media, and Twitter. The types of pairs are between <paper, news> and <paper, tweet>. Each pair is labeled for the degree of information similarity in the findings described by each sentence, on a scale from 1-5. This is called the Information Matching Score (IMS). The data was curated from S2ORC… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/spiced.spider-sql-promptsdatasette-spike-fara
Datasette spike — FARA Active Foreign Principals
For: CoS → WordPress Guru (doctorparadox.net embed/link)Built: 2026-09-17 (ET)Status: DATA half ready — public SQLite + Datasette Lite URL
Why this dataset
Doctor Paradox already centers corruption / foreign influence / authoritarian-adjacent reporting (Corruption Tracker, Corruption Daily cards). FARA filings are the federal public ledger of who lobbies in the U.S. on behalf of foreign principals.
We use the… See the full description on the dataset page: https://huggingface.co/datasets/doctorparadox/datasette-spike-fara.spiritual-indication-1a1e78
spiritual-indication-1a1e78
Synthetic products test data: 54 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/yeongceolgim/spiritual-indication-1a1e78.oil-spillsma-spinal-stimulation-figure-data
SMA Spinal-Stimulation Figure Data
This repository contains the figure-source data for a first-in-human study of epidural spinal-cord stimulation in three adults with SMA type III.
All data.zip is the 28,192,745-byte Zenodo v1 archive. It holds 28 CSV and spreadsheet files covering main Figures 2–6 and Extended Data Figures 2–8 and 10, including gait, electromyography, strength, and stimulation-response measurements. The accompanying MANIFEST.csv records the DOI, published MD5… See the full description on the dataset page: https://huggingface.co/datasets/YannisTevissen/sma-spinal-stimulation-figure-data.igaming-job-market
iGaming Job Market — open vacancies index by SpinHire
Snapshot date: 2026-09-18 · Open jobs: 6559 · Companies: 575 · License: CC BY 4.0
Live job postings in the iGaming industry (online casino, sports betting, game studios, affiliates,
payments, compliance) aggregated by SpinHire from employer career pages, ATS feeds
and public channels. The index is refreshed every 6 hours; jobs that disappear at the source are archived.
This dataset is a point-in-time export of the public API… See the full description on the dataset page: https://huggingface.co/datasets/Spinhire/igaming-job-market.spiritual-development-4b59d7
spiritual-development-4b59d7
Synthetic products test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting… See the full description on the dataset page: https://huggingface.co/datasets/Karen-Williams/spiritual-development-4b59d7.drive_statsangry-spirit-abdb80
angry-spirit-abdb80
Synthetic sensors test data: 30 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/minjunbag/angry-spirit-abdb80.Spica
Spica Dataset
Overview
This dataset contains image captioning data split into training, validation, and test sets. All samples in the split datasets (train, val, test) are contained in the spica_all_data.csv file.
Data Format
This dataset is provided as a single archive file:
spica.zip (13GB)
Contains approximately 75K files
Dataset Statistics
File
Samples
Unique Images
spica_all_data.csv
333,397
75,535
spica_train.csv
309,386
69,280… See the full description on the dataset page: https://huggingface.co/datasets/hiranohachiman/Spica.sql-spider-kaggledbqa-with-contextwikisql_and_spiderspidergene_expression_omnibus_nlpannotations_creators:
no-annotation
languages:
-English
All data pulled from Gene Expression Omnibus website. tab separated file with GSE number followed by title and abstract text.
spider_th
Spider Thai Dataset
Thai translation of the official Spider benchmark (A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task).
Dataset Description
This dataset contains Thai translations of the Spider text-to-SQL benchmark, translated from the official Spider data source.
Source
Original Dataset: Spider Benchmark
Paper: Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/spider_th.market-structural-fragility-spike-detection-v0.1
What this dataset tests
Detect structural fragility spikes before regime breaks.
The system must evaluatecross-asset relationshipsvolatility distortionsliquidity stresspositioning crowding.
Required outputs
fragility score
spike flag
time-to-break estimate
crowded trade pressure index
liquidity thinning index
Why it matters
Crashes are structural events.They occur when relationships compressand liquidity withdraws simultaneously.
This dataset trains… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/market-structural-fragility-spike-detection-v0.1.PragMegaPlusafrica-synth-energy-oilgas-spills-nigeria
Africa Synth Energy Oilgas Spills Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-oilgas-spills-nigeria.spidersr_spill_knowledgespider_wclinical-quad-pvalue-margin-secondary-endpoints-language-publication-pressure-spin-event-v0.1What this repo does
This dataset models statistical spin formation in clinical trial narratives. It predicts when the interaction between weak p-value margin, many secondary endpoints, high language intensity, and publication pressure indicates a high probability of a spin event where the narrative frames a weak result as strong.
Core quad
pvalue_margin_index
secondary_endpoint_count
language_intensity_index
publication_pressure_index
Prediction target
label_spin_event
Row structure
Each row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-pvalue-margin-secondary-endpoints-language-publication-pressure-spin-event-v0.1.spider_train_compact5_11testclinical-quad-ddi-polypharmacy-drift-exposure-spike-acute-safety-event-v0.1Clinical Quad DDI Polypharmacy Drift Exposure Spike Acute Safety Event v0.1
Each row is a patient snapshot.
Core quad
DDI riskPolypharmacy driftExposure spikeAcute safety event
Target
label_acute_safety_event_next_14d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.
The sample demonstrates the geometry.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-ddi-polypharmacy-drift-exposure-spike-acute-safety-event-v0.1.
