datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trial-v0-20250313
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/team-wonders/trial-v0-20250313.tripitaka-siamrath
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกภาษาไทยฉบับสยามรัฏฐ จำนวน 45 เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม 1 (754 หน้า): พระวินัยปิฎก เล่ม ๑ มหาวิภังค์ ปฐมภาค
เล่ม 2 (717 หน้า): พระวินัยปิฎก เล่ม ๒ มหาวิภังค์ ทุติภาค
เล่ม 3 (328 หน้า): พระวินัยปิฎก เล่ม ๓ ภิกขุณี วิภังค์
เล่ม 4 (304 หน้า): พระวินัยปิฎก เล่ม ๔ มหาวรรคภาค ๑
เล่ม 5 (278… See the full description on the dataset page: https://huggingface.co/datasets/uisp/tripitaka-siamrath.rebus-dataset
|🔄 🚍| Re-Bus: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
Understanding Rebus Puzzles requires a variety of skills such as image recognition, cognitive skills, commonsense reasoning, and multi-step reasoning, making this a challenging task for current Vision-Language Models. In this paper, we present Re-Bus, a large and diverse benchmark of 1,333 English Rebus Puzzles containing different artistic… See the full description on the dataset page: https://huggingface.co/datasets/TrishanuDas/rebus-dataset.pali-tripitaka-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๕ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
...
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินย. มหาวิภงฺโค (๑)
เล่ม ๒: วินย. มหาวิภงฺโค (๒)
เล่ม ๓: วินย. ภิกฺขุนีวิภงฺโค
เล่ม ๔: วินย. มหาวคฺโค (๑)
เล่ม ๕: วินย. มหาวคฺโค (๒)
เล่ม ๖: วินย. จุลฺลวคฺโค (๑)
เล่ม ๗: วินย. จุลฺลวคฺโค (๒)
เล่ม ๘: วินย.… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-tripitaka-thai-script-siamrath-version.trilemma-of-truth
Dataset Card for Trilemma of Truth (ToT) Dataset
🧾 Dataset Summary
The Trilemma of Truth (ToT) dataset serves as a benchmark for evaluating veracity probes across three distinct statement types:
Factually true statements.
Factually false statements.
Neither-valued statements are defined as those for which the language model lacks sufficient evidence to assign a truth value (see formal definition below).
The dataset includes three domain configurations:… See the full description on the dataset page: https://huggingface.co/datasets/carlomarxx/trilemma-of-truth.tri-attia-fast-charge-2020-raw
Closed-loop optimization of extreme fast charging for batteries
BSEBench status: raw_mirror_pending_validation
This dataset repository is a raw mirror of the TRI Energy & Materials Data / data.matr.io project "Closed-loop optimization of extreme fast charging for batteries using machine learning". The source describes commercial lithium-ion phosphate (LFP)/graphite cells cycled under fast-charging conditions for the associated Nature publication.
Source page:… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/tri-attia-fast-charge-2020-raw.clinical-trials-xml-2018-2024nyc_taxi_trip_2024_p1_sampletripclick-training
TripClick Baselines with Improved Training Data
Establishing Strong Baselines for TripClick Health Retrieval Sebastian Hofstätter, Sophia Althammer, Mete Sertkan and Allan Hanbury
https://arxiv.org/abs/2201.00365
tl;dr We create strong re-ranking and dense retrieval baselines (BERTCAT, BERTDOT, ColBERT, and TK) for TripClick (health ad-hoc retrieval). We improve the – originally too noisy – training data with a simple negative sampling policy. We achieve large gains over BM25 in the… See the full description on the dataset page: https://huggingface.co/datasets/sebastian-hofstaetter/tripclick-training.pali-tripitaka-thai-script-siamrath-version
📚 พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฐ (๔๕ เล่ม)
ต้นฉบับสามารถเข้าถึงได้ที่: 84000 พระธรรมขันธ์ 84000.org พระไตรปิฎก
🧾 รายการพระไตรปิฎก
📘 เล่ม ๑–๘: วินัยปิฎก
เล่ม ๑: มหาวิภงฺโค (๑)
เล่ม ๒: มหาวิภงฺโค (๒)
เล่ม ๓: ภิกฺขุนีวิภงฺโค
เล่ม ๔: มหาวคฺโค (๑)
เล่ม ๕: มหาวคฺโค (๒)
เล่ม ๖: จุลฺลวคฺโค (๑)
เล่ม ๗: จุลฺลวคฺโค (๒)
เล่ม ๘: ปริวาโร
📗 เล่ม ๙–๒๕: สุตตันตปิฎก
เล่ม ๙–๑๑: ทีฆนิกาย
เล่ม ๑๒–๑๔: มัชฌิมนิกาย
เล่ม ๑๕–๑๙: สังยุตตนิกาย… See the full description on the dataset page: https://huggingface.co/datasets/mgprogm/pali-tripitaka-thai-script-siamrath-version.multi-destination-trip-dataset
Intro
Booking.com provides a unique dataset based on millions of real anonymized bookings to encourage the research on sequential recommendation problems.
Many travelers go on trips which include more than one destination. Our mission at Booking.com is to make it easier for everyone to experience the world, and we can help to do that by providing real-time recommendations for what their next in-trip destination will be. By making accurate predictions, we help deliver a frictionless… See the full description on the dataset page: https://huggingface.co/datasets/Booking-com/multi-destination-trip-dataset.clinical-trial-outcomes-predictions
Clinical Trial Outcomes Prediction Dataset
A dataset of 1,366 binary forecasting questions about clinical trial outcomes, automatically generated and labeled using Lightning Rod Labs' Future-as-Label methodology.
Dataset Description
This dataset contains questions about pharmaceutical clinical trials from 2023-2024, paired with verified outcomes (success/failure). Each question asks whether a specific trial will meet its endpoints, receive FDA approval, or complete by a… See the full description on the dataset page: https://huggingface.co/datasets/3rdSon/clinical-trial-outcomes-predictions.PhotonicDeepBeam-dataset
PhotonicDeepBeam Dataset
This repository hosts the datasets used in the QuantBeam project — an adaptation of the DeepBeam architecture towards Photonic-Aware Neural Networks (PANN) for millimeter-wave (mmWave) beam classification.
⚠️ The datasets are NOT ours. They are the original experimental datasets collected and published by the DeepBeam / WiNES Lab team (Polese, Restuccia, Melodia — Northeastern University). We re-host them here solely to facilitate reproducibility of our… See the full description on the dataset page: https://huggingface.co/datasets/triani/PhotonicDeepBeam-dataset.bjj-kimura-lesson001-trial
BJJ Kimura from Side Control — Research Trial (sampled across the action arc)
Tier: Research / Evaluation (free, CC BY-NC-SA 4.0)
Source: RTK Motion Intelligence Platform · api.rtkmotion.io
Full commercial dataset: rtk-training/bjj-kimura-lesson001 (gated)
A temporally-sampled trial subset of a full 4D motion-capture session
of a Brazilian Jiu-Jitsu Kimura submission from side control, demonstrated
by a Former IBJJF World Champion (anonymized) with a training partner.… See the full description on the dataset page: https://huggingface.co/datasets/rtk-training/bjj-kimura-lesson001-trial.phi_so101_8bin_v1_trim
phi_so101_8bin_v1 — opening-pause trim table
Companion to BrutalCaesar/phi_so101_8bin_v1.
This is not a dataset. It is a 119-row table plus the script that produced it. The original
dataset is unmodified and remains authoritative. Applying this table excludes each episode's
pre-teleop dead air as a chunk start point, without deleting a single frame from disk.
Why
Every episode begins with the arm sitting still while the operator has not yet moved the leader.… See the full description on the dataset page: https://huggingface.co/datasets/Parv-09/phi_so101_8bin_v1_trim.202407-citibike-tripdatachina-myeloma-clinical-trials
China Multiple Myeloma Clinical Trials — Open Dataset
Multiple myeloma clinical trials registered in China, curated from official NMPA / CDE filings by the China Myeloma Digital Network (CMDN), an independent non-profit patient advocacy organisation.
This is a mirror. The citable version of record lives at doi.org/10.5281/zenodo.22690814; the source repository is chinamyeloma/china-myeloma-clinical-trials; the documentation is at chinamyeloma.org.
Why this exists… See the full description on the dataset page: https://huggingface.co/datasets/chinamyeloma/china-myeloma-clinical-trials.GDP-Per-Capita_Gov-Expenditure_TradeAbout Dataset
Edit
This dataset provides a comprehensive view of key macroeconomic indicators across various entities (countries or regions) over time. It includes annual data for the following variables:
Entity: The name of the country or region for which the data is recorded.
Code: A standardized three-letter country or region code, facilitating easier identification and merging with other datasets.
Year: The calendar year for which the economic indicators are reported.
GDP per capita:… See the full description on the dataset page: https://huggingface.co/datasets/tripathyShaswata/GDP-Per-Capita_Gov-Expenditure_Trade.clinical-quad-unblinding-sae-cluster-media-leak-trial-halt-decision-v0.1Clinical Quad Unblinding SAE Cluster Media Leak Trial Halt Decision v0.1
Each row is a site weekly snapshot.
Core quad
Emergency unblindingSAE clusterMedia leak riskTrial halt decision risk
Target
label_trial_halt_risk_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
flow-edit-cube-triple-tkgrid-seeds30003-40004
Edit-placement (t x K x eb) campaign — OGBench cube-triple, task 2 — seeds 30003 & 40004
Seed scope: this repository contains only seeds 30003 and 40004. It is not the
full seed set for this campaign — seeds 10001 and 20002 were trained on separate hardware
and are not included here. Any per-cell mean computed from this repo alone is an n=2
estimate; see Caveats.
290 training runs from the uedit_place agent: a grid over where in the flow a
value-driven edit is applied (t), how… See the full description on the dataset page: https://huggingface.co/datasets/jaehyeokdoo2/flow-edit-cube-triple-tkgrid-seeds30003-40004.clinical-trial-outcomes-2020plus
Clinical Trial Outcomes (2020+) with Normalized Endpoints
124,790 normalized endpoints across 14,170 clinical studies that started on or after
2020-01-01 and have posted results on ClinicalTrials.gov.
Snapshot: 2026-09-08. Source: ClinicalTrials.gov API v2 (U.S. National Library of Medicine).
Built with ctgov — the same normalizer, released as an
MIT-licensed package with zero dependencies. So this snapshot is not a dead artifact: you can
re-run it against the live registry, or… See the full description on the dataset page: https://huggingface.co/datasets/GooseWithStories/clinical-trial-outcomes-2020plus.grading-question-triage-datasetdepression-trials-pubmed-evidence
MatrixNorm TrialEvidence — Depression
Structured, evidence-linked clinical-trial + biomedical-literature data extracted by MatrixNorm. Every record is provenance-stamped (source_url, raw_hash), carries its evidence trail (reasoning_atoms), and is honesty-gated — a claim is only statistically_supported when a p-value/CI/effect was present in the source, and normalized values keep their verbatim source_original_value.
Quality audit: PASS — 0 error(s), 0 warning(s), 5 note(s). See… See the full description on the dataset page: https://huggingface.co/datasets/williamTLmiller/depression-trials-pubmed-evidence.epa-tri-toxic-release-inventory
EPA Toxic Release Inventory — 2022–2023 reporting years
Rebuilt September 19, 2026 from the two pinned EPA national Basic Data CSV files.
This is a dated historical snapshot; future updates are not included. The archive
contains 158,687 TRI form records, representing 23,081 distinct TRI facilities.
It does not contain the previously advertised 1987-onward history.
This repository contains the deterministic 1,000-row sample. The complete dated
CSV/Parquet package is available at… See the full description on the dataset page: https://huggingface.co/datasets/claritystorm/epa-tri-toxic-release-inventory.Tripadvisor_stadium_reviews_P5_schoolsglp1-clinical-trials-2026
GLP-1 Clinical Trials Index — 2026
Publisher: Ozari Health | ozarihealth.comLicense: CC-BY-4.0Last updated: May 2026Rows: 22 trialsHuggingFace: https://huggingface.co/datasets/Ozarihealth/glp1-clinical-trials-2026
About This Dataset
A structured index of every major published GLP-1 clinical trial as of May 2026. This dataset consolidates trial-level data from peer-reviewed publications into a single comparable format — covering semaglutide, tirzepatide, liraglutide, oral… See the full description on the dataset page: https://huggingface.co/datasets/Ozarihealth/glp1-clinical-trials-2026.africa-synth-cancer-triple-negative-breast-cancer-gene-all
TNBC Gene Expression Profiles - Sub-Saharan African Women | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cancer-triple-negative-breast-cancer-gene-all.pre-data-room-triage-workbook
Pre-data-room triage workbook
This open workbook helps corporate venture, corporate development, innovation,
and M&A teams turn an inbound opportunity into a documented first-pass evidence
agenda before committing to a full data-room process.
It is a triage aid, not a valuation, investment recommendation, legal review,
or replacement for full due diligence. Do not put confidential target data in
the public dataset; download the blank workbook and use it in the team's own… See the full description on the dataset page: https://huggingface.co/datasets/mheilimo/pre-data-room-triage-workbook.clinical-trial-basin-enrichment-and-recruitment-v0.1What this dataset tests
Whether a system can design basin selective recruitment.
It must:
pick the eligible basin
write inclusion and exclusion rules
estimate signal gain and dilution risk
Required outputs
eligible_basin_definition
recruitment_filter_rules
signal_amplification_index
heterogeneity_dilution_risk
expected_trial_coherence_gain
exclusion_rationale
Use case
Trial rescue.
Smaller trials with stronger signals.
Reduced washout from basin mixing.
legal-limitation-period-trigger-tolling-coherence-risk-v0.1What this dataset does
You receive
cause and dates
rules
trigger analysis
tolling or standstill
issue date
You decide
coherent
or
incoherent
Daily use
limitation QC
urgent filing flag
tolling gap detection
escalation trigger
