datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Linguistic-Diagnostics-Pragmatics
LINDSEA Pragmatics
LINDSEA Pragmatics is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, pragmatics in particular, for Indonesian.
Supported Tasks and Leaderboards
LINDSEA Pragmatics is designed for evaluating chat or instruction-tuned large language models (LLMs).
Languages
Indonesian (id)
Dataset Details
LINDSEA Pragmatics only has an Indonesian (id) split, with additional splits containing… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Pragmatics.Linguistic-Diagnostics-Syntax
LINDSEA Syntax
LINDSEA Syntax is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, syntax in particular, for Indonesian.
Supported Tasks and Leaderboards
LINDSEA Syntax is designed for evaluating chat or instruction-tuned large language models (LLMs).
Languages
Indonesian (id)
Dataset Details
LINDSEA Syntax only has an Indonesian (id) split, with additional splits containing fewshot examples. Below… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Syntax.grokking-diagnostics-runs
Grokking Diagnostics Runs
Per-run training records and aggregate fits backing:
Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
Lucky Verma. Independent Researcher. 2026.
Paper ·
DOI ·
PDF ·
Code
Contents
The paper provenance indexes 1,792 paper-run records: 1,442 records from the
main paper-integrated run tree plus 350 cross-architecture scope-probe records.
This dataset repository also includes convenience subset mirrors, so the… See the full description on the dataset page: https://huggingface.co/datasets/lucky-verma/grokking-diagnostics-runs.Linguistic-Diagnostics-Syntax-Judge
LINDSEA Syntax
LINDSEA Syntax is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, syntax in particular, for Indonesian.
Supported Tasks and Leaderboards
LINDSEA Syntax is designed for evaluating chat or instruction-tuned large language models (LLMs).
Languages
Indonesian (id)
Dataset Details
Data Sources
Data Source
License
Language/s
Split/s
CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Syntax-Judge.crossarch-1b-diagnostics
Cross-architecture mergeability diagnostics for five ~1B monolingual LMs
A third model family for the mergeability project, alongside Goldfish and Beetle/MergeBench.
Read this first: what "merging" means for these five models
Five independently trained ~1B monolingual models were requested: Pythia-1.4B (EN),
Zh-Pythia-1.4B (ZH), Tucano-1b1 (PT), Bielik-1.5B-v3 (PL), Minerva-1B (IT). They differ in
architecture family, hidden dimension (1536 vs 2048), depth (16 /… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/crossarch-1b-diagnostics.painting-restoration-eval-diagnosticsglue_diagnostics
Overview
Original dataset available here.
Dataset curation
Filled in the empty rows of columns "lexical semantics", "predicate-argument structure",
"logic", "knowledge" with empty string "".
Labels are encoded as follows
{"entailment": 0, "neutral": 1, "contradiction": 2}
Code to create dataset
import pandas as pd
from datasets import Features, Value, ClassLabel, Dataset
df = pd.read_csv("<path to file>/diagnostic-full.tsv", sep="\t")
# column names to… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/glue_diagnostics.lexiconerror-diagnostics
LexiconError Diagnostics
LexiconError Diagnostics is a 16,474-record, provenance-preserving corpus of programming-language compiler diagnostics, linter rules, runtime exceptions, infrastructure failures, and accelerator-runtime faults. It is a reference dataset, not a claim of exhaustive or fully editorially verified coverage.
Dataset details
Release version: 2026.08
Records: 16,474 across 44 languages and 45 tools
Records marked verified: 41; generated registry… See the full description on the dataset page: https://huggingface.co/datasets/Magnexis/lexiconerror-diagnostics.mi-064-local-openbook-diagnostics-staging
MI-064 Bounded Diagnostic Methodology v3.1 — Bridge Review
This is a sanitized public bridge-review package for documenting MI-064 Taiwan-alignment diagnostic methodology, comparability corrections, and the Phase3.1 TW-v4a bridge diagnostic.
Status: public_sanitized_append_only_update_authorized__phase4_final_revision_reconciled
Read this first
Traditional-Chinese visual reader guide: click the diagram above or open the HF Static Space.
繁中 Dataset Card:… See the full description on the dataset page: https://huggingface.co/datasets/rickytzai/mi-064-local-openbook-diagnostics-staging.p1b-mcmc-diagnostics
P1B MCMC Chain Diagnostics
Title: MCMC chain diagnostics and convergence CSV files
Description
This dataset contains the top-level MCMC chain diagnostic and convergence summary files
produced by the Cobaya MCMC sampler for the P1B cosmological parameter inference runs.
It corresponds to Appendix A, item 1 of the paper
"P1B — Technical Verification Companion to the ECH Spin-Torsion Program" (v1B.0.75).
Files include convergence CSVs (R-1 statistics per chain)… See the full description on the dataset page: https://huggingface.co/datasets/bamfai/p1b-mcmc-diagnostics.glue_diagnostics
Citation
@inproceedings{wang2019glue, title={{GLUE}: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding}, author={Wang, Alex and Singh, Amanpreet and Michael, Julian and Hill, Felix and Levy, Omer and Bowman, Samuel R.}, note={In the Proceedings of ICLR.}, year={2019}}
progressive-8k-30nfe-diagnosticsafrica-synth-tuberculosis-tb-laboratory-diagnostics-all
TB Laboratory Diagnostics | Africa (World Health Organization)
Size category: 10K<n<100K - Formats: csv - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-tuberculosis-tb-laboratory-diagnostics-all.automotive-diagnostics-v2
Automotive Diagnostics AI Dataset v2
Description
Training dataset for Automotive Diagnostics AI Assistant.
Built from real internet sources using automated text parsing.
No manual data entry used in dataset construction.
Data Sources
OBD-II documentation from GitHub repositories
UDS ISO 14229 protocol reference documents
ISO 26262 functional safety standard reference
NHTSA US Government vehicle complaint API
Dataset Structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/RRK1987/automotive-diagnostics-v2.spanish_diagnosticsassay-cupel-pose-error-v2-r8-all-frame-diagnostics-rltcm-diagnostics
Diagnostics · 中医诊法脉学 💰 (Commercial Dataset)
This is a commercial dataset. A free 3-work sample is provided below; the
full dataset is available for licensing/purchase.
📧 To purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — derived from public-domain classical works.
What you get
Pulse/tongue/inspection + pattern diagnosis: 脉经·濒湖脉学·舌鉴 (脉学·辩证诊治)
42 public-domain works of classical Traditional Chinese Medicine, as clean full… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/tcm-diagnostics.eval_servo_diagnosticsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 210,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/adungus/eval_servo_diagnostics.crag-mm-diagnosticsdiagnostic-scope-differential-control-radiology-v01Diagnostic Scope and Differential Control v01
What this dataset is
This dataset evaluates whether a system respects the diagnostic limits of a radiologic study and avoids collapsing the differential diagnosis prematurely.
You give the model:
Imaging findings
A clinical prompt
A report level claim
You ask one question.
Is this conclusion
within the diagnostic scope
of the image
Why this matters
Radiology supports diagnosis.
It rarely delivers certainty.
Common failure patterns:
Treating… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/diagnostic-scope-differential-control-radiology-v01.diagnostics-new-with-categories
Diagnostics Dataset with Categories
This is a diagnostic dataset for testing various linguistic and semantic phenomena.
Dataset Structure
The dataset contains 14 subtasks, each testing different linguistic phenomena.
Usage
from datasets import load_dataset
# Load a specific subtask
dataset = load_dataset("your_username/diagnostics-new-with-categories", "Anaphora")
# Load all configs
all_configs = load_dataset("your_username/diagnostics-new-with-categories"… See the full description on the dataset page: https://huggingface.co/datasets/Junrui1202/diagnostics-new-with-categories.clinical-icu-demand-staff-bed-diagnostics-quad-coherence-risk-v0.1What this repo is for
This dataset tests whether a model can detect quad coupling coherence risk in hospital critical care flow.
It measures alignment between four signals
patient acuity and ICU demand
staffing coverage
available ICU beds
diagnostic turnaround time
You label each case
coherent when the four signals align and escalation completes
incoherent when any node drifts enough to block escalation or trigger system strain
What it predicts
ICU overflow events
ED boarding spikes
delayed… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-icu-demand-staff-bed-diagnostics-quad-coherence-risk-v0.1.healthcare-diagnostics-turnaround-coherence-risk-v0.1What this repo is for
detect diagnostic backlog early
predict treatment delays
flag imaging capacity issues
support throughput planning
reduce length of stay
clinical-icu-demand-staff-bed-diagnostics-quad-coherence-risk-v0.2Clinical ICU Demand Staff Bed Diagnostics Quad Coherence Risk v0.2
What this dataset does
It tests whether a model can detect ICU system coherence failure under load.
Four coupled nodes
patient_acuity_signal
icu_demand_signal
staffing_coverage_signal
available_icu_beds_signal
Timing node
diagnostic_turnaround_signal
Escalation loop fields
escalation_attempted_signal
escalation_completed_signal
delay_reason_documented_signal
interim_mitigation_signal
Task
Given the clinical… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-icu-demand-staff-bed-diagnostics-quad-coherence-risk-v0.2.industrial-machinery-safety-diagnostics-preview
⚙️ Industrial Machinery Safety & Diagnostics Dataset (Enterprise Preview)
Overview
This repository contains a 150-row verified preview sample of our proprietary enterprise dataset designed for industrial RAG applications, diagnostic assistant tuning, and machinery safety compliance models.
The full core dataset is grounded in official European machinery safety standards, CNC diagnostic procedures, and hydraulic equipment maintenance documentation.
💡… See the full description on the dataset page: https://huggingface.co/datasets/Moravax/industrial-machinery-safety-diagnostics-preview.icml26-repro-conditional-coverage-diagnostics
Conditional Coverage Diagnostics reproduction
Independent scaled checks for ERT definitions, classifier power, convergence,
the over/under decomposition, and the k-fold estimator. Run with:
uv run --with-requirements requirements.txt python reproduce.py
The author_code/ checkout is retained for provenance. reproduce.py does not
import it.
Vehicle_Diagnostics_LLM_Training_Sample
Vehicle Diagnostic Sample Dataset
🧩 Dataset Summary
This dataset contains a sample subset of structured vehicle diagnostic logs generated for various vehicle types and subsystems, such as transmissions, battery systems, brakes, and engines. Each entry includes detailed parameters such as fault codes, performance metrics, measurements, temporal trends, and maintenance recommendations.
This subset (500 examples) is meant to demonstrate the structure and potential use cases… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Vehicle_Diagnostics_LLM_Training_Sample.Oncology-Companion-Diagnosticsmicroplex-us-diagnostics
Microplex US Diagnostics
Stable diagnostics registry for Microplex-US artifact bundles.
Layout
runs/<run_id>/manifest.json
runs/<run_id>/policyengine_native_scores.json
runs/<run_id>/pe_us_data_rebuild_native_audit.json
runs/<run_id>/pe_native_target_diagnostics.json
latest.json
run_registry.jsonl
The runs/<run_id>/... paths are immutable once published. latest.json and run_registry.jsonl are mutable discovery files.
nemotron-car-diagnostics-datasets
