structure
table-transformer-structure-recognitiontable-transformer-structure-recognition-v1.1-allMoLFormer-XL-both-10pctOsmosis-Structure-0.6Bppt-pythia-1b-structured-seed3407-stage2ppt-pythia-1b-structured-seed3408-stage2table-transformer-structure-recognition-v1.1-finppt-pythia-1b-appendix-structured-seed3407-stage2
structured-wikipedia
Dataset Card for Wikimedia Structured Wikipedia
Quick Links
Wikimedia Enterprise
Structured Contents Documentation
Data Dictionary
Wikimedia Attribution Framework
Meta-Wiki Discussion
Dataset Summary
Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API.
This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/structured-wikipedia.ndl-core-structured-data
NDL Core – Structured Data
Overview
NDL Core – Structured Data is a curated collection of structured UK public sector datasets, converted into Apache Parquet format for efficient analytics and machine learning workflows.
This repository is part of the broader NDL Core Corpus, which combines both textual and structured data sourced from authoritative UK government and public sector platforms.
Textual sources (e.g. GOV.UK, Hansard, legislation.gov.uk) are hosted separately… See the full description on the dataset page: https://huggingface.co/datasets/theodi/ndl-core-structured-data.brain-structureA collection of T1-weighted .nii.gz structural MRI scans in a BIDS-like arrangement,
with JSON sidecar metadata indicating train/validation/test splits.tcren_structures
isalgo/tcren_structures
TCR:peptide:MHC structure sets and benchmarks for TCRen2 (structure-based prediction of TCR
recognition). Fetch with tcren / the manuscript scripts/bootstrap_data.py.
Contents rule: structures as .gz/.tar.gz (LFS) and .txt/.md descriptions only —
no notebooks, figures, or analysis tables.
Layout
folder
task
contents
Native2026/, Canonical2026/
derivation / ergodicity
non-redundant TCR:pMHC structures (.gz)
Native2022/… See the full description on the dataset page: https://huggingface.co/datasets/isalgo/tcren_structures.sharegpt-structured-output-json
ShareGPT-Formatted Dataset for Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.viral-protein-structuresFolded/Extracted structures from PDB, AF2, ESMAtlas, and additional structures folded via AlphaFold2 on the Kempner Institute H100 GPUs.

