CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01wikimedia /structured-wikipedia Dataset Card for Wikimedia Structured Wikipedia Quick Links Wikimedia Enterprise Structured Contents Documentation Data Dictionary Wikimedia Attribution Framework Meta-Wiki Discussion Dataset Summary Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API. This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/structured-wikipedia.text10M<n<100M394 likes14k downloads4mo agoHugging Face02theodi /ndl-core-structured-data NDL Core – Structured Data Overview NDL Core – Structured Data is a curated collection of structured UK public sector datasets, converted into Apache Parquet format for efficient analytics and machine learning workflows. This repository is part of the broader NDL Core Corpus, which combines both textual and structured data sourced from authoritative UK government and public sector platforms. Textual sources (e.g. GOV.UK, Hansard, legislation.gov.uk) are hosted separately… See the full description on the dataset page: https://huggingface.co/datasets/theodi/ndl-core-structured-data.100M<n<1B0 likes11k downloads8mo agoHugging Face03radiata-ai /brain-structureA collection of T1-weighted .nii.gz structural MRI scans in a BIDS-like arrangement, with JSON sidecar metadata indicating train/validation/test splits.image-classification9 likes4.7k downloads2y agoHugging Face04isalgo /tcren_structures isalgo/tcren_structures TCR:peptide:MHC structure sets and benchmarks for TCRen2 (structure-based prediction of TCR recognition). Fetch with tcren / the manuscript scripts/bootstrap_data.py. Contents rule: structures as .gz/.tar.gz (LFS) and .txt/.md descriptions only — no notebooks, figures, or analysis tables. Layout folder task contents Native2026/, Canonical2026/ derivation / ergodicity non-redundant TCR:pMHC structures (.gz) Native2022/… See the full description on the dataset page: https://huggingface.co/datasets/isalgo/tcren_structures.0 likes1.4k downloads1mo agoHugging Face05Arun63 /sharegpt-structured-output-json ShareGPT-Formatted Dataset for Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.texttext-generationn<1K7 likes1.4k downloads2y agoHugging Face06agarosegirls /viral-protein-structuresFolded/Extracted structures from PDB, AF2, ESMAtlas, and additional structures folded via AlphaFold2 on the Kempner Institute H100 GPUs. 1 likes1.3k downloads6mo agoHugging Face07nvidia /Nemotron-RL-Instruction-Following-Structured-Outputs-v2 Dataset Description: Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema. Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.texttext-generation10K<n<100K7 likes1.2k downloads4mo agoHugging Face08open-athena /a3-rl-laion_nemotron-gym-instruction-following-structuredtext10K<n<100K0 likes1.1k downloads4mo agoHugging Face09xX-its-amit-Xx /pxr-structure-pose-pool PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands) Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2) structure-prediction challenge, released openly with per-pose labels so the community can reuse the compute already spent — and, we hope, crack the problem this data makes visible. What's here poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand (resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.tabular1K<n<10K0 likes938 downloads2mo agoHugging Face10jablonkagroup /nomad_structure Dataset Details Dataset Description A subset from NOMAD dataset, which is a database of DFT computed results of materials. This subset consists of cif structures of around 0.5 million bulk stable materials and their geometric and structural information. All materials in this dataset are modeled using Density Functional Theory using GGA functional. Curated by: License: CC BY 4.0 Dataset Sources original data source Citation BibTeX:… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nomad_structure.tabular10M<n<100M0 likes794 downloads1y agoHugging Face11anonymous-structured-agent /structured-file-audit-benchmark Paper Data Release This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them. Contents datasets/ Benchmark data and per-task manifests for the three paper-facing splits. datasets/sc_flat/data SC-Flat is derived from DaBench, augmented with a replayable perturbation injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.texttable-question-answering1 likes658 downloads2mo agoHugging Face12structure-epflai /neurips-spectraThe dataset from Albert's et al, downloaded from zenodo. It's on here for easier access and organisation. text100K<n<1M0 likes651 downloads6mo agoHugging Face13StructureCloud /MISATO_MDtext10K<n<100K0 likes608 downloads2mo agoHugging Face14nvidia /Nemotron-RL-instruction_following-structured_outputs Dataset Description: The Nemotron-RL-instruction_following-structured_outputs dataset tests the ability of the model to follow output formatting instructions under schema constraints under the JSON format. Each problem consists of three components: The document, output formatting Instruction (Schema), and question. The dataset varies the difficulty of each problem by varying the location of instructions, the comprehensiveness of instructions, the complexity of the schema, and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-instruction_following-structured_outputs.text1K<n<10K40 likes563 downloads8mo agoHugging Face15proteinea /secondary_structure_predictiontext10K<n<100K4 likes547 downloads4y agoHugging Face16helioom /3dvlm-structured3d_subset Structured3D Subset (3DVLM) A small, fast-to-download slice of the Structured3D synthetic indoor dataset, converted to a uniform posed-RGB-D format for quick model test-runs. This is a subset: 100 scenes (randomly sampled, seed 0) from collection 00, using the pre-rendered full (furnished) perspective views. Across the 100 scenes there are 2,198 frames (3–49 per scene). These are photorealistic synthetic renders with perfect dense ground-truth depth and exact camera poses — no… See the full description on the dataset page: https://huggingface.co/datasets/helioom/3dvlm-structured3d_subset.3ddepth-estimation1K<n<10K0 likes515 downloads3mo agoHugging Face17jiahaozhang2003 /beacon-secondary-structure BEACON — Secondary_structure_prediction RNA secondary-structure prediction data with nucleotide-level pair matrices. Official data from the shared BEACON/RNABenchmark Drive folder: https://drive.google.com/drive/folders/19ddrwI8ycvIxkgSV3gDo_VunLofYd4-6?hl=en. This repository is the standardized Hugging Face publication of the official task data. The data/ directory is the canonical viewer-friendly layer, and the original file contents and source names are preserved for… See the full description on the dataset page: https://huggingface.co/datasets/jiahaozhang2003/beacon-secondary-structure.textother10K<n<100K0 likes482 downloads2mo agoHugging Face18StructureCloud /OMat240 likes471 downloads2mo agoHugging Face19DiffSynth-Studio /ImagePulseV2-Edit-Structure ImagePulseV2 Dataset - Image Structure The ImagePulseV2 dataset is a collection we constructed for training the Diffusion Templates series of models. It comprises multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB. Open-source code: DiffSynth-Studio Technical report: arXiv Project homepage: GitHub Documentation: English Version, Chinese Version Online demo: ModelScope Studio Model… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Structure.image100K<n<1M0 likes369 downloads5mo agoHugging Face20terminusresearch /pseudo-camera-10k-structured-json pseudo-camera-10k, structured JSON captions The 9,997 training images from bghira/pseudo-camera-10k, recaptioned into the structured JSON caption schema that Ideogram 4 consumes. The images are unchanged: free photographs from world class photographers, Lanczos-resized so the shorter edge is 1024px, nothing upsampled. The original dataset carries short CogVLM prose captions. This one replaces them with one JSON object per image describing the scene at three levels: an overall… See the full description on the dataset page: https://huggingface.co/datasets/terminusresearch/pseudo-camera-10k-structured-json.imagetext-to-image10K<n<100K0 likes332 downloads12d agoHugging Face21Antix5 /structure-heavy-token-quality-datasettext1K<n<10K0 likes315 downloads28d agoHugging Face22jungtaekkim /datasets-nanophotonic-structures0 likes314 downloads3y agoHugging Face23open-athena /nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes302 downloads2mo agoHugging Face24Aregay01 /structured-wikipedia Dataset Card for Wikimedia Structured Wikipedia Quick Links Wikimedia Enterprise Structured Contents Documentation Data Dictionary Wikimedia Attribution Framework Meta-Wiki Discussion Dataset Summary Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API. This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/Aregay01/structured-wikipedia.text10M<n<100M0 likes298 downloads4mo agoHugging Face25domofon /structured-cpt Structured CPT - JSON + SQL pretrain documents SmolLM2-1.7B continued-pretraining shard of structured documents. Each document is a <task> / <input> / <output> block whose <output> is a canonical JSON object, terminated by the SmolLM2 end-of-text token ``. Sources: source description rows shards repeat sql_bmc2 b-mc2 sql-create-context -> JSON (4 keys, stub explanation) 392,885 1 5 sql_gretelai gretelai synthetic_text_to_sql -> JSON (4 keys) 529,255 1 5… See the full description on the dataset page: https://huggingface.co/datasets/domofon/structured-cpt.texttext-generation1M<n<10M0 likes298 downloads16d agoHugging Face26PDBEurope /protein_structure_NER_model_v3.1 Overview This data was used to train model: https://huggingface.co/PDBEurope/BiomedNLP-PubMedBERT-ProteinStructure-NER-v3.1 There are 20 different entity types in this dataset: "bond_interaction", "chemical", "complex_assembly", "evidence", "experimental_method", "gene", "mutant", "oligomeric_state", "protein", "protein_state", "protein_type", "ptm", "residue_name", "residue_name_number","residue_number", "residue_range", "site", "species", "structure_element", "taxonomy_domain"… See the full description on the dataset page: https://huggingface.co/datasets/PDBEurope/protein_structure_NER_model_v3.1.0 likes290 downloads2y agoHugging Face27Panhapich /bank-statement-structure-recognition Synthetic Bank Statement Table Structure Dataset A synthetically generated collection of bank statement images with pixel-perfect, automatically-produced bounding box annotations for table structure recognition (TSR). 🔑 In one sentence: fake bank statements + auto-generated YOLO labels for every table cell, built so you can train table-detection models (TATR, DETR, YOLO) without manual annotation. At a Glance Task Object Detection → Table… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/bank-statement-structure-recognition.imageobject-detection10K<n<100K1 likes289 downloads3mo agoHugging Face28PDBEurope /protein_structure_NER_model_v2.1 Overview This data was used to train model: https://huggingface.co/PDBEurope/BiomedNLP-PubMedBERT-ProteinStructure-NER-v2.1 There are 20 different entity types in this dataset: "bond_interaction", "chemical", "complex_assembly", "evidence", "experimental_method", "gene", "mutant", "oligomeric_state", "protein", "protein_state", "protein_type", "ptm", "residue_name", "residue_name_number","residue_number", "residue_range", "site", "species", "structure_element", "taxonomy_domain"… See the full description on the dataset page: https://huggingface.co/datasets/PDBEurope/protein_structure_NER_model_v2.1.0 likes283 downloads2y agoHugging Face29open-athena /nemotron-gym-instruction-following-structured-minimax-m27-131k-tracestext1K<n<10K0 likes245 downloads4mo agoHugging Face30GoktugD /turkish-structured-summarization-1.5m Turkish Structured Summarization 1.5M v2 Üç cümlelik kurgusal operasyon kayıtları ve kısa Türkçe özetleri. Doğrulanmış boyut Train: 1,470,000 Validation: 15,000 Test: 15,000 Toplam: 1,500,000 Ana görev sütunları: id, document, summary, domain Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-structured-summarization-1.5m.textsummarization1M<n<10M0 likes236 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.