CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01xX-its-amit-Xx /pxr-structure-pose-pool PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands) Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2) structure-prediction challenge, released openly with per-pose labels so the community can reuse the compute already spent — and, we hope, crack the problem this data makes visible. What's here poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand (resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.tabular1K<n<10K0 likes938 downloads2mo agoHugging Face02jablonkagroup /nomad_structure Dataset Details Dataset Description A subset from NOMAD dataset, which is a database of DFT computed results of materials. This subset consists of cif structures of around 0.5 million bulk stable materials and their geometric and structural information. All materials in this dataset are modeled using Density Functional Theory using GGA functional. Curated by: License: CC BY 4.0 Dataset Sources original data source Citation BibTeX:… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nomad_structure.tabular10M<n<100M0 likes794 downloads1y agoHugging Face03anonymous-structured-agent /structured-file-audit-benchmark Paper Data Release This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them. Contents datasets/ Benchmark data and per-task manifests for the three paper-facing splits. datasets/sc_flat/data SC-Flat is derived from DaBench, augmented with a replayable perturbation injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.texttable-question-answering1 likes658 downloads2mo agoHugging Face04safelegalaidata /eu-ai-act-structured EU AI Act, structured Regulation (EU) 2024/1689 (the Artificial Intelligence Act) as tables: every article, recital, annex and definition, 677 obligations coded by actor, risk tier, application date and penalty basis, plus milestones, national competent authorities and fine tiers. Built 2026-09-08 by SafeLegalAI (Cognesio LLP) from the official English texts served by the Publications Office of the European Union (Cellar): the consolidated text as of 27 July 2026 (CELEX… See the full description on the dataset page: https://huggingface.co/datasets/safelegalaidata/eu-ai-act-structured.tabular1K<n<10K1 likes232 downloads14d agoHugging Face05fredzzp /uniref50-sorted-structure-tokentabular10M<n<100M1 likes196 downloads10mo agoHugging Face06sxiong /MLR_structured_trajectory Reasoning Trajectories with Step-Level Annotations This dataset contains structured reasoning trajectories introduced in (ICLR 2026) Enhancing Language Model Reasoning with Structured Multi-Level Modeling. Compared with the full-trajectory release, this dataset is a cleaned and segmented version designed for research on hierarchical reasoning, trajectory supervision, and multi-step policy training. Each example contains the original prompt and response fields together with a… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/MLR_structured_trajectory.tabular10K<n<100K1 likes169 downloads3mo agoHugging Face07Synthyra /PDB-Monomeric-Structure-ESMFold2 PDB-Monomeric-Structure-ESMFold2 Monomeric, protein-only PDB structure dataset for minimum ESMFold2-style training. Each row is one eligible single-chain biological assembly with a canonical amino-acid sequence input and all-atom protein labels in atom37. Labels atom37_positions: residue x 37 x 3 coordinates, with zeros for missing atoms. atom37_mask: residue x 37 resolved-atom mask. aatype, residue_index, auth_seq_id, insertion_code, residue_name, ca_mask.… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/PDB-Monomeric-Structure-ESMFold2.tabular100K<n<1M0 likes157 downloads3mo agoHugging Face08lamm-mit /protein-secondary-structure-netsurfp NetSurfP-3.0 Secondary-Structure Splits This dataset repo contains NetSurfP-derived protein secondary-structure labels converted for Protein-I-JEPA probe training and evaluation. Source page: https://services.healthtech.dtu.dk/services/NetSurfP-3.0/5-Dataset.php Profile: hhblits Labels are Q3 per-residue labels: H: helix E: beta strand C: coil/other .: ignored residue for loss and accuracy Splits Split Rows JSONL TSV train 10348 train.jsonl tsv/train.tsv… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/protein-secondary-structure-netsurfp.tabular10K<n<100K0 likes137 downloads4mo agoHugging Face09jamesdborin /Nemotron-Agents-Tools-and-Structured-Task-Execution-prompt-only Agents, Tools and Structured Task Execution Prompt-Only This dataset combines prompt-only datasets by capability theme for distillation experiments. It contains 675,882 unique prompts from 808,884 raw rows; 133,002 exact canonical duplicates were removed. Rows retain the canonical prompt-extraction columns and add source_repo_id for provenance. Deduplication uses normalized system_prompt, prompt, tools, and schema_str, with the first row in manifest order retained. Original… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-Agents-Tools-and-Structured-Task-Execution-prompt-only.tabular100K<n<1M0 likes132 downloads2mo agoHugging Face10dhruveshpatel /openclassgen-structured-v1 OpenClassGen Structured v1 Derived from mrahman2025/OpenClassGen (Rahman et al. 2025, arXiv:2504.15564). License: CC BY 2.0 (same as upstream). Keep repository_name and file_path when redistributing. Underlying GitHub repos may carry additional software licenses. gold_code is upstream human_written_code. We add parsed fields, body-span indices, and a Variant-3 prompt/target pair (v3_prompt_text / v3_target_text). No unit tests. Splits are repository-disjoint (train /… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/openclassgen-structured-v1.tabulartext-generation100K<n<1M0 likes130 downloads2mo agoHugging Face11StructureCloud /OMol25tabular1M<n<10M0 likes125 downloads2mo agoHugging Face12StructureCloud /HiQBindtabular10K<n<100K0 likes87 downloads2mo agoHugging Face13narendra747 /structured-vitalsimage10K<n<100K0 likes83 downloads5mo agoHugging Face14StructureCloud /MatPEStabular100K<n<1M0 likes83 downloads2mo agoHugging Face15letrinhan /vn-provinces-land-use-structure Vietnam land use structure by locality Vietnam land-use composition shares (%) of total natural area as of 31 December - agricultural production, forestry, special-use, and residential - for 2018 and 2020-2023. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Comparison Color key Files… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-land-use-structure.tabularn<1K0 likes78 downloads1d agoHugging Face16structure-epflai /nmrexptabular1M<n<10M0 likes74 downloads5mo agoHugging Face17Haeryz /putusan-structured-extraction Putusan structured-extraction dataset Built 2026-07-08T23:02:28+00:00 by notebooks/build_dataset.py (seed 3407). Indonesian court-decision (putusan) extractive-structuring dataset over three corpora (Anak, Asusila, TPPO). Each row is one model extraction of one source document into 31 canonical sections of verbatim spans. Empty sections were completed from sibling model extractions of the same document where available (cross_model_fill_json records per-section donor provenance).… See the full description on the dataset page: https://huggingface.co/datasets/Haeryz/putusan-structured-extraction.tabulartext-generation1K<n<10K0 likes71 downloads2mo agoHugging Face18NagaYu /rebar-structure 🧱 Rebar Structure A corpus for restoring the heading hierarchy of Japanese documents from flat text and for evaluating structure-aware chunking. Each record is a flattened document, its gold heading tree (positions + depths), and a damaged variant simulating PDF/text extraction. Code: https://github.com/NagaYu/rebar Model: https://huggingface.co/NagaYu/rebar-heading-classifier Demo (Space): https://huggingface.co/spaces/NagaYu/rebar Why it exists The same… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/rebar-structure.tabulartoken-classificationn<1K0 likes59 downloads10d agoHugging Face19seeaman /economic-index-structured Economic Index - Structured & Cleaned Dataset This dataset is a cleaned, structured version of the Anthropic Economic Index, organized for easy integration with persona-based scenario generation pipelines. Dataset Description The Anthropic Economic Index tracks how people use Claude AI for work-related tasks. This structured version extracts and organizes the key information into easy-to-use tables. Original Data Period: August 4-11, 2025Source: Anthropic Economic… See the full description on the dataset page: https://huggingface.co/datasets/seeaman/economic-index-structured.tabular1K<n<10K0 likes53 downloads7mo agoHugging Face20structure-epflai /MassSpecGym MassSpecGym provides a dataset and benchmark for the discovery and identification of new molecules from MS/MS spectra. The provided challenges abstract the process of scientific discovery of new molecules from biological and environmental samples into well-defined machine learning problems. Please refer to the MassSpecGym GitHub page and the paper for details. tabular100K<n<1M0 likes46 downloads5mo agoHugging Face21StructureCloud /AlexMP20tabular100K<n<1M0 likes45 downloads2mo agoHugging Face22AmelieSchreiber /toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-001 ToricBLM dataset state: toricblm-structure-priority-balanced-3day-20260709T185034Z epoch 001 This dataset repo records the exact local training-data state visible to the dynamic epoch launcher. It intentionally stores manifests and audit records rather than duplicating large Parquet shards. Special checkpoint: toricblm-structure-priority-balanced-3day-20260709T185034Z_epoch_001_special_structure_current_step_002000.pt Checkpoint repo: AmelieSchreiber/ToricGT_160M_FoT Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-001.tabulartext-generationn<1K0 likes45 downloads2mo agoHugging Face23tomyimkc /repro-on-structured-state-space-duality-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes44 downloads2mo agoHugging Face24quguanni /scp-foundation-structuredtabular10K<n<100K1 likes43 downloads7mo agoHugging Face25AmelieSchreiber /toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-002 ToricBLM dataset state: toricblm-structure-priority-balanced-3day-20260709T185034Z epoch 002 This dataset repo records the exact local training-data state visible to the dynamic epoch launcher. It intentionally stores manifests and audit records rather than duplicating large Parquet shards. Special checkpoint: toricblm-structure-priority-balanced-3day-20260709T185034Z_epoch_002_special_structure_delta_step_002750.pt Checkpoint repo: AmelieSchreiber/ToricGT_160M_FoT Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-002.tabulartext-generationn<1K0 likes43 downloads2mo agoHugging Face26barbara-ai /administrative-structure-912fd9 administrative-structure-912fd9 Synthetic products test data: 36 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number… See the full description on the dataset page: https://huggingface.co/datasets/barbara-ai/administrative-structure-912fd9.tabularn<1K0 likes42 downloads11d agoHugging Face27genbio-ai /sample-structure-dataset Sample dataset for PETAL model This dataset is a sample dataset to test the functionalities of the PETAL model (encoder and decoder). It is based on CASP15 dataset, see https://predictioncenter.org/casp15/ https://github.com/Bhattacharya-Lab/CASP15 The registries folder contains the registry of CASP15 dataset (a csv file with filename, pdb_id, etc.) tabularn<1K0 likes41 downloads2y agoHugging Face28dargason /structure184-five-method-cofolding Structure184 five-method PXR cofolding dataset This repository contains 92,000 PXR–ligand cofolded models: 184 ligands 20 seed positions 5 samples per seed Boltz2, Chai-1, ESMFold2, OpenFold3, and Protenix 18,400 models per method Models use the governed prepared top solution-state ligand representation. Coordinates are harmonized to the 1NRL PXR frame using a 162-Cα core. The original generated coordinates are preserved up to that rigid alignment. Repository… See the full description on the dataset page: https://huggingface.co/datasets/dargason/structure184-five-method-cofolding.tabular100K<n<1M0 likes39 downloads2mo agoHugging Face29leonli66 /stage3-synthetic-structured-retrieval Stage 3 Synthetic Structured-Retrieval Agents Native search-tool trajectories generated by Qwen/Qwen3-235B-A22B-Instruct-2507 for LCLM Stage-3 agent post-training. The default config contains only traces that passed programmatic evidence and answer verification. Harvest Accepted traces: 82 Native search calls: 179 Compressed tool-observation traces: 41 Uncompressed traces: 41 Task-ID overlap between pilot and collection batch: 0 Family counts are 20 latest-state… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-synthetic-structured-retrieval.tabulartext-generationn<1K0 likes39 downloads1mo agoHugging Face30jihye54 /medical-structure-f19c39 medical-structure-f19c39 Synthetic weather test data: 39 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/jihye54/medical-structure-f19c39.tabularn<1K0 likes38 downloads11d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.