datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bootstrap-latent-thought-dataThis dataset is associated with the paper Reasoning to Learn from Latent Thoughts. It contains data used for pretraining language models with a focus on improving data efficiency by modeling and inferring latent thoughts underlying the text generation process, such as on reasoning-intensive math corpus. An expectation-maximization algorithm is developed for models to self-improve their self-generated thoughts and data efficiency.
pod-bootstrapgraphical-bootstrap-correlator-dataset
Graphical Bootstrap Correlator Dataset
This dataset contains large-scale graph-structured data arising from high-order perturbative computations of four-point correlators in planar $\mathcal{N}=4$ super Yang--Mills theory.
The data consists of denominator graphs (d-graphs) appearing in the graphical bootstrap formulation of correlators. Each graph is associated with a binary label indicating whether it contributes to the correlator at a given perturbative order.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Gabriele-dian/graphical-bootstrap-correlator-dataset.graphical-bootstrap-correlator-dataset
Graphical Bootstrap Correlator Dataset
This dataset contains large-scale graph-structured data arising from high-order perturbative computations of four-point correlators in planar $\mathcal{N}=4$ super Yang--Mills theory.
The data consists of denominator graphs (d-graphs) appearing in the graphical bootstrap formulation of correlators. Each graph is associated with a binary label indicating whether it contributes to the correlator at a given perturbative order.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/anonymous314/graphical-bootstrap-correlator-dataset.synth-bootstrap-trialguile-bootstrapmaskscore-rung-1-bootstrap
MaskScore Rung 1 — Bootstrap (5 of 8 stubs)
Walking-skeleton implementation of MaskScore Rung 1. Five of the eight MASKSCORE.md
stubs are filled with real content from a synthetic ANNY bootstrap (rest pose + rank1
identity + rank5 perturbation). Text, Speech, and Video stubs are deferred to Rung 2 —
the bootstrap has no transcript, no audio, and no video, and CLAUDE.md's ETNF rule
forbids putting a null in for the missing input.
Each stub ships as three ZSTD-compressed parquets:… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/maskscore-rung-1-bootstrap.De_Novo_Drug_Design.BootstrapDetails of AF3D_pLDDT_PChem3D_Shapes_Energy_Binding_2025_03_14 data:
De novo drug design:
Energy minimization is used to generate new drug molecules from scratch based on the 3D structure of the target receptor
Energy minimization, also known as geometry optimization.
9329 Rows of data provided out of total of 14109 rows,
Columns filtered to remove null values.
Duplicate "PDB_protein_key" mapped to same duplicate "AF 3D AtomicData"(along with all data in rows) due to the many to many… See the full description on the dataset page: https://huggingface.co/datasets/AICanada/De_Novo_Drug_Design.Bootstrap.aleafiate-bootstrapbootstrapvue-standardSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: BootstrapVue 2.23
Documentation Data Source Link: https://bootstrap-vue.org/docs/
Data Source License: https://github.com/bootstrap-vue/bootstrap-vue/blob/dev/LICENSE
Data Source Authors: BootstrapVue Team
AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai
drivevla-w0-ar-bootstrap
DriveVLA-W0 NAVSIM AR Action Expert bootstrap
This directory is a server-side bootstrap bundle, not a model or dataset
mirror. It contains reproducible instructions, pinned-source metadata, server
download/verification code, training profiles, and local static-check reports.
Local-stage boundary
The local stage was run through WSL2 Ubuntu-22.04. The Linux source checkout,
revision lock, source archive, source audit, and CPU-only static smoke test have
completed.… See the full description on the dataset page: https://huggingface.co/datasets/wannac1/drivevla-w0-ar-bootstrap.bootstrap_sms_v2_repeat_1hgt-bootstrap-v2-synthetic
HGT Bootstrap V2 Synthetic Pairs
259,896 synthetic mutation pairs (129,948 train + 129,948 eval) generated via ICI-DC using the bootstrap S2 checkpoint.
Generation Details
Model A: Nhoodie/omni-dna-sad-mutation-bootstrap (S2 bootstrap checkpoint)
Model B: LongSafari/hyenadna-tiny-1k-seqlen-hf
Source sequences: 4,998 unique sequences from 8 taxonomic domains
Gap intervals: [3, 4, 5, 6] (codon distance between gaps)
Seeds: Train=[42, 137, 7, 23, 31, 89, 53], Eval=[101, 157… See the full description on the dataset page: https://huggingface.co/datasets/Nhoodie/hgt-bootstrap-v2-synthetic.hexo-bootstrap-corpus
Hexo Human Corpus
Encoding-free corpus of decisive human Hex Tac Toe games — hexagonal grid,
six-in-a-row to win (player 1 opens with 1 move, then both players play 2 moves
per turn; the board is theoretically infinite).
Each line is one game as a raw axial move list + outcome. Nothing about any
neural-network encoding is baked in — no planes, no fixed board size, no action
space. Read it with the stdlib json module and build whatever representation
you want.
Files… See the full description on the dataset page: https://huggingface.co/datasets/timmyburn/hexo-bootstrap-corpus.jade-propainter-bootstrapbootstrap_sms_v2ppl-synthesis-sft-bootstrap
SynthStats PPL Synthesis SFT Bootstrap
This dataset contains natural-language modelling prompts paired with
probabilistic programs, written in the probabilistic programming
languages PyMC (Python) and LazyPPL (Haskell), for supervised
fine-tuning (SFT).
Each row has these fields:
prompt: natural-language modelling task.
reasoning_trace: modelling rationale for the program.
completion: one fenced program block.
complexity: coarse task complexity label.
metadata: runtime… See the full description on the dataset page: https://huggingface.co/datasets/SynthStats/ppl-synthesis-sft-bootstrap.hgt-bootstrap-v1-synthetic
HGT Bootstrap V1 Synthetic Pairs
8,112 synthetic mutation pairs generated via ICI-DC (Interleaved Codon Insertion — Double Consensus) using the SAD coeff1.5 checkpoint as Model A and HyenaDNA-tiny-1k as Model B.
Generation Details
Model A: Nhoodie/omni-dna-sad-mutation (SAD coeff1.5 checkpoint)
Model B: LongSafari/hyenadna-tiny-1k-seqlen-hf (Legacy DC)
Source sequences: 1,014 unique sequences from 8 taxonomic domains
Gap intervals: [3, 4, 5, 6] (codon distance between… See the full description on the dataset page: https://huggingface.co/datasets/Nhoodie/hgt-bootstrap-v1-synthetic.Llama2-7B-generic-predictions-starwars-bootstrap-coef-10-once-augmentedbootstrap_sms
Dataset Card for "bootstrap_sms"
More Information needed
eval-mentions-bootstrap-v2
davanstrien/eval-mentions-bootstrap-v2
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over /input/cleaned-cards-quality.parquet.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
/input/cleaned-cards-quality.parquet (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
benchmark name, evaluation metric
Confidence threshold
0.6
Samples processed
5000
Total entities… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/eval-mentions-bootstrap-v2.Llama2-7B-generic-predictions-starwars-bootstrap-coef-10-oncemodel-cards-ml-metadata-bootstrap
davanstrien/model-cards-ml-metadata-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over librarian-bots/model_cards_with_metadata.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
librarian-bots/model_cards_with_metadata (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
base model name, context length, training method, training dataset name, benchmark name… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/model-cards-ml-metadata-bootstrap.tarsier-bootstrappercurated-v2-combined-dpo-bootstrapAcitvityNet-Captions-bootstrapped-5KMeta-Llama-3-8B-generic-predictions-starwars-bootstrap-coef-10-onceeval-mentions-bootstrap
davanstrien/eval-mentions-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over /input/cleaned-cards.parquet.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
/input/cleaned-cards.parquet (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
benchmark name, evaluation dataset, evaluation metric
Confidence threshold
0.6
Samples processed
10000
Total entities… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/eval-mentions-bootstrap.hexo-bootstrap-corpus
Hexo Human Corpus
Encoding-free corpus of 6,902 decisive human Hex Tac Toe games — hexagonal
grid, six-in-a-row to win (player 1 opens with 1 move, then both players play 2
moves per turn; the board is theoretically infinite).
Each line is one game as a raw axial move list + outcome. Nothing about any
neural-network encoding is baked in — no planes, no fixed board size, no action
space. Read it with the stdlib json module and build whatever representation
you want.… See the full description on the dataset page: https://huggingface.co/datasets/Yulolam/hexo-bootstrap-corpus.runpod-minimax-h3-bootstrap
