datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bootstrap-latent-thought-dataThis dataset is associated with the paper Reasoning to Learn from Latent Thoughts. It contains data used for pretraining language models with a focus on improving data efficiency by modeling and inferring latent thoughts underlying the text generation process, such as on reasoning-intensive math corpus. An expectation-maximization algorithm is developed for models to self-improve their self-generated thoughts and data efficiency.
synth-bootstrap-trialbootstrap_sms_v2_repeat_1bootstrap_sms_v2bootstrap_sms
Dataset Card for "bootstrap_sms"
More Information needed
model-cards-ml-metadata-bootstrap
davanstrien/model-cards-ml-metadata-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over librarian-bots/model_cards_with_metadata.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
librarian-bots/model_cards_with_metadata (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
base model name, context length, training method, training dataset name, benchmark name… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/model-cards-ml-metadata-bootstrap.curated-v2-combined-dpo-bootstrapAcitvityNet-Captions-bootstrapped-5Kgrpo-5-sft-bootstrapbootstrap_agreement_long_17training-methods-bootstrap
davanstrien/training-methods-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over /input/cleaned-cards.parquet.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
/input/cleaned-cards.parquet (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
training method
Confidence threshold
0.7
Samples processed
10000
Total entities extracted
4278
Inference device
cuda… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/training-methods-bootstrap.bootstrap_oai_ptbootstrap_leslie_long_80grpo-5-sft-bootstrap-qwen3-4b-thinking-2507
BRlkl/grpo-5-sft-bootstrap-thinking
Derived from BRlkl/grpo-5-sft-bootstrap-qwen3-4b-thinking-2507.
This version repairs the plan column only for rows where blacklisted = true.
Transformation:
Parse the plan JSON.
Read walk[0].
If walk[0] is the wrapped prompt form:
CONVERSATION_HISTORY: [Empty] ... Generate {"walk":[...]} for NEW_USER_MESSAGE.
then replace it with just the embedded NEW_USER_MESSAGE text.
Leave all non-blacklisted rows unchanged.
Audit summary:
Total rows:… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/grpo-5-sft-bootstrap-qwen3-4b-thinking-2507.bootstrap_oai_pt_thinkbootstrap_leslie_long_94bootstrap_leslie_long_98bootstrap_prevalence_long_6bootstrap_prevalence_long_17bootstrap_agreement_long_43sam3-ls-bootstrap-demo
davanstrien/sam3-ls-bootstrap-demo
Bootstrap dataset produced by running facebook/sam3 over a small set of test images and storing the predictions in a Label Studio project for review.
This is a proof-of-concept artifact demonstrating an end-to-end "unlabeled images → bootstrapped dataset" workflow on Hugging Face infrastructure. The predictions in this dataset are SAM3 outputs — not human-reviewed.
Workflow
Images imported into Label Studio project 20 on… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/sam3-ls-bootstrap-demo.eval-mentions-bootstrap
davanstrien/eval-mentions-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over /input/cleaned-cards.parquet.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
/input/cleaned-cards.parquet (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
benchmark name, evaluation dataset, evaluation metric
Confidence threshold
0.6
Samples processed
10000
Total entities… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/eval-mentions-bootstrap.AcitvityNet-Captions-bootstrapped-smolbootstrap_leslie_long_3bootstrap_leslie_long_67bootstrap_prevalence_long_8bootstrap_prevalence_long_10bootstrap_prevalence_long_93bootstrap_agreement_long_20bootstrap_agreement_long_22
