datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synth-bootstrap-trialmodel-cards-ml-metadata-bootstrap
davanstrien/model-cards-ml-metadata-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over librarian-bots/model_cards_with_metadata.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
librarian-bots/model_cards_with_metadata (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
base model name, context length, training method, training dataset name, benchmark name… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/model-cards-ml-metadata-bootstrap.AcitvityNet-Captions-bootstrapped-5Kgrpo-5-sft-bootstraptraining-methods-bootstrap
davanstrien/training-methods-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over /input/cleaned-cards.parquet.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
/input/cleaned-cards.parquet (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
training method
Confidence threshold
0.7
Samples processed
10000
Total entities extracted
4278
Inference device
cuda… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/training-methods-bootstrap.eval-mentions-bootstrap
davanstrien/eval-mentions-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over /input/cleaned-cards.parquet.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
/input/cleaned-cards.parquet (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
benchmark name, evaluation dataset, evaluation metric
Confidence threshold
0.6
Samples processed
10000
Total entities… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/eval-mentions-bootstrap.bootstrap_oai_ptbootstrap_oai_pt_thinkgrpo-5-sft-bootstrap-qwen3-4b-thinking-2507
BRlkl/grpo-5-sft-bootstrap-thinking
Derived from BRlkl/grpo-5-sft-bootstrap-qwen3-4b-thinking-2507.
This version repairs the plan column only for rows where blacklisted = true.
Transformation:
Parse the plan JSON.
Read walk[0].
If walk[0] is the wrapped prompt form:
CONVERSATION_HISTORY: [Empty] ... Generate {"walk":[...]} for NEW_USER_MESSAGE.
then replace it with just the embedded NEW_USER_MESSAGE text.
Leave all non-blacklisted rows unchanged.
Audit summary:
Total rows:… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/grpo-5-sft-bootstrap-qwen3-4b-thinking-2507.AcitvityNet-Captions-bootstrapped-smol
