datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Inkling-Small-NVFP4-Regenerated-Collection
Inkling-Small-NVFP4 Regenerated Collection
On-policy training data for a DSpark
speculative-decoding drafter targeting
thinkingmachines/Inkling-Small-NVFP4.
Every assistant response here was regenerated by Inkling-Small-NVFP4 itself over prompts
drawn from Magpie + UltraChat, so the completions reflect the target model's own
distribution rather than the datasets' original responses. This is what makes the data
on-policy for drafter training: the drafter learns to predict the… See the full description on the dataset page: https://huggingface.co/datasets/orestis-z/Inkling-Small-NVFP4-Regenerated-Collection.calib-agentic-sample-Qwen3.8-27B-QUASAR-NVFP4
calib-agentic-sample — Qwen3.8-27B-QUASAR-NVFP4
The exact calibration sample used to quantize lm_head in
digi-texx/Qwen3.8-27B-FULL-NVFP4.
Published so the quantization is reproducible: this is not a representative extract, it is the
documents the quantizer actually saw.
Lineage
11 public agentic / tool-calling datasets
-> digi-texx/calib-agentic-normalized 6,044,537 rows (unified schema)
-> digi-texx/calib-agentic-curated 5,835,723 rows… See the full description on the dataset page: https://huggingface.co/datasets/digi-texx/calib-agentic-sample-Qwen3.8-27B-QUASAR-NVFP4.
