datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NextNightglm5-next-tiny-fidelity-root-v1
glm5_next random CPU fixture root
A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/glm5-next-tiny-random-bf16.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm5-next-tiny-fidelity-root-v1.MVEmo
MVEmo
This is the dataset repository for the paper: Bridging Categorical and Dimensional Affect: The MVEmo Multi-Task Benchmark for Music-Related Emotion Recognition.
Dataset Details
Dataset Description
MVEmo is a large-scale multimodal dataset that consists of 11,764 music video samples with both static and dynamic emotion annotations for music-related emotion recognition (MRER). It consists of the following key features:
Basic Information: title, artist… See the full description on the dataset page: https://huggingface.co/datasets/NEXTLab-ZJU/MVEmo.nexttoken-pmkisan-domain-sft-data
NextToken pmkisan domain SFT data (v1)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA,
banking products, insurance, savings instruments, etc.).
Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.wish-engine-toolcall-next-v3-strict-general
wish-engine-toolcall-next-v3-strict-general
Wish-engine implementor next-step tool-calling dataset (v3 strict generalization subset, dynamic aliases).
Splits
train.jsonl: 8081980 bytes
validation.jsonl: 1008543 bytes
test.jsonl: 997791 bytes
Schema
Rows are JSONL with at least:
id
messages (chat format with assistant tool_calls)
tool_name
metadata fields (mode, status, trajectory_*)
Notes
Tool names are dynamically aliased per sample.
A tool… See the full description on the dataset page: https://huggingface.co/datasets/sahilmob/wish-engine-toolcall-next-v3-strict-general.wish-engine-toolcall-next-v2
wish-engine-toolcall-next-v2
Wish-engine implementor next-step tool-calling dataset (v2 canonicalized format).
Splits
train.jsonl: 12598457 bytes
validation.jsonl: 1731252 bytes
test.jsonl: 1558954 bytes
Schema
Rows are JSONL with at least:
id
messages (chat format with assistant tool_calls)
tool_name
metadata fields (mode, status, trajectory_*)
Generated from wish-engine benchmark artifacts by:
scripts/build-hf-toolcall-datasets-v2.mjs
wish-engine-toolcall-next-v2-strict
wish-engine-toolcall-next-v2-strict
Wish-engine implementor next-step tool-calling dataset (v2 strict canonical subset).
Splits
train.jsonl: 5447561 bytes
validation.jsonl: 679818 bytes
test.jsonl: 672961 bytes
Schema
Rows are JSONL with at least:
id
messages (chat format with assistant tool_calls)
tool_name
metadata fields (mode, status, trajectory_*)
Generated from wish-engine benchmark artifacts by:
scripts/build-hf-toolcall-datasets-v2.mjs
wish-engine-toolcall-next-v1-strict
wish-engine-toolcall-next-v1-strict
Wish-engine implementor tool-call next-step dataset (strict write-positive subset, v1).
Splits
train.jsonl: 5179089 bytes
validation.jsonl: 557088 bytes
test.jsonl: 772103 bytes
Schema
Rows are JSONL with at least:
id
messages (chat format with assistant tool_calls where applicable)
metadata fields (mode, status, tool_name, etc.)
Generated from wish-engine benchmark artifacts by:
scripts/build-hf-toolcall-datasets-v1.mjs
wish-engine-toolcall-next-v1
wish-engine-toolcall-next-v1
Wish-engine implementor tool-call next-step supervision dataset (v1).
Splits
train.jsonl: 12148281 bytes
validation.jsonl: 1366467 bytes
test.jsonl: 1658634 bytes
Schema
Rows are JSONL with at least:
id
messages (chat format with assistant tool_calls where applicable)
metadata fields (mode, status, tool_name, etc.)
Generated from wish-engine benchmark artifacts by:
scripts/build-hf-toolcall-datasets-v1.mjs
nextjs1000wish-engine-toolcall-next-v3-general
wish-engine-toolcall-next-v3-general
Wish-engine implementor next-step tool-calling dataset (v3 generalization, dynamic tool aliases).
Splits
train.jsonl: 19689428 bytes
validation.jsonl: 2711092 bytes
test.jsonl: 2444482 bytes
Schema
Rows are JSONL with at least:
id
messages (chat format with assistant tool_calls)
tool_name
metadata fields (mode, status, trajectory_*)
Notes
Tool names are dynamically aliased per sample.
A tool roster is… See the full description on the dataset page: https://huggingface.co/datasets/sahilmob/wish-engine-toolcall-next-v3-general.heart-arrhythmias-symptoms
