datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
midashenglm-gen-training-latents
ModelsLab/midashenglm-gen-training-latents
Precomputed audio latents for fine-tuning
mispeech/midashenglm-gen,
paired with six-view prompts in the exact format the model was trained on.
This is not an audio dataset and not a caption dataset. Each record is the
output of the model's frozen DashengTokenizer encoder — 768-dimensional latents
at 25 Hz, stored float16 — next to the tagged prompt string built from the
source metadata.
Why it exists
The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.yoruba-cfm-latentsLatentSkill
LatentSkill Data
This dataset repository contains the data released for LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents.
Code: https://github.com/yuaofan0-oss/LatentSkillPaper: https://arxiv.org/abs/2606.06087Checkpoint repository: https://huggingface.co/AofaYu71/LatentSkill
Contents
skill_pretrain/
train.jsonl
val.jsonl
skill_ift/
train.json
search_test/
2wikimultihopqa_test.jsonl
bamboogle_test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/AofaYu71/LatentSkill.baar-tinysd-latents-cachesn38r4b-subsn38r6-u208-subsn38-submission-r3dsn38-submission-r3esn38r4-sub
