datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
inference-scratchscratch-archiveasr_rnnt_eou_from_scratch
NeMo ASR-EOU 训练脚本解读与论文出处梳理
目标文件:examples/asr/asr_eou/speech_to_text_rnnt_eou_train.py链接:https://github.com/NVIDIA-NeMo/NeMo/blob/main/examples/asr/asr_eou/speech_to_text_rnnt_eou_train.py
这份脚本本身是一个 训练入口脚本(Hydra + PyTorch Lightning),核心功能是:按配置创建 EncDecRNNTBPEEOUModel,并支持从已有 .nemo 初始化、添加/训练 adapter,以及在“词表扩展(新增 <EOU>/<EOB>)”时做权重迁移。
下面按“它用到的技术点 → 在代码/配置里怎么体现 → 原始论文出处”总结。
1) ASR-EOU:把“端点/话轮信息”并入 ASR(<EOU>, <EOB>)
它做什么:
除了输出转写文本外,还让模型在时间轴上预测:
EOU:End Of Utterance(一句话结束)… See the full description on the dataset page: https://huggingface.co/datasets/echodict/asr_rnnt_eou_from_scratch.10232025_scratchThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"total_videos": 0,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {
"observation.image.ego_global": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/10232025_scratch.hf-scratchpadmy-scratch-ai-extensionstokenizer-scratch
load_data.py
Dataset Summary
A music dataset with audio text modality, stored in tfrecord format.
Preprocessing & Augmentation
Preprocessing: aggressive
Augmentation: autoaugment
Splits & Sampling
Split strategy: leave one out
Sampling: active
Quality & Labeling
Quality filtering: strict
Labeling: pseudo label
Files
load_data.py — main artifact of this repository
License
See the… See the full description on the dataset page: https://huggingface.co/datasets/frromano/tokenizer-scratch.transformer-from-scratch-tutorial
Implementing Transformer from Scratch: A Step-by-Step Guide
This repository provides a detailed guide and implementation of the Transformer architecture from the "Attention Is All You Need" paper. The implementation focuses on understanding each component through clear code, comprehensive testing, and visual aids.
For implementions of more recent architectural innovations from DeepSeek, see the Related Implementations section.
Quick Start
View the complete… See the full description on the dataset page: https://huggingface.co/datasets/bird-of-paradise/transformer-from-scratch-tutorial.scratch-dump-0913
scratch dump
Temporary files. No documentation, no support.
eval_smolvla_SCRATCH_lego_yellow_violet_RANDOMThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 23,
"total_frames": 16001,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:23"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/the-fexy/eval_smolvla_SCRATCH_lego_yellow_violet_RANDOM.generator_scratchScratchMath
ScratchMath
Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math
AIED 2026 — 27th International Conference on Artificial Intelligence in Education
Overview
ScratchMath is a multimodal benchmark for evaluating whether MLLMs can analyze handwritten mathematical scratchwork produced by real students. Unlike existing math benchmarks that focus on problem-solving accuracy, ScratchMath targets error diagnosis — identifying… See the full description on the dataset page: https://huggingface.co/datasets/songdj/ScratchMath.details_pszemraj__pythia-31m-KI_v1-2048-scratch
Dataset Card for Evaluation run of pszemraj/pythia-31m-KI_v1-2048-scratch
Dataset Summary
Dataset automatically created during the evaluation run of model pszemraj/pythia-31m-KI_v1-2048-scratch on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_pszemraj__pythia-31m-KI_v1-2048-scratch.relay-ko-s2-scratch-g4-150h-R0-arelay-ko-s2-scratch-g4-53h-R0-arelay-ko-s2-scratch-g4-100h-R0-brelay-ko-s2-scratch-g4-53h-R0-brelay-ko-s2-scratch-g4-100h-R0-arelay-ko-s2-scratch-g4-150h-R0-blibero-smolvla-spatial_scratch_residualThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 432,
"total_frames": 52970,
"total_tasks": 10,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:432"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HollyTan/libero-smolvla-spatial_scratch_residual.Nacidcette dataset n'est le plat de résistance principal tiède de nos modèles.
inference-scratch-llmdetails_pszemraj__pythia-31m-simplepile-lite-2048-scratch-2e
Dataset Card for Evaluation run of pszemraj/pythia-31m-simplepile-lite-2048-scratch-2e
Dataset Summary
Dataset automatically created during the evaluation run of model pszemraj/pythia-31m-simplepile-lite-2048-scratch-2e on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_pszemraj__pythia-31m-simplepile-lite-2048-scratch-2e.details_pszemraj__pythia-31m-simplewiki-scratch-bf16
Dataset Card for Evaluation run of pszemraj/pythia-31m-simplewiki-scratch-bf16
Dataset Summary
Dataset automatically created during the evaluation run of model pszemraj/pythia-31m-simplewiki-scratch-bf16 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_pszemraj__pythia-31m-simplewiki-scratch-bf16.shiying-scratch-archive-publicdetails_pszemraj__pythia-31m-goodwiki-deduped-2048-scratch
Dataset Card for Evaluation run of pszemraj/pythia-31m-goodwiki-deduped-2048-scratch
Dataset Summary
Dataset automatically created during the evaluation run of model pszemraj/pythia-31m-goodwiki-deduped-2048-scratch on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_pszemraj__pythia-31m-goodwiki-deduped-2048-scratch.unpermuted_mixture_2M-scratch-ctx16-51000ood-role-count-scratchkuka_heat_solvability_scratch_20260908_212121This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 10,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"ee_x.pos",
"ee_y.pos",
"ee_z.pos",
"ee_wx.pos",
"ee_wy.pos",
"ee_wz.pos"… See the full description on the dataset page: https://huggingface.co/datasets/ar0s/kuka_heat_solvability_scratch_20260908_212121.RedbullDonnez cette dataset a nos SLM,
c'est comme donner du RedBull a un enfant de 5 ans,
ça se mets a sauter partout.
recette de la dataset :
mes anciennes croyances d'enfants,
mélanges a quelques absurdités.
