datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StreamingBench
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.Zhoulifeng-Streaming-Dataset
Zhoulifeng-Streaming-Dataset (峰哥直播语料 SFT 数据清洗流水线)
本项目旨在将“峰哥亡命天涯”的原始直播语音转录文本(ASR Transcripts)清洗、提纯、去重并重写,最终构建出高质量的大模型指令微调(SFT)和思维链(CoT)数据集。
本流水线包含 11 个核心处理步骤,通过规则过滤、统计学过滤、语义降噪以及大模型(LLM)深度重写等手段,将原始的五万多条粗糙问答,提纯为高逻辑密度、高“峰味”浓度的高质量训练集。
🗂️ 核心处理流程 (Data Pipeline)
数据清洗呈现严格的“漏斗状”递减,以下是完整的清洗流程序号与功能说明:
Step 01: 语音文本标点恢复 (Punctuation Generation)
脚本: 01_punctuation_generate.py
输入/输出: 生成 01_punc_transcripts/ 目录
说明: 针对原始 ASR 语音识别出的无断句纯文本,利用模型恢复正确的标点符号,为后续的语义切分和 QA… See the full description on the dataset page: https://huggingface.co/datasets/hzb29/Zhoulifeng-Streaming-Dataset.hls_streaming_mediaproof-pile-2-streaming
ArXiv | Models | Data | Code | Blog | Sample Explorer
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, Sean Welleck
The Proof-Pile-2 is a 55 billion token dataset of mathematical and scientific documents. This dataset was created in order to train the Llemma 7B and Llemma 34B models. It consists of three subsets:
arxiv (29B tokens): the ArXiv subset of RedPajama
open-web-math (15B tokens): The OpenWebMath… See the full description on the dataset page: https://huggingface.co/datasets/xavierdurawa/proof-pile-2-streaming.streaming_vlmstreamingbenchstreaming_vlmcombined-dataset-streamingvietspeech-train-streamingtsa-throughput-streaming-test4Spatial-TTT-Data-Streamingoakink2-vitra-streaming-v1
OakInk2 → VITRA Stage-1 (complete audited release)
This repository contains all 627 physical OakInk2-TaMF sequences converted to
VITRA Stage-1. Every sequence source pair is pinned to
kelvin34501/OakInk-v2 revision 21705616140d726607027e70d58b7837f442ffd8, aligned by exact
frame identity, converted across the four calibrated views, checked by geometry
and every-frame RGB audits, smoke-tested through the VITRA loader, uploaded,
and verified at an immutable commit before local… See the full description on the dataset page: https://huggingface.co/datasets/MIT-Media-Lab/oakink2-vitra-streaming-v1.StreamingOmniDatasets
StreamingOmniDatasets v0.5.0 Streaming CoT Mixed
Four training-ready configs with slim viewer schemas. Structured CoT stores concise, auditable causal state updates rather than private model thinking. Full duplex, barge-in, and simultaneous listen/speak are intentionally deferred.
nemotron-3.5-asr-streaming-0.6bLAION-Mobile-streamingtsa-throughput-streaming-test2combined-dataset-streaming-large-testcombined-dataset-streaming-large-9tsa-throughput-streaming-teststreamingvlm-task-aware-vqa-suite-rowwise
StreamingVLM Task-Aware VQA Suite (Row-wise)
Viewer-ready evaluation samples for task-aware visual-token sensitivity experiments. Every config contains 3,000 deterministic manifest-order samples with the image embedded in each row.
Task groups
task_group
Dataset configs
coarse_object_presence
pope, repope, hpope
general_scene_understanding
vqav2, gqa
fine_grained_visual_evidence
gqa_attribute, mmbench
textual_fine_grained_evidence
textvqa, docvqa… See the full description on the dataset page: https://huggingface.co/datasets/kfkas/streamingvlm-task-aware-vqa-suite-rowwise.omnimcp_data_kafka_streaming_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_data_kafka_streaming_teaser.StreamingBench-Slice
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/jayzhu486/StreamingBench-Slice.combined-dataset-streaming-largesmollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming
A corpus of high quality fine tuning data meant for fine tuning various HelixLM models
Dataset Composition:
A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ...
Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning.
Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.egooops_streamingStreamingVLMDemo_merged_finalstreaming-phi-deidentification-benchmark
Streaming PHI De-Identification Benchmark
Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk.
This dataset captures that. Every record is fully synthetic. It… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.Streaming_Omnistreaming_vlm_fix_1SANA-Streaming-example-training-dataset
SANA-Streaming Example Training Dataset
This repository contains 1,000 aligned reverse video-editing pairs for the
public SANA-Streaming bidirectional V2V training recipe.
License and Terms
This dataset is made available for non-commercial research use only under the
terms in LICENSE. See NOTICE.md for the redistributed content covered by
those terms.
The videos, prompts, annotations, and metadata are all subject to the
non-commercial research-only terms. Do not… See the full description on the dataset page: https://huggingface.co/datasets/Efficient-Large-Model/SANA-Streaming-example-training-dataset.
