CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mjuicem /StreamingBench StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding 🏠 Project Page | 📄 arXiv Paper | 📦 Dataset | 🏅Leaderboard StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟 [NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output. [NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.imagequestion-answering1K<n<10K13 likes12k downloads1y agoHugging Face02hzb29 /Zhoulifeng-Streaming-Dataset Zhoulifeng-Streaming-Dataset (峰哥直播语料 SFT 数据清洗流水线) 本项目旨在将“峰哥亡命天涯”的原始直播语音转录文本(ASR Transcripts)清洗、提纯、去重并重写,最终构建出高质量的大模型指令微调(SFT)和思维链(CoT)数据集。 本流水线包含 11 个核心处理步骤,通过规则过滤、统计学过滤、语义降噪以及大模型(LLM)深度重写等手段,将原始的五万多条粗糙问答,提纯为高逻辑密度、高“峰味”浓度的高质量训练集。 🗂️ 核心处理流程 (Data Pipeline) 数据清洗呈现严格的“漏斗状”递减,以下是完整的清洗流程序号与功能说明: Step 01: 语音文本标点恢复 (Punctuation Generation) 脚本: 01_punctuation_generate.py 输入/输出: 生成 01_punc_transcripts/ 目录 说明: 针对原始 ASR 语音识别出的无断句纯文本,利用模型恢复正确的标点符号,为后续的语义切分和 QA… See the full description on the dataset page: https://huggingface.co/datasets/hzb29/Zhoulifeng-Streaming-Dataset.textfeature-extraction1K<n<10K7 likes3.9k downloads7mo agoHugging Face03DeliberatorArchiver /hls_streaming_media0 likes2.6k downloads2y agoHugging Face04xavierdurawa /proof-pile-2-streaming ArXiv | Models | Data | Code | Blog | Sample Explorer Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, Sean Welleck The Proof-Pile-2 is a 55 billion token dataset of mathematical and scientific documents. This dataset was created in order to train the Llemma 7B and Llemma 34B models. It consists of three subsets: arxiv (29B tokens): the ArXiv subset of RedPajama open-web-math (15B tokens): The OpenWebMath… See the full description on the dataset page: https://huggingface.co/datasets/xavierdurawa/proof-pile-2-streaming.text-generation10B<n<100B1 likes2.1k downloads3y agoHugging Face05xrorrim /streaming_vlmvideo0 likes1.4k downloads1y agoHugging Face06freesky /streamingbenchvideon<1K0 likes472 downloads10mo agoHugging Face07StreamingVLM /streaming_vlm0 likes404 downloads1y agoHugging Face08RuiLu /combined-dataset-streamingtext10M<n<100M0 likes300 downloads10mo agoHugging Face09aiai-laboratory /vietspeech-train-streamingtext100K<n<1M0 likes288 downloads2mo agoHugging Face10JohnVitz /tsa-throughput-streaming-test4tabular100K<n<1M0 likes254 downloads7mo agoHugging Face11THU-SI /Spatial-TTT-Data-Streaming1 likes253 downloads7mo agoHugging Face12MIT-Media-Lab /oakink2-vitra-streaming-v1 OakInk2 → VITRA Stage-1 (complete audited release) This repository contains all 627 physical OakInk2-TaMF sequences converted to VITRA Stage-1. Every sequence source pair is pinned to kelvin34501/OakInk-v2 revision 21705616140d726607027e70d58b7837f442ffd8, aligned by exact frame identity, converted across the four calibrated views, checked by geometry and every-frame RGB audits, smoke-tested through the VITRA loader, uploaded, and verified at an immutable commit before local… See the full description on the dataset page: https://huggingface.co/datasets/MIT-Media-Lab/oakink2-vitra-streaming-v1.imagevideo-classificationn<1K0 likes247 downloads1mo agoHugging Face13skyzhou06 /StreamingOmniDatasets StreamingOmniDatasets v0.5.0 Streaming CoT Mixed Four training-ready configs with slim viewer schemas. Structured CoT stores concise, auditable causal state updates rather than private model thinking. Full duplex, barge-in, and simultaneous listen/speak are intentionally deferred. audio10K<n<100K0 likes242 downloads2mo agoHugging Face14echodict /nemotron-3.5-asr-streaming-0.6b0 likes238 downloads3mo agoHugging Face15sumathiselvan /LAION-Mobile-streaming0 likes226 downloads1mo agoHugging Face16JohnVitz /tsa-throughput-streaming-test2tabular10K<n<100K0 likes214 downloads7mo agoHugging Face17RuiLu /combined-dataset-streaming-large-testtext1M<n<10M0 likes209 downloads10mo agoHugging Face18RuiLu /combined-dataset-streaming-large-9text10M<n<100M0 likes208 downloads10mo agoHugging Face19JohnVitz /tsa-throughput-streaming-testtabular10K<n<100K0 likes204 downloads7mo agoHugging Face20kfkas /streamingvlm-task-aware-vqa-suite-rowwise StreamingVLM Task-Aware VQA Suite (Row-wise) Viewer-ready evaluation samples for task-aware visual-token sensitivity experiments. Every config contains 3,000 deterministic manifest-order samples with the image embedded in each row. Task groups task_group Dataset configs coarse_object_presence pope, repope, hpope general_scene_understanding vqav2, gqa fine_grained_visual_evidence gqa_attribute, mmbench textual_fine_grained_evidence textvqa, docvqa… See the full description on the dataset page: https://huggingface.co/datasets/kfkas/streamingvlm-task-aware-vqa-suite-rowwise.image10K<n<100K0 likes197 downloads2mo agoHugging Face21emgena /omnimcp_data_kafka_streaming_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_data_kafka_streaming_teaser.text-generationn<1K0 likes194 downloads5d agoHugging Face22jayzhu486 /StreamingBench-Slice StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding 🏠 Project Page | 📄 arXiv Paper | 📦 Dataset | 🏅Leaderboard StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟 [NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output. [NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/jayzhu486/StreamingBench-Slice.imagequestion-answering1K<n<10K1 likes190 downloads6mo agoHugging Face23RuiLu /combined-dataset-streaming-largetext1M<n<10M0 likes178 downloads10mo agoHugging Face24david-thrower /smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming A corpus of high quality fine tuning data meant for fine tuning various HelixLM models Dataset Composition: A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ... Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning. Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.texttext-generation1M<n<10M0 likes162 downloads4mo agoHugging Face25Pramish /egooops_streamingvideon<1K0 likes153 downloads1y agoHugging Face26xrorrim /StreamingVLMDemo_merged_finalvideon<1K0 likes152 downloads11mo agoHugging Face27vkatg /streaming-phi-deidentification-benchmark Streaming PHI De-Identification Benchmark Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk. This dataset captures that. Every record is fully synthetic. It… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.documentn<1K0 likes135 downloads6mo agoHugging Face28shamot /Streaming_Omnigatedvideon<1K2 likes135 downloads25d agoHugging Face29xrorrim /streaming_vlm_fix_1videon<1K0 likes124 downloads1y agoHugging Face30Efficient-Large-Model /SANA-Streaming-example-training-dataset SANA-Streaming Example Training Dataset This repository contains 1,000 aligned reverse video-editing pairs for the public SANA-Streaming bidirectional V2V training recipe. License and Terms This dataset is made available for non-commercial research use only under the terms in LICENSE. See NOTICE.md for the redistributed content covered by those terms. The videos, prompts, annotations, and metadata are all subject to the non-commercial research-only terms. Do not… See the full description on the dataset page: https://huggingface.co/datasets/Efficient-Large-Model/SANA-Streaming-example-training-dataset.video-to-video1K<n<10K1 likes116 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.