datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
combined-dataset-streamingtsa-throughput-streaming-test4vietspeech-train-streamingcombined-dataset-streaming-large-testtsa-throughput-streaming-test2StreamingOmniDatasets
StreamingOmniDatasets v0.5.0 Streaming CoT Mixed
Four training-ready configs with slim viewer schemas. Structured CoT stores concise, auditable causal state updates rather than private model thinking. Full duplex, barge-in, and simultaneous listen/speak are intentionally deferred.
combined-dataset-streaming-large-9tsa-throughput-streaming-teststreamingvlm-task-aware-vqa-suite-rowwise
StreamingVLM Task-Aware VQA Suite (Row-wise)
Viewer-ready evaluation samples for task-aware visual-token sensitivity experiments. Every config contains 3,000 deterministic manifest-order samples with the image embedded in each row.
Task groups
task_group
Dataset configs
coarse_object_presence
pope, repope, hpope
general_scene_understanding
vqav2, gqa
fine_grained_visual_evidence
gqa_attribute, mmbench
textual_fine_grained_evidence
textvqa, docvqa… See the full description on the dataset page: https://huggingface.co/datasets/kfkas/streamingvlm-task-aware-vqa-suite-rowwise.combined-dataset-streaming-largesmollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming
A corpus of high quality fine tuning data meant for fine tuning various HelixLM models
Dataset Composition:
A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ...
Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning.
Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.tsa-throughput-streaming-test3StreamingBenchcombined-dataset-streaming-large-7combined-dataset-streaming-large-3combined-dataset-streaming-large-5combined-dataset-streaming-small-test8combined-dataset-streaming-small-test6combined-dataset-streaming-small-test14streaming-gebd-causal
Audited Kinetics-GEBD Causal Metadata
This metadata-only release converts the publicly released Kinetics-GEBD
annotations into an auditable 24 FPS causal training representation. It does
not redistribute Kinetics or YouTube video bytes.
Splits
Hub split
Official source file
Records
Meaning
train
k400_train_raw_annotation.pkl
18,808
Public GEBD training annotations
validation
k400_val_raw_annotation.pkl
18,815
Public Kinetics-GEBD validation… See the full description on the dataset page: https://huggingface.co/datasets/kfkas/streaming-gebd-causal.speed_video_streaming_threadThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 1720,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imstevenpmwork/speed_video_streaming_thread.combined-dataset-streaming-small-test7venc_streaming_high_res_chunk_branch10ep12secThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 10,
"total_frames": 3116,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 50,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imstevenpmwork/venc_streaming_high_res_chunk_branch10ep12sec.combined-dataset-streaming-small-testmodel0-boundary-streaming
Model Card: Streaming Terminal Log Boundary Predictor (Phi-4 LoRA)
🤖 Model Details
Base Model: unsloth/Phi-4-unsloth-bnb-4bit (14B Parameters)
Architecture: LoRA Adapters (PEFT)
Task: Binary Classification (Terminal Event Boundary Detection)
Quantization: 4-bit (bitsandbytes)
Language: English / Bash / Terminal XML
🎯 Intended Use
This model serves as "Model 0" for the Winter 2026 iteration of the AutoDocs project. Its primary function is to segment a… See the full description on the dataset page: https://huggingface.co/datasets/librocubic/model0-boundary-streaming.StreamingCoT
Streaming Reasoning Math Train/Eval
This dataset is built for streaming real-time reasoning. It provides math
reasoning examples where a model should reason while the input is being read,
rather than waiting for the complete problem context before starting to think.
The dataset is organized into two Hugging Face configs:
Train: supervised fine-tuning data for learning streaming reasoning traces.
Eval: held-out evaluation data with three benchmark splits.
The training reasoning… See the full description on the dataset page: https://huggingface.co/datasets/JunlongTong/StreamingCoT.streaming-smokeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 5,
"total_frames": 50,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/irvinh/streaming-smoke.venc_streaming_thread_rerun_moving_1min3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 1712,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imstevenpmwork/venc_streaming_thread_rerun_moving_1min3.venc_streaming_med_res2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 1715,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imstevenpmwork/venc_streaming_med_res2.venc_streaming_high_res_chunk_branch10ep12sec_mybranchThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 10,
"total_frames": 3395,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 50,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imstevenpmwork/venc_streaming_high_res_chunk_branch10ep12sec_mybranch.
