datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SparseVideoNav
SparseVideoNav Datasets
This repository contains the real-world navigation datasets released with OpenDriveLab/SparseVideoNav:
BVN: Beyond-the-View Navigation.
IFN: Instruction-Following Navigation.
Project links:
Project page: https://opendrivelab.com/SparseVideoNav
GitHub: https://github.com/OpenDriveLab/SparseVideoNav
Paper: https://arxiv.org/abs/2602.05827
Dataset Summary
SparseVideoNav studies real-world vision-language navigation with sparse future… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab/SparseVideoNav.sparse-reward-long-tasks
Sparse Reward Long Tasks
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/sparse-reward-long-tasks.walnut-edip-sparse20
Scitomo Walnut EDIP Sparse-20 prepared dataset
This is a derived Scitomo Sparse-20 preparation of the public Walnut-1
cone-beam X-ray CT acquisition. It is not the original Walnut archive, not a
published EDIP reconstruction, and not a blessed Scitomo result. The package
contains measured projections, corrected vector-cone geometry, and the
published AGD-50 evaluation reference used by the maintained Scitomo Walnut
EDIP evidence workflow.
Provenance and attribution… See the full description on the dataset page: https://huggingface.co/datasets/scitomo/walnut-edip-sparse20.multilingual-automotive-sparse
Multilingual Automotive Sparse Retrieval Corpus
A synthetic training corpus for fine-tuning multilingual sparse retrieval
models (SPLADE-family) on automotive, supply-chain, and geopolitics content.
Designed to teach a model cross-lingual alignment: a Japanese query should
retrieve the right English passage, and vice versa.
Corpus contents
File
Rows
Purpose
concepts.jsonl
4995
Raw concept records
train_triplets.jsonl
134880
Flattened (query, positive… See the full description on the dataset page: https://huggingface.co/datasets/cp500/multilingual-automotive-sparse.Dr.Sparse-RL-train-562
Dr.Sparse SpGEMM training pool (562 matrices)
The complete RL / selector training pool of Dr.Sparse (branch v2): 562 SuiteSparse matrices in the
harness .bin layout (int32 rows, cols, nnz; int32 row_ptr; int32 col_ind; float32 values; float32 x),
laid out as level1_small/ (91), level2_medium/ (273), level3_large/ (198); the huge tier is deliberately left out of training and evaluation.
Every matrix has a cuSPARSE SpGEMM reference (C = AA, or AA^T when rectangular) on an H200;… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-RL-train-562.stack-v2-sparse-classes-10k
Stack v2 Sparse Python Classes 10k
This is a 10,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments.
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters.
Splits
train.jsonl: 9,000
val.jsonl: 500
test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-10k.stack-v2-sparse-classes-75kplus
Stack v2 Sparse Python Classes 75kplus
This is a frozen snapshot with 75829 samples for Diffusion + Autoregressive hybrid code generation experiments.
Splits
train.jsonl: 74829
val.jsonl: 500
test.jsonl: 500
all.jsonl: 75829
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-75kplus.stack-v2-sparse-classes-36k
Stack v2 Sparse Python Classes 36k
This is a 36,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments.
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters.
Splits
train.jsonl: 35,000
val.jsonl: 500
test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-36k.SparsePlug-Compression-Benchmarks
SparsePlug Hardware Profiler & Compression Benchmarks
Author: Brad Wallace (coo@koba42.com)Source Code: github.com/tensorrent/prime-fheLicense: Apache-2.0
Component Description
Hardware-adaptive sparsity profiling, TrinityWasm compression, and zero-degradation motif throughput benchmarks at 0%, 75%, and 90% sparsity.
Verified Benchmark Performance
Metric
Value
Vector Dim
384
Step Latency Ns
1420
Full Evaluation Ms
0.55
Throughput… See the full description on the dataset page: https://huggingface.co/datasets/K42COO/SparsePlug-Compression-Benchmarks.WebVoyager-GAIA-SparseAttention-Offline-Qwen3-VL-30B-A3B
Sparse Attention for Web Agents — Offline Replay Records
Step-level records from an offline replay study comparing three sparse-attention methods against
full attention on browser-agent trajectories.
Model: Qwen3-VL-30B-A3B-Instruct (48 layers, 128 experts / 8 active, GQA 32:4, head_dim 128)
Hardware: NVIDIA GB10 (sm_121)
Tasks: 50 complete trajectories sampled from WebVoyager + GAIA — 332 steps, of which 190 emit an
element index. Every task was completed successfully by the… See the full description on the dataset page: https://huggingface.co/datasets/shiqihe/WebVoyager-GAIA-SparseAttention-Offline-Qwen3-VL-30B-A3B.resume_sparser
