datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VCR-Bench
VCR-Bench ( A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning)
🌐 Homepage | 🤗 Dataset | 🤗 Paper | 📖 arXiv | GitHub
Dataset Details
As shown in the figure below, current video benchmarks often lack comprehensive annotations of CoT steps, focusing only on the accuracy of final answers during model evaluation while neglecting the quality of the reasoning process. This evaluation approach makes it difficult to comprehensively evaluate model’s… See the full description on the dataset page: https://huggingface.co/datasets/VLM-Reasoning/VCR-Bench.vcl-vibebench
VCL VibeBench
Stop trusting benchmark slides. Run it yourself.
Practical AI model prompts from Vibe Coder's Life.
This dataset mirrors the open-source GitHub suites. It is not a blended intelligence leaderboard. There is no LLM judge. 9/12 on Score means nine JavaScript helpers compiled and passed hidden unit tests — not “75% smart.”
Configs
Config
What it is
Rows
fun
Fun 1.0 — short copy-paste prompts
10
dev
Dev 1.1 — coding / debugging prompts
10… See the full description on the dataset page: https://huggingface.co/datasets/kondasviktor/vcl-vibebench.vcell-coexpr-baselines
ConvergeCELL Coexpression Baselines — v0.1.0
Per-tissue gene–gene correlation matrices trained on the
ConvergeCELL Pseudobulk Panel
(panel size pb20, restricted to 8,000 HVGs).
Each baseline is a fitted :class:CoexpressionBaselineModel — predicts gene
expression via Y = (X − μ)/σ @ C @ σ + μ. Useful as the co-expression
floor any foundation model must beat to claim it has learned anything
beyond gene–gene correlation structure.
Baselines (12)
Tissue
File
#… See the full description on the dataset page: https://huggingface.co/datasets/nicolas-lynn/vcell-coexpr-baselines.top-funded-startups-2026
The 250 Most-Funded Startups of 2026
The 250 private startups that raised the most venture capital in calendar year 2026,
ranked by total disclosed new funding.
Columns: rank, company, slug, raised_2026_usd, rounds, last_round,
lead_investor, country, industry, website, url
Methodology. Source-verified against primary announcements (Reuters, Bloomberg,
TechCrunch, company press releases). Duplicate rounds de-duplicated;
valuations-mistaken-for-raises corrected. Amounts are whole… See the full description on the dataset page: https://huggingface.co/datasets/indexed-vc/top-funded-startups-2026.openve_q5_vcos_tcfg1_step200_shift3c7-vctk-retrain
C7 VCTK Retrain
Seam-3 C7-retrain VCTK content-disjoint evaluation.
Design
Retrain from scratch on content-reserved train_4000 (40 text_ids R reserved for novel eval)
Two arms within run:
control_no_lora: LoRA attached but frozen, AF+DT trainable
lora_trainable: q/v r16 LoRA trainable, AF+DT trainable
Eval on same 25 heldout speakers with:
content_seen: 512 windows, text_ids NOT in R (seen during training)
content_novel: 512 windows, text_ids IN R (unseen… See the full description on the dataset page: https://huggingface.co/datasets/aistocrat/c7-vctk-retrain.vctk-tokens
MintTTS Pre-tokenized Audio Tokens
Pre-extracted audio codec tokens for TTS training.
Source
Dataset: sanchit-gandhi/vctk
Codec: MOSS-Audio-Tokenizer-Nano
Codec sample rate: 48,000 Hz (stereo)
Frame rate: 12.5 Hz (1 frame = 80ms)
Stats
Metric
Value
Total samples
88,156
Total audio hours
81.5h
Codebooks
16
Avg frames/sample
41.6
Avg duration
3.3s
Format
JSONL file (manifest.jsonl) where each line is:
{
"text": "The… See the full description on the dataset page: https://huggingface.co/datasets/somu9/vctk-tokens.
