datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stack-v3-train
🥞 The Stack v3
What is it?
What is being released
How to download and use it
Dataset statistics
Dataset structure
Dataset creation
Considerations for using the data
Additional information
What is it?
The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/TechnoBaptist/stack-v3-train.HARD-TIME
HARD-TIME
HARD-TIME evaluates whether video-language models can localize moments in time and avoid answers that are
not supported by the video. This repository contains benchmark annotations for the evaluation tasks; it does
not include source videos, transcripts, or evidence artifacts. Access to the annotations does not grant any
license or reuse rights for the underlying third-party content.
Configs
Config
Rows
Description
temporal_retrieval
4,812… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/HARD-TIME.golden-vault-v0
Golden Vault
Golden Vault is a curated long-video understanding dataset with dense video descriptions,
audio transcripts, timestamp-grounded questions, evidence-grounded questions, multi-hop
questions, and contrastive unanswerable questions.
VAULT stands for Video-Audio Understanding over Long Timelines.
Dataset At A Glance
This release is intentionally pruned. The QA/evaluation configs are preserved in full,
while broad caption configs are limited to a 500-video… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/golden-vault-v0.
