vida
Datasets
All datasets matching “vida”VidaForge-3M
3.14 million scene-level video clips with multi-level captions, camera labels, semantic tags, quality signals, and duplicate groups.
Paper
·
VidaForge Code
·
Project Blog
·
Source Dataset
Overview
VidaForge-3M is a large-scale video pretraining dataset produced with
VidaForge, an open data pipeline for
building and studying video foundation model pretraining data. The pipeline and
dataset are described in the paper
VidaForge: Open Research Infrastructure… See the full description on the dataset page: https://huggingface.co/datasets/VidaForge/VidaForge-3M.VID-ADarXiv: https://arxiv.org/abs/2603.13964
vi-dataset-for-pretrain
Dataset Card for "vi-dataset-for-pretrain"
This is a combination of multiple Vietnamese dataset for pretraining CLMs such as GPT, GPT2, etc.
The dataset consists of:
vietgpt/covid_19_news_vi
hieunguyen1053/binhvq-news-corpus
oscar (unshuffled_deduplicated_vi)
vietgpt/wikipedia_vi
Dataset info
Splits
N.o examples
Size
Train
23,891,116
77.36 GB
Validation
1,257,428
4.06 GB
Total
25,148,544
81.43 GB
vi-dataset-for-pretrain
Dataset Card for "vi-dataset-for-pretrain"
This is a combination of multiple Vietnamese dataset for pretraining CLMs such as GPT, GPT2, etc.
The dataset consists of:
vietgpt/covid_19_news_vi
hieunguyen1053/binhvq-news-corpus
oscar (unshuffled_deduplicated_vi)
vietgpt/wikipedia_vi
Dataset info
Splits
N.o examples
Size
Train
23,891,116
77.36 GB
Validation
1,257,428
4.06 GB
Total
25,148,544
81.43 GB
CARE-PDOverview
Please read carefully the terms and conditions and any accompanying documentation at
neurips2025.care-pd.ca/terms-of-use
before you download and/or use the CARE-PD dataset.
Project page:
https://neurips2025.care-pd.ca/
CARE-PD is the largest publicly available archive of 3D mesh gait data for Parkinson's Disease (PD) and the first to include data collected across multiple sites.
The dataset aggregates 9 cohorts from 8 clinical sites, including 362 participants… See the full description on the dataset page: https://huggingface.co/datasets/vida-adl/CARE-PD.ViDAS
ViDAS Dataset
Abstract
We present a novel dataset aimed at advancing danger analysis and assessment by addressing the challenge of quantifying danger in video content and identifying how human-like a Large Language Model (LLM) evaluator is for the same. This is achieved by compiling a collection of 100 YouTube videos featuring various events. Each video is annotated by human participants who provided danger ratings on a scale from 0 (no danger to humans) to 10… See the full description on the dataset page: https://huggingface.co/datasets/pranked03/ViDAS.
