aurora
Datasets
All datasets matching “aurora”AuroraCap-trainset
AuroraCap Trainset
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: AuroraCap Model
Huggingface: VDC Benchmark
Huggingface: Trainset
Features
We use over 20 million high-quality image/video-text pairs to train AuroraCap in three stages.
Pretraining stage. We first align visual features with the word embedding space of LLMs. To achieve this, we freeze the pretrained ViT and LLM, training solely the vision-language connector.
Vision stage. We… See the full description on the dataset page: https://huggingface.co/datasets/wchai/AuroraCap-trainset.gta-files-auroraaurora_rollout_betaaurora-training-data
Aurora Video-Editing Training Data
The video-editing data used to train the editor in "Aurora: Unified Video
Editing with a Tool-Using Agent"
(arXiv:2605.18748).
Code: github.com/yeates/Aurora.
This release contains only the subsets reported in Table 1 of the paper.
The data is packaged as WebDataset
tar shards so the HuggingFace dataset viewer renders each sample's video next to
its text prompt. Media is stored uncompressed (byte-identical to the clips used
in training);… See the full description on the dataset page: https://huggingface.co/datasets/yeates/aurora-training-data.epstein-files-ocr-datasets-1-8-early-release
Epstein Files OCR — Datasets 1–8 (Early Release)
Work in Progress (WIP)
This is an early publication. We are actively working on improving OCR quality and expanding coverage.
Dataset Summary
This dataset contains page-level OCR output (as Markdown) from a public release of documents related to Jeffrey Epstein / the Epstein case.
Each Markdown file represents one scanned page converted to text using an automated OCR pipeline. The dataset is designed for:
Question… See the full description on the dataset page: https://huggingface.co/datasets/aurora2424/epstein-files-ocr-datasets-1-8-early-release.AURORA
Read the paper here: https://arxiv.org/abs/2407.03471. IMPORTANT: Please check out our GitHub repository for more instructions on how to also access the Something-Something-Edit subdataset, which we can't publish directly: https://github.com/McGill-NLP/AURORA
