datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fav_db_test_7FAVOR
A Comprehensive Benchmark for Fine-Grained Video Motion Understanding
🔥 News
2025.09.18 🎉 FAVOR-Bench has been accepted by NeurIPS 2025 Datasets and Benchmarks Track!
2025.03.19 🌟 We released Favor-Bench, a new benchmark for fine-grained video motion understanding that spans both ego-centric and third-person perspectives with comprehensive evaluation including both close-ended QA tasks and open-ended descriptive tasks!
Introduction
Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/FAVOR-Bench/FAVOR.fav_db_test_1fav_db_test_9fav_db_test_19fav_db_test_4fava-data
FAVA Datasets
FAVA datasets include: annotation data and training data.
Dataset Details
Annotation Data
The annotation dataset includes 460 annotated passages identifying and editing errors using our hallucination taxonomy. This dataset was used for the fine-grained error detection task, using the annotated passages as the gold passages.
Training Data
The training data includes 35k training instances of erroneous input and corrected output pairs… See the full description on the dataset page: https://huggingface.co/datasets/fava-uw/fava-data.fav_db_test_11fav_db_test_17fav_db_test_12fav_db_test_16fav_db_test_15fav_db_test_8fav_db_test_13faviqtupuri_english
French-English to Tupuri Translation Dataset
Dataset Summary
This dataset provides sentence-level translations from French and English into Tupuri, a low-resource language spoken in northern Cameroon and southwestern Chad. It aims to support research in African language translation and NLP for underrepresented languages.
Supported Tasks and Leaderboards
Machine Translation: French/English → Tupuri.
Multilingual modeling: Useful for zero-shot or low-resource… See the full description on the dataset page: https://huggingface.co/datasets/Favourez/tupuri_english.titan-safety-check-catalog
Titan safety check catalog
This dataset packages Titan's shipped pre-trade safety-check catalog as small structured data:
gates.csv
gates.json
Related note: article-verifiable-receipts.md explains Titan's verifiable decision receipt boundary.
Each row includes:
order: cascade order from the shipped gate pipeline
gate_id: backend/API identifier
label: operator-facing display label
category: public Operator Guide group
applies_to: signal class from the shipped pipeline… See the full description on the dataset page: https://huggingface.co/datasets/favlo/titan-safety-check-catalog.fav_db_test_14dante-faves-allCOIG-CQIA-fullI-Favour-The-Villaness-JSONL-PT-1fav_db_test_5fava-data
FAVA Datasets
FAVA datasets include: annotation data and training data.
Dataset Details
Annotation Data
The annotation dataset includes 460 annotated passages identifying and editing errors using our hallucination taxonomy. This dataset was used for the fine-grained error detection task, using the annotated passages as the gold passages.
Training Data
The training data includes 35k training instances of erroneous input and corrected output pairs… See the full description on the dataset page: https://huggingface.co/datasets/nkhuggme/fava-data.dante-faves-all-with-titlesfav_db_test_0fav_db_test_2fav_db_test_3fav_db_test_6fav_db_test_10fav_db_test_18
