datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
av_sql_preprocessed_data
Dataset Card for Preprocessed Text-to-SQL Benchmarks
This repository contains preprocessed data for several text-to-SQL benchmarks, as presented in the paper AV-SQL: Decomposing Complex Text-to-SQL Queries with Agentic Views.
The official code for the AV-SQL framework can be found on GitHub: pminhtam/AV-SQL.
Dataset Summary
This repository contains preprocessed data for several text-to-SQL benchmarks:
BIRD
KaggleDBQA
Spider
sciencebenchmark
BEAVER
Spider2-Lite… See the full description on the dataset page: https://huggingface.co/datasets/griffith-bigdata/av_sql_preprocessed_data.dpv2v-preprocessedvisdial-fga-preprocessed
VisDial v1.0, preprocessed for Factor Graph Attention
The preprocessed VisDial v1.0 files used by Factor Graph Attention
(CVPR'19) — code at idansc/fga.
Evaluation is done on VisDialv1.0.
Short description:
VisDial v1.0 contains 1 dialog with 10 question-answer pairs (starting from an image caption) on ~130k images
from COCO-trainval and Flickr, totalling ~1.3 million question-answer pairs.
These are the tokenized, integer-indexed versions of those dialogs: every question… See the full description on the dataset page: https://huggingface.co/datasets/Idan/visdial-fga-preprocessed.dpv2v-preprocessed2ainavox-kazakh-preprocessed
AinaVox Kazakh TTS preprocessed training artifacts
Precomputed training artifacts used for the AinaVox Kazakh IndexTTS-2
experiments. This repository is intended to avoid repeating the expensive
feature-extraction stage when reproducing or extending the training runs.
The binary artifacts are stored in the public Hugging Face Storage Bucket
ruslawik/ainavox-kazakh-preprocessed-data.
This dataset repository contains the documentation and source integrity
manifest.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/ruslawik/ainavox-kazakh-preprocessed.preprocessed_json_patients_symptoms_to_diagnosis22-7-2025_eda_filtered_preprocessed_enPreprocessedMIXED
