CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01impresso-project /wiki_comparable_corpus_en_de_hi_it_ko_zh Multilingual Wikipedia Comparable Corpus (en, de, it, ko, hi, zh) This dataset is a document-level comparable corpus of Wikipedia articles across 6 languages: English (en), German (de), Italian (it), Korean (ko), Hindi (hi), and Chinese (zh). The key property is alignment across languages: entries are topic-matched such that, for a given index i, dataset["en"][i] is comparable to dataset["de"][i], dataset["it"][i], … (and likewise via the aligned_id field). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/wiki_comparable_corpus_en_de_hi_it_ko_zh.tabular10K<n<100K1 likes30 downloads8mo agoHugging Face02JINIAC /Hindi_comparabletext100K<n<1M0 likes22 downloads2y agoHugging Face03arbml /comparable_arabizi Dataset Card for comparable_arabizi Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source… See the full description on the dataset page: https://huggingface.co/datasets/arbml/comparable_arabizi.text10K<n<100K0 likes17 downloads3y agoHugging Face04arbml /Comparable_Wikipediatext10K<n<100K0 likes15 downloads4y agoHugging Face05rlogger /ood-vqa-mscoco-paper-comparable Paper-comparable VQA OOD pool 400 unique MSCOCO / VQAv2 pairs (eval draw 250, seed 20260730). Paper-comparable reconstruction — not the unreleased Gulati & Raval 250 IDs. Built 2026-08-17. imagevisual-question-answeringn<1K0 likes4 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.