comparable
wiki_comparable_corpus_en_de_hi_it_ko_zh
Multilingual Wikipedia Comparable Corpus (en, de, it, ko, hi, zh)
This dataset is a document-level comparable corpus of Wikipedia articles across 6 languages: English (en), German (de), Italian (it), Korean (ko), Hindi (hi), and Chinese (zh).
The key property is alignment across languages: entries are topic-matched such that, for a given index i, dataset["en"][i] is comparable to dataset["de"][i], dataset["it"][i], … (and likewise via the aligned_id field).
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/wiki_comparable_corpus_en_de_hi_it_ko_zh.Open-Canopy-Small-Comparable
Open-Canopy: Towards Very High Resolution Forest Monitoring - Small Comparable Subset
This is not the official repository for the dataset "Open-Canopy". For the original work, please visit the following links:
Datapaper : Pre-print on arXiv: https://arxiv.org/abs/2407.09392.
Code : https://github.com/fajwel/Open-Canopy
Dataset link : https://huggingface.co/datasets/AI4Forest/Open-Canopy.
Small Comparable Subset
The original Open-Canopy dataset is approx.… See the full description on the dataset page: https://huggingface.co/datasets/pierreadorni/Open-Canopy-Small-Comparable.Hindi_comparablecomparable_arabizi
Dataset Card for comparable_arabizi
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source… See the full description on the dataset page: https://huggingface.co/datasets/arbml/comparable_arabizi.Comparable_Wikipediaood-vqa-mscoco-paper-comparable
Paper-comparable VQA OOD pool
400 unique MSCOCO / VQAv2 pairs (eval draw 250, seed 20260730).
Paper-comparable reconstruction — not the unreleased Gulati & Raval 250 IDs.
Built 2026-08-17.
