datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
owid_charts_en_beirThis is a copy of https://huggingface.co/datasets/jinaai/owid_charts_en reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai" for… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/owid_charts_en_beir.owid_charts_en
Our World In Data (OWID) Evaluation Dataset
We sampled a set of ~5k charts and articles from Our World In Data to produce this evaluation set.
The set is comprised of 666 queries (text snippets from the articles in reference to the charts), and a total of 1k unique charts.
This particular dataset is a subsample of 1000 random charts from the full dataset which can be found here.
The text_description column contains OCR text extracted from the images using EasyOCR.… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/owid_charts_en.owid
Our World in Data (Recaptioned)
Chart images, short data-insight posts, and narrative articles from Our World in Data, prepared as image-text data by the Swiss AI Initiative vision team. OWID publishes research and visualizations on global living conditions and development, covering topics like health, energy, climate, poverty, food and democracy.
The release has three files.
grapher_charts.parquet has 3,487 standalone charts with generated captions
data_insights.parquet has… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/owid.owid_charts_en_deprecated
Our World In Data (OWID) Evaluation Dataset
We sampled a set of ~5k charts and articles from Our World In Data to produce this evaluation set.
The set is comprised of 666 queries (text snippets from the articles in reference to the charts), and a total of 1k unique charts.
This particular dataset is a subsample of 1000 random charts from the full dataset which can be found here.
The text_description column contains OCR text extracted from the images using EasyOCR.… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/owid_charts_en_deprecated.
