datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GRAB
GRAB: A Challenging GRaph Analysis Benchmark for Large Multimodal Models
GRAB consists of 3 splits: GRAB, GRAB-real and GRAB-lite. This is the dataset for GRAB.
Dataset Summary
Large multimodal models (LMMs) have exhibited proficiencies across many visual tasks. Although numerous benchmarks exist to evaluate model performance, they increasingly have insufficient headroom and are unfit to evaluate the next generation of frontier LMMs.
To overcome this, we… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/GRAB.bls_cpi
Changelog
2025-01-18
I decided that I'll name the column survey, instead of consumer. I'll also set the value to the description, instead of the code.
I didn't realize that the pandas version of "Use this dataset" includes the filename. I'll remove the date from the filename. So that people do not have to change their code.
I have been updating this using the UI, but I will create a script in Spaces to update this from the BLS site.
2025-01-12
While using… See the full description on the dataset page: https://huggingface.co/datasets/robert-co/bls_cpi.PatternNet
Dataset Card for "PatternNet"
Licensing Information
For research purposes.
Citation Information
PatternNet: A benchmark dataset for performance evaluation of remote sensing image retrieval
@article{zhou2018patternnet,
title = {PatternNet: A benchmark dataset for performance evaluation of remote sensing image retrieval},
author = {Zhou, Weixun and Newsam, Shawn and Li, Congmin and Shao, Zhenfeng},
year = 2018,
journal = {ISPRS… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/PatternNet.x_dataset_041134
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_041134.x_dataset_0405200
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0405200.zerobench
ZeroBench: An Impossible* Visual Benchmark for Contemporary Large Multimodal Models
🌐 Project Page | 📄 Paper | GitHub
*Given the recent rapid progress on benchmarks, we do not imagine ZeroBench will remain "impossible" for long!
v3 changelog:
23/12/2025 -- The updated version of the main questions and subquestions is now on the main branch. Previous version available on v2 branch.
Question 5:
question_text: added Report left, right, and total in kg, writing "kg" after each number.… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/zerobench.NWPU-RESISC45
Dataset Card for "NWPU-RESISC45"
Licensing Information
[CC-BY-SA]
Citation Information
Remote sensing image scene classification: Benchmark and state of the art
@article{cheng2017remote,
title = {Remote sensing image scene classification: Benchmark and state of the art},
author = {Cheng, Gong and Han, Junwei and Lu, Xiaoqiang},
year = 2017,
journal = {Proceedings of the IEEE},
publisher = {IEEE},
volume = 105… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/NWPU-RESISC45.roberta_pretrain
Dataset Card for RoBERTa Pretrain
Dataset Summary
This is the concatenation of the datasets used to Pretrain RoBERTa.
The dataset is not shuffled and contains raw text. It is packaged for convenicence.
Essentially is the same as:
from datasets import load_dataset, concatenate_datasets
bookcorpus = load_dataset("bookcorpus", split="train")
openweb = load_dataset("openwebtext", split="train")
cc_news = load_dataset("cc_news", split="train")
cc_news =… See the full description on the dataset page: https://huggingface.co/datasets/gsgoncalves/roberta_pretrain.x_dataset_041213
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_041213.SciFIBench
SciFIBench
Jonathan Roberts, Kai Han, Neil Houlsby, and Samuel Albanie
NeurIPS 2024
Note: This repo has been updated to add two splits ('General_Figure2Caption' and 'General_Caption2Figure') with an additional 1000 questions. The original version splits are preserved and have been renamed as follows: 'Figure2Caption' -> 'CS_Figure2Caption' and 'Caption2Figure' -> 'CS_Caption2Figure'.
Dataset Summary
The SciFIBench (Scientific Figure… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/SciFIBench.x_dataset_0406135
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0406135.x_dataset_0401151
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0401151.CLRS
Dataset Card for "CLRS"
Licensing Information
For academic purposes.
Citation Information
CLRS: Continual Learning Benchmark for Remote Sensing Image Scene Classification
@article{s20041226,
title = {CLRS: Continual Learning Benchmark for Remote Sensing Image Scene Classification},
author = {Li, Haifeng and Jiang, Hao and Gu, Xin and Peng, Jian and Li, Wenbo and Hong, Liang and Tao, Chao},
year = 2020,
journal = {Sensors}… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/CLRS.x_dataset_0410139
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0410139.x_dataset_040752
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_040752.x_dataset_040849
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_040849.x_dataset_0409154
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0409154.roberta-datax_dataset_040484
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_040484.x_dataset_0403203
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0403203.wikitext-103-raw-v1_sents_min_len10_max_len30_princeton-nlp_sup-simcse-roberta-largeAID_MultiLabel
Dataset Card for "AID_MultiLabel"
Licensing Information
CC0: Public Domain
Citation Information
Imagery:
AID: A benchmark data set for performance evaluation of aerial scene classification
Multilabels:
Relation Network for Multi-label Aerial Image Classification
@article{xia2017aid,
title = {AID: A benchmark data set for performance evaluation of aerial scene classification},
author = {Xia, Gui-Song and Hu, Jingwen and Hu, Fan and Shi, Baoguang… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/AID_MultiLabel.zerobench_no_answersGID
Dataset Card for "GID"
Licensing Information
Public domain.
Citation Information
Land-cover classification with high-resolution remote sensing images using transferable deep models
@article{GID2020,
title = {Land-cover classification with high-resolution remote sensing images using transferable deep models},
author = {Tong, Xin-Yi and Xia, Gui-Song and Lu, Qikai and Shen, Huanfeng and Li, Shengyang and You, Shucheng and Zhang, Liangpei},
year… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/GID.recipe_RL_data_roberta-base
Dataset Description
Structure
Consists of 5 fields
Each row corresponds to a policy - sequence of actions, given an initial <START> state, and corresponding rewards at each step.
Fields
steps, step_attn_masks, rewards, actions, dones
Field descriptions
steps (List of lists of Ints) - tokenized step tokens of all the steps in the policy sequence (here we use the roberta-base tokenizer, as roberta-base would be used to encode each step of a recipe)… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousSub/recipe_RL_data_roberta-base.RSI-CB256
Dataset Card for "RSI-CB256"
Licensing Information
For academic purposes.
Citation Information
Exploring Models and Data for Remote Sensing Image Caption Generation
@article{lu2017exploring,
title = {Exploring Models and Data for Remote Sensing Image Caption Generation},
author = {Lu, Xiaoqiang and Wang, Binqiang and Zheng, Xiangtao and Li, Xuelong},
journal = {IEEE Transactions on Geoscience and Remote Sensing},
volume = 56… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/RSI-CB256.recipe_RL_data_ONLY_CORRECT_roberta-basebdappv
BDAPPV — Aerial Images of Rooftop Photovoltaic Installations
BDAPPV is a dataset of aerial images of rooftop PV installations in France and Belgium,
with segmentation masks and installation metadata. Images are provided by two aerial
imagery providers (Google and IGN), making it suitable for both segmentation/classification
benchmarks and distribution shift evaluation across imagery sources.
Paper: Kasmi et al., Scientific Data, 2023 — arXiv:2209.03726
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Roberthowe/bdappv.GRAB-real
GRAB: A Challenging GRaph Analysis Benchmark for Large Multimodal Models
GRAB consists of 3 splits: GRAB, GRAB-real and GRAB-lite. This is the dataset for GRAB-real.
Dataset Summary
The 'real' split of the GRAB benchmark is comprised of 1114 real questions covering 4 real tasks: noise, screenshots, whiteboard, sketches.Large multimodal models (LMMs) have exhibited proficiencies across many visual tasks. Although numerous benchmarks exist to evaluate model… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/GRAB-real.cc12m_princeton-nlp_sup-simcse-roberta-large
