datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chessdbcn
Static dumps of chessdb.cn (cdb)
This is a collection of static dumps of chessdb.cn, the largest online database
of chess positions and openings. The dumps can be probed with the help of cdbdirect.
Statistics
20240814
20241114
20250608
20251115
20260702
Positions
44262943988
46456113101
48454315961
53035759834
63996865341
Scored moves
88251700855
92617595285
96605581079
128743456819
151652263234
Size
821GB
808GB
846GB
997GB
1.2TB
roberta-pt-checkpointsGenshin
Genshin Dataset
The dataset contains different angles images of 64 characters captured from Genshin game manually, 20 pictures for each character.
The dataset is intended for Genshin character lora model training.
Dataset Details
Character Proportion
The character should occupy a large proportion of the image. There should not be too much background, and ideally, the image should be just wrapped around the body (screenshots can be taken if necessary).… See the full description on the dataset page: https://huggingface.co/datasets/RobertLau/Genshin.SATIN
Dataset Card for SATIN
Dataset Summary
SATIN (SATellite ImageNet) is a metadataset containing 27 constituent satellite and aerial image datasets spanning 6 distinct tasks: Land Cover, Land Use,
Hierarchical Land Use, Complex Scenes, Rare Scenes, and False Colour Scenes. The imagery is globally distributed, comprised of resolutions spanning 5 orders
of magnitude, multiple fields of view sizes, and over 250 distinct class labels. Presented at ICCV '23 TNGCV Workshop.… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/SATIN.GRAB
GRAB: A Challenging GRaph Analysis Benchmark for Large Multimodal Models
GRAB consists of 3 splits: GRAB, GRAB-real and GRAB-lite. This is the dataset for GRAB.
Dataset Summary
Large multimodal models (LMMs) have exhibited proficiencies across many visual tasks. Although numerous benchmarks exist to evaluate model performance, they increasingly have insufficient headroom and are unfit to evaluate the next generation of frontier LMMs.
To overcome this, we… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/GRAB.x_dataset_041134
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_041134.PatternNet
Dataset Card for "PatternNet"
Licensing Information
For research purposes.
Citation Information
PatternNet: A benchmark dataset for performance evaluation of remote sensing image retrieval
@article{zhou2018patternnet,
title = {PatternNet: A benchmark dataset for performance evaluation of remote sensing image retrieval},
author = {Zhou, Weixun and Newsam, Shawn and Li, Congmin and Shao, Zhenfeng},
year = 2018,
journal = {ISPRS… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/PatternNet.NWPU-RESISC45
Dataset Card for "NWPU-RESISC45"
Licensing Information
[CC-BY-SA]
Citation Information
Remote sensing image scene classification: Benchmark and state of the art
@article{cheng2017remote,
title = {Remote sensing image scene classification: Benchmark and state of the art},
author = {Cheng, Gong and Han, Junwei and Lu, Xiaoqiang},
year = 2017,
journal = {Proceedings of the IEEE},
publisher = {IEEE},
volume = 105… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/NWPU-RESISC45.zerobench
ZeroBench: An Impossible* Visual Benchmark for Contemporary Large Multimodal Models
🌐 Project Page | 📄 Paper | GitHub
*Given the recent rapid progress on benchmarks, we do not imagine ZeroBench will remain "impossible" for long!
v3 changelog:
23/12/2025 -- The updated version of the main questions and subquestions is now on the main branch. Previous version available on v2 branch.
Question 5:
question_text: added Report left, right, and total in kg, writing "kg" after each number.… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/zerobench.x_dataset_0405200
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0405200.x_dataset_041213
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_041213.roberta_pretrain
Dataset Card for RoBERTa Pretrain
Dataset Summary
This is the concatenation of the datasets used to Pretrain RoBERTa.
The dataset is not shuffled and contains raw text. It is packaged for convenicence.
Essentially is the same as:
from datasets import load_dataset, concatenate_datasets
bookcorpus = load_dataset("bookcorpus", split="train")
openweb = load_dataset("openwebtext", split="train")
cc_news = load_dataset("cc_news", split="train")
cc_news =… See the full description on the dataset page: https://huggingface.co/datasets/gsgoncalves/roberta_pretrain.x_dataset_0401151
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0401151.SciFIBench
SciFIBench
Jonathan Roberts, Kai Han, Neil Houlsby, and Samuel Albanie
NeurIPS 2024
Note: This repo has been updated to add two splits ('General_Figure2Caption' and 'General_Caption2Figure') with an additional 1000 questions. The original version splits are preserved and have been renamed as follows: 'Figure2Caption' -> 'CS_Figure2Caption' and 'Caption2Figure' -> 'CS_Caption2Figure'.
Dataset Summary
The SciFIBench (Scientific Figure… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/SciFIBench.x_dataset_0406135
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0406135.mostx_dataset_0410139
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0410139.x_dataset_040752
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_040752.x_dataset_040849
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_040849.CLRS
Dataset Card for "CLRS"
Licensing Information
For academic purposes.
Citation Information
CLRS: Continual Learning Benchmark for Remote Sensing Image Scene Classification
@article{s20041226,
title = {CLRS: Continual Learning Benchmark for Remote Sensing Image Scene Classification},
author = {Li, Haifeng and Jiang, Hao and Gu, Xin and Peng, Jian and Li, Wenbo and Hong, Liang and Tao, Chao},
year = 2020,
journal = {Sensors}… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/CLRS.x_dataset_0409154
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0409154.yanxiaroberta-datax_dataset_040484
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_040484.x_dataset_0403203
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_0403203.roberta-tokenized-dataAID_MultiLabel
Dataset Card for "AID_MultiLabel"
Licensing Information
CC0: Public Domain
Citation Information
Imagery:
AID: A benchmark data set for performance evaluation of aerial scene classification
Multilabels:
Relation Network for Multi-label Aerial Image Classification
@article{xia2017aid,
title = {AID: A benchmark data set for performance evaluation of aerial scene classification},
author = {Xia, Gui-Song and Hu, Jingwen and Hu, Fan and Shi, Baoguang… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/AID_MultiLabel.wikitext-103-raw-v1_sents_min_len10_max_len30_princeton-nlp_sup-simcse-roberta-largezerobench_no_answersGID
Dataset Card for "GID"
Licensing Information
Public domain.
Citation Information
Land-cover classification with high-resolution remote sensing images using transferable deep models
@article{GID2020,
title = {Land-cover classification with high-resolution remote sensing images using transferable deep models},
author = {Tong, Xin-Yi and Xia, Gui-Song and Lu, Qikai and Shen, Huanfeng and Li, Shengyang and You, Shucheng and Zhang, Liangpei},
year… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/GID.
