datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
x_dataset_44311
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_44311.lsqb-sf1x_dataset_7480
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_7480.x_dataset_14253
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_14253.LaDe-D
1. About Dataset
LaDe is a publicly available last-mile delivery dataset with millions of packages from industry.
It has three unique characteristics: (1) Large-scale. It involves 10,677k packages of 21k couriers over 6 months of real-world operation.
(2) Comprehensive information, it offers original package information, such as its location and time requirements, as well as task-event information, which records when and where the courier is while events such as task-accept and… See the full description on the dataset page: https://huggingface.co/datasets/Cainiao-AI/LaDe-D.x_dataset_40563
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_40563.x_dataset_21893
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_21893.LaDe-P
1. About Dataset
LaDe is a publicly available last-mile delivery dataset with millions of packages from industry.
It has three unique characteristics: (1) Large-scale. It involves 10,677k packages of 21k couriers over 6 months of real-world operation.
(2) Comprehensive information, it offers original package information, such as its location and time requirements, as well as task-event information, which records when and where the courier is while events such as task-accept and… See the full description on the dataset page: https://huggingface.co/datasets/Cainiao-AI/LaDe-P.x_dataset_51674
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_51674.surdoc-liticos
SURDOC — material lítico arqueológico de Chile
Metadatos e imágenes de 8,737 fichas del Sistema Único de Registro
Documental del Patrimonio Cultural de Chile (Surdoc.cl).
Captura: 2026-08-31. Es una captura fechada, no "todo lo que hay en
Surdoc": la fuente cambia. Un corpus comparable perdió 16 registros en 18 días
entre dos capturas del mismo filtro. No asumas estabilidad entre versiones.
Marco institucional
Este dataset se publica en el marco del Convenio de… See the full description on the dataset page: https://huggingface.co/datasets/LADUC/surdoc-liticos.x_dataset_21318
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_21318.x_dataset_41147
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_41147.indian-english-lady-embeddings-v3x_dataset_2447
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_2447.x_dataset_55395
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_55395.x_dataset_36129
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_36129.x_dataset_41414
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_41414.ldbc-csrLADaS
LADaS: Layout Analysis Dataset with Segmonto
Dataset Details
LADaS, created by the ALMANaCH team-project at Inria,
continued in partnership with other researchers, is a multidocuments diachronic layout analysis
dataset. This dataset includes:
Monographs from the Bibliothèque Nationale de France (17th century - today);
PhD Thesis, in various fields (not only STEM, 20th-21st century);
Selling Catalogs (for manuscripts and art pieces), in various fields (18th-20th… See the full description on the dataset page: https://huggingface.co/datasets/almanach/LADaS.x_dataset_17682
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_17682.imnet1k_ladybug_ladybeetle_lady_beetle_ladybird_ladybird_beetlesmall-kgs
Small Knowledge Graphs (small-kgs)
A collection of small knowledge graphs in multiple formats for graph ML research and development.
Dataset Structure
This dataset contains knowledge graphs in three formats:
graph-std
Compressed Sparse Row (CSR) graph format with parquet files containing:
Node data (persons, organizations, events, concepts, places, knowledge graphs)
Edge indices and indirection pointers
Node/edge type mappings
Schema definition… See the full description on the dataset page: https://huggingface.co/datasets/ladybugdb/small-kgs.x_dataset_12949
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_12949.GameBoyGhost-LADX
GameBoyGhost-LADX
A dataset of 16,777,216 emulator actions across 8,133 episodes, collected
for the GameBoyGhost research project in Link's Awakening DX. Source and
experiment documentation: GameBoyGhost.
Contents
raw/: 64 Parquet shards, retaining full episodes and all committed rows.
Each Parquet row group contains one complete source episode.
metadata/episodes.parquet: episode provenance, source checksums, shard and
row-group lookup, original curation… See the full description on the dataset page: https://huggingface.co/datasets/foxmedik/GameBoyGhost-LADX.omnigen2-azimuth-ladder-anny-20260901
omnigen2-azimuth-ladder-anny-20260901
An image-edit ladder in the EditScore dataset shape: a candidate measured against a baseline
on the same prompts, one row per (source, edited, instruction) with the per-pair
measurement beside the images. The baseline is OmniGen2; the candidate is the same model
after a camera-control LoRA.
This ladder has no EditScore score. The runs measured recovered azimuth — where the
body actually faces in the generated view — not EditScore's pf / sc /… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/omnigen2-azimuth-ladder-anny-20260901.x_dataset_24095
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_24095.delta_learning_model_ladderLAD-training-1m-512x_dataset_47139
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_47139.tatoeba-ladino
Ladino sentences with English, Turkish and Spanish translations from Tatoeba platform (http://tatoeba.org)
Corpus also includes English-Ladino pairs converted to Turkish-Ladino and Spanish-Ladino using machine translation.
License: CC-BY
Curation of this dataset was done as part of project "Judeo-Spanish: Connecting the two ends of the Mediterranean" carried out by Col·lectivaT and Sephardic Center of Istanbul within the framework of the “Grant Scheme for Common Cultural Heritage:… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/tatoeba-ladino.
