datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GlobalGeoTree
GlobalGeoTree Dataset
GlobalGeoTree is a comprehensive global dataset for tree species classification, comprising 6.3 million geolocated tree occurrences spanning 275 families, 2,734 genera, and 21,001 species across hierarchical taxonomic levels. Each sample is paired with Sentinel-2 image time series and 27 auxiliary environmental variables.
Dataset Structure
This repository contains three main components:
1. GlobalGeoTree-6M
Training dataset with around 6M… See the full description on the dataset page: https://huggingface.co/datasets/yann111/GlobalGeoTree.srtm-3-arc-second-global
SRTM 3 Arc-Second Global
Raw ASCII heightmaps of the Earth's surface labelled according to latitude and longitude.
Mission Description
The Shuttle Radar Topography Mission (SRTM) was flown aboard the space shuttle Endeavour February 11-22, 2000. The National Aeronautics and Space Administration (NASA) and the National Geospatial-Intelligence Agency (NGA) participated in an international project to acquire radar data which were used to create the first near-global set of… See the full description on the dataset page: https://huggingface.co/datasets/mpatrick1991/srtm-3-arc-second-global.GlobalGeoTree
GlobalGeoTree Dataset
GlobalGeoTree is a comprehensive global dataset for tree species classification, comprising 6.3 million geolocated tree occurrences spanning 275 families, 2,734 genera, and 21,001 species across hierarchical taxonomic levels. Each sample is paired with Sentinel-2 image time series and 27 auxiliary environmental variables.
Dataset Structure
This repository contains three main components:
1. GlobalGeoTree-6M
Training dataset with around 6M… See the full description on the dataset page: https://huggingface.co/datasets/cindycui/GlobalGeoTree.NayanaDocs-Global-45k-webdataset
Nayana-DocOCR Global Annotated Dataset
Dataset Description
This is a large-scale multilingual document OCR dataset containing approximately 400GB of images with comprehensive annotations across multiple global languages and English. The dataset is stored in WebDataset format using TAR archives for efficient streaming and processing.
Available Language Subsets
Arabic (ar): Available
German (de): Available
Russian (ru) : Available
Spanish (es): Available
French… See the full description on the dataset page: https://huggingface.co/datasets/Nayana-cognitivelab/NayanaDocs-Global-45k-webdataset.
