datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text_coordinates_regions
Dataset Card for Multilingual Geo-Tagged Social Media Posts (by 123 world regions)
Dataset Summary
The "Regions" dataset is a multilingual corpus that encompasses textual data from the 123 most populated regions worldwide, with each region's data organized into separate .json files. This dataset consists of approximately 500,000 text samples, each paired with its geographic coordinates.
Key Features:
Textual Data: The dataset contains 500,000 text samples.… See the full description on the dataset page: https://huggingface.co/datasets/yachay/text_coordinates_regions.text_coordinates_seasons
Dataset Card for Geo-Tagged Social Media Posts with Timestamps
Dataset Summary
The "Seasons" dataset is a collection of over 600,000 social media posts spanning 12 months and encompassing 15 distinct time zones. It focuses on six countries: Cuba, Iran, Russia, North Korea, Syria, and Venezuela, with each post containing textual content, timestamps, and geographical coordinates. The dataset's primary objective is to investigate the correlation between the timing of posts… See the full description on the dataset page: https://huggingface.co/datasets/yachay/text_coordinates_seasons.spatial-awareness-coordinates
