soncco/historical-twitter-trends-spanish-locations-2020-2021
Historical Twitter Trends Across 11 Spanish-Speaking Locations (2020–2021) Dataset summary This dataset contains 5,726,446 historical observations of trending topics displayed by TweetDeck across 11 Spanish-speaking countries and territories. The observations cover the period from July 15, 2020, to September 22, 2021. The collection software was configured to run every 15 minutes and recorded the trend text, its position in the TweetDeck trends column, the… See the full description on the dataset page: https://huggingface.co/datasets/soncco/historical-twitter-trends-spanish-locations-2020-2021.
Historical Twitter Trends Across 11 Spanish-Speaking Locations (2020–2021)
Dataset summary
This dataset contains 5,726,446 historical observations of trending topics displayed by TweetDeck across 11 Spanish-speaking countries and territories. The observations cover the period from July 15, 2020, to September 22, 2021.
The collection software was configured to run every 15 minutes and recorded the trend text, its position in the TweetDeck trends column, the approximate tweet volume displayed by the platform, the selected location, and the observation time.
The dataset was originally collected as the first operational stage of a broader research strategy: identify relevant trends, retrieve related tweets, follow links to news sources, and subsequently support the construction of a Spanish-language corpus for fake-news research.
This is not a fake-news dataset. It does not contain tweets, user accounts, news articles, URLs, or veracity labels. It is a longitudinal dataset of platform-generated trending-topic observations.
Resumen en español
Este dataset contiene 5,726,446 observaciones históricas de tendencias mostradas por TweetDeck en 11 países y territorios hispanohablantes, recopiladas entre el 15 de julio de 2020 y el 22 de septiembre de 2021.
El recolector fue configurado para ejecutarse cada 15 minutos. Cada registro contiene la ubicación, el texto original de la tendencia, su posición, el volumen aproximado mostrado por TweetDeck y la fecha de observación. El dataset no contiene tuits, usuarios, noticias ni etiquetas de veracidad; por tanto, no debe interpretarse como un corpus de noticias falsas.
Dataset details
- Curator: Braulio Andrés Soncco Pimentel
- Version: 1.0.0
- License: Open Data Commons Attribution License (ODC-By) 1.0
- Collection source: TweetDeck trends columns
- Scraper source code: github.com/soncco/trends-scraper
- Dataset repository: huggingface.co/datasets/soncco/historical-twitter-trends-spanish-locations-2020-2021
- Related publication: Fake News in Spanish: Towards the Building of a Corpus Based on Twitter
Relationship with the related publication
The related paper proposed using Twitter as a mediator for identifying relevant topics and news sources during the construction of a Spanish-language fake-news corpus. The present dataset corresponds to the historical trend-collection stage developed after that initial work.
The later stages proposed in the paper—retrieving related tweets, identifying linked news sources, and assigning veracity labels—were not performed on this dataset. Its publication is intended to make the collected trend history available for other research questions and for researchers who may wish to continue or adapt that strategy.
Geographic and temporal coverage
The complete UTC observation range is:
2020-07-15 16:06:08 UTC to 2021-09-22 05:37:01 UTCDataset statistics
Unique trend strings are counted exactly and case-sensitively. For example, differently capitalized versions are counted separately.
The 15-minute interval was the configured collection frequency, not a guarantee of continuous coverage. Interruptions occurred during the collection period. The longest observed interruption was approximately 13 days.
Collection process
The public trends-scraper performed the following process:
- Started an authenticated browser session with Selenium.
- Opened TweetDeck with trends columns configured for the selected locations.
- Read the displayed trend elements for every configured column.
- Recorded the trend text and its position in the displayed order.
- Parsed the approximate tweet volume when TweetDeck displayed it.
- Generated an observation timestamp using the fixed UTC offset configured for the location.
- Stored the observations in a MySQL table.
- Repeated the process on a nominal 900-second interval.
The observations were obtained from the TweetDeck user interface rather than from a historical trends API.
Data processing
The original MySQL table contained the following columns:
id, location, hashtag, tweets_counter, position, trend_dateThe published dataset applies these transformations:
Missing tweet volumes
The original scraper stored 0 when TweetDeck did not display a tweet-volume element. Consequently, those zeros did not represent a verified volume of zero tweets. During publication they were converted to null.
When available, values such as 11.4K Tweets or 2.1M Tweets were converted by the scraper into approximate integer counts. tweet_volume should therefore be treated as an approximate platform-provided display value, not as an independently measured exact count.
Timestamp normalization
The original database stored a timezone-naive datetime, but the scraper generated that value using a fixed offset configured for each location. The published field observed_at_utc reverses that transformation:
observed_at_utc = observed_at_local − configured UTC offsetThis reconstruction preserves the collection instant intended by the original software. However, observed_at_local should not always be interpreted as the location's official civil time because fixed offsets do not account for daylight-saving transitions or multiple time zones within a country. Use observed_at_utc for cross-location temporal analysis.
Data structure
The dataset is distributed as Zstandard-compressed Parquet files, organized by location and year:
data/
├── AR/
│ ├── trends-2020.parquet
│ └── trends-2021.parquet
├── CL/
├── CO/
├── EC/
├── ES/
├── GT/
├── MX/
├── PA/
├── PE/
├── PR/
└── VE/The repository also contains manifest.json, which records source counts, partition counts, file sizes, transformations, configured offsets, and validation results.
Data fields
Example record:
{
"source_record_id": 1,
"location": "Argentina",
"country_code": "AR",
"trend_name": "#MatrimonioIgualitario",
"is_hashtag": true,
"tweet_volume": 11400,
"rank": 1,
"observed_at_local": "2020-07-15T13:06:08",
"utc_offset_minutes": -180,
"observed_at_utc": "2020-07-15T16:06:08Z",
"year": 2020,
"month": 7
}Loading the dataset
Install the Hugging Face datasets library:
python -m pip install datasetsComplete dataset
from datasets import load_dataset
dataset = load_dataset(
"soncco/historical-twitter-trends-spanish-locations-2020-2021",
"all",
split="train",
)
print(dataset)
print(dataset[0])One location
from datasets import load_dataset
peru = load_dataset(
"soncco/historical-twitter-trends-spanish-locations-2020-2021",
"peru",
split="train",
)Available configurations:
all, argentina, chile, colombia, ecuador, guatemala, mexico,
panama, peru, puerto_rico, spain, venezuelaStreaming
from datasets import load_dataset
stream = load_dataset(
"soncco/historical-twitter-trends-spanish-locations-2020-2021",
"all",
split="train",
streaming=True,
)
for row in stream.take(5):
print(row)Loading one Parquet file with pandas
import pandas as pd
url = (
"hf://datasets/soncco/"
"historical-twitter-trends-spanish-locations-2020-2021/"
"data/PE/trends-2021.parquet"
)
df = pd.read_parquet(url)Potential uses
The dataset may support research involving:
- temporal evolution and persistence of trending topics;
- comparison of trends across countries and territories;
- cross-location diffusion of topics;
- event and anomaly detection;
- emerging vocabulary and hashtag analysis;
- historical studies of social-media attention during 2020–2021;
- selection of relevant topics for subsequent news or misinformation research;
- construction and evaluation of time-series analysis methods.
Out-of-scope uses
The dataset should not be used by itself to:
- classify news as true or false;
- analyze the sentiment or content of individual tweets;
- identify which users promoted a topic;
- estimate the opinions of the general population;
- infer causal relationships from trend rankings;
- treat
tweet_volumeas an exact measure of all tweets about a topic.
Limitations and biases
- Platform selection bias: Twitter users are not representative of the general population.
- Algorithmic opacity: Twitter/TweetDeck determined which topics appeared and in which order; the ranking algorithm is not documented in this dataset.
- Incomplete temporal coverage: the scraper was scheduled every 15 minutes, but server, network, browser, authentication, or platform interruptions produced gaps.
- Approximate volumes: displayed
KandMvalues were converted to integers and should not be treated as exact counts. - Missing volumes: 13.76% of observations have no reported tweet volume.
- Repeated observations: a trend can appear in consecutive snapshots. These are expected longitudinal observations, not necessarily erroneous duplicates.
- Unlabeled language: the locations are primarily Spanish-speaking, but individual trend strings may contain other languages, names, abbreviations, or platform-specific expressions.
- Fixed timezone offsets: configured offsets do not fully represent daylight-saving changes or all time zones within a country.
- No tweet-level context: no posts, users, URLs, interactions, or semantic context are included.
- Historical platform behavior: the dataset describes Twitter/TweetDeck during 2020–2021 and should not be assumed to represent the present-day behavior of X.
Ethical considerations
The dataset contains aggregate trending-topic observations and does not include tweet text, account identifiers, profile data, or direct interaction records. Nevertheless, trend names can contain names of people, organizations, political groups, events, or sensitive topics.
Researchers should avoid using aggregate trends to infer sensitive attributes about individuals or populations and should consider the social and political context of each location when interpreting results.
License
This database is made available under the Open Data Commons Attribution License (ODC-By) 1.0.
The license applies to the database compilation, organization, documentation, and derived metadata provided by the dataset curator. No ownership is claimed over third-party trademarks or platform-generated content.
Twitter, TweetDeck, and X are trademarks of their respective owners. This dataset is not affiliated with, authorized by, or endorsed by X Corp. Users are responsible for ensuring that their use complies with applicable laws and relevant third-party terms.
Citation
If you use this dataset, please cite both the dataset and, when relevant, the related publication.
Dataset
@dataset{soncco2026historicaltrends,
author = {Soncco Pimentel, Braulio Andrés},
title = {Historical Twitter Trends Across 11 Spanish-Speaking
Locations (2020--2021)},
year = {2026},
publisher = {Hugging Face},
version = {1.0.0},
url = {https://huggingface.co/datasets/soncco/historical-twitter-trends-spanish-locations-2020-2021}
}Related publication
@incollection{soncco2020fakenews,
author = {Soncco Pimentel, Braulio Andres and
Portugal, Roxana L. Q.},
title = {Fake News in Spanish: Towards the Building of a Corpus
Based on Twitter},
booktitle = {Information Management and Big Data},
pages = {333--339},
year = {2020},
publisher = {Springer},
doi = {10.1007/978-3-030-46140-9_32}
}Changelog
Version 1.0.0
- Initial public release.
- 5,726,446 validated observations.
- 11 location configurations.
- Parquet files partitioned by location and year.
- Missing displayed volumes represented as
null. - UTC observation timestamps reconstructed from the original fixed offsets.
