CoolFace
Datasetpublic

soncco/historical-twitter-trends-spanish-locations-2020-2021

Historical Twitter Trends Across 11 Spanish-Speaking Locations (2020–2021) Dataset summary This dataset contains 5,726,446 historical observations of trending topics displayed by TweetDeck across 11 Spanish-speaking countries and territories. The observations cover the period from July 15, 2020, to September 22, 2021. The collection software was configured to run every 15 minutes and recorded the trend text, its position in the TweetDeck trends column, the… See the full description on the dataset page: https://huggingface.co/datasets/soncco/historical-twitter-trends-spanish-locations-2020-2021.

sourceHugging Faceodc-byupdated 5d agoView on Hugging Face
1likes88downloads
Dataset Card

Historical Twitter Trends Across 11 Spanish-Speaking Locations (2020–2021)

Dataset summary

This dataset contains 5,726,446 historical observations of trending topics displayed by TweetDeck across 11 Spanish-speaking countries and territories. The observations cover the period from July 15, 2020, to September 22, 2021.

The collection software was configured to run every 15 minutes and recorded the trend text, its position in the TweetDeck trends column, the approximate tweet volume displayed by the platform, the selected location, and the observation time.

The dataset was originally collected as the first operational stage of a broader research strategy: identify relevant trends, retrieve related tweets, follow links to news sources, and subsequently support the construction of a Spanish-language corpus for fake-news research.

This is not a fake-news dataset. It does not contain tweets, user accounts, news articles, URLs, or veracity labels. It is a longitudinal dataset of platform-generated trending-topic observations.

Resumen en español

Este dataset contiene 5,726,446 observaciones históricas de tendencias mostradas por TweetDeck en 11 países y territorios hispanohablantes, recopiladas entre el 15 de julio de 2020 y el 22 de septiembre de 2021.

El recolector fue configurado para ejecutarse cada 15 minutos. Cada registro contiene la ubicación, el texto original de la tendencia, su posición, el volumen aproximado mostrado por TweetDeck y la fecha de observación. El dataset no contiene tuits, usuarios, noticias ni etiquetas de veracidad; por tanto, no debe interpretarse como un corpus de noticias falsas.

Dataset details

Relationship with the related publication

The related paper proposed using Twitter as a mediator for identifying relevant topics and news sources during the construction of a Spanish-language fake-news corpus. The present dataset corresponds to the historical trend-collection stage developed after that initial work.

The later stages proposed in the paper—retrieving related tweets, identifying linked news sources, and assigning veracity labels—were not performed on this dataset. Its publication is intended to make the collected trend history available for other research questions and for researchers who may wish to continue or adapt that strategy.

Geographic and temporal coverage

LocationISO codeRowsConfigured UTC offset
ArgentinaAR537,296UTC−03:00
ChileCL521,704UTC−04:00
ColombiaCO523,477UTC−05:00
EcuadorEC528,161UTC−05:00
GuatemalaGT497,383UTC−06:00
MexicoMX527,774UTC−05:00
PanamaPA501,321UTC−05:00
PeruPE520,971UTC−05:00
Puerto RicoPR536,938UTC−04:00
SpainES526,665UTC+02:00
VenezuelaVE504,756UTC−04:00
Total5,726,446

The complete UTC observation range is:

text
2020-07-15 16:06:08 UTC to 2021-09-22 05:37:01 UTC

Dataset statistics

MetricValue
Total observations5,726,446
Locations11
Parquet files22
Exact unique trend strings167,491
Observations beginning with #1,328,835 (23.21%)
Observations without reported tweet volume787,685 (13.76%)
Location-level snapshots identified by rank = 1291,329
Median interval between snapshots within a location15 minutes
Observed rank range1–67

Unique trend strings are counted exactly and case-sensitively. For example, differently capitalized versions are counted separately.

The 15-minute interval was the configured collection frequency, not a guarantee of continuous coverage. Interruptions occurred during the collection period. The longest observed interruption was approximately 13 days.

Collection process

The public trends-scraper performed the following process:

  1. 1.Started an authenticated browser session with Selenium.
  2. 2.Opened TweetDeck with trends columns configured for the selected locations.
  3. 3.Read the displayed trend elements for every configured column.
  4. 4.Recorded the trend text and its position in the displayed order.
  5. 5.Parsed the approximate tweet volume when TweetDeck displayed it.
  6. 6.Generated an observation timestamp using the fixed UTC offset configured for the location.
  7. 7.Stored the observations in a MySQL table.
  8. 8.Repeated the process on a nominal 900-second interval.

The observations were obtained from the TweetDeck user interface rather than from a historical trends API.

Data processing

The original MySQL table contained the following columns:

text
id, location, hashtag, tweets_counter, position, trend_date

The published dataset applies these transformations:

Original fieldPublished fieldTransformation
idsource_record_idRenamed; original value preserved
locationlocationPreserved
—country_codeAdded from the location configuration
hashtagtrend_nameRenamed because many trends are not hashtags
—is_hashtagtrue when trend_name begins with #
tweets_countertweet_volumeRenamed; zero converted to null
positionrankRenamed
trend_dateobserved_at_localOriginal stored timestamp preserved
—utc_offset_minutesAdded from the scraper configuration
—observed_at_utcReconstructed from the configured fixed offset
—year, monthDerived from observed_at_utc

Missing tweet volumes

The original scraper stored 0 when TweetDeck did not display a tweet-volume element. Consequently, those zeros did not represent a verified volume of zero tweets. During publication they were converted to null.

When available, values such as 11.4K Tweets or 2.1M Tweets were converted by the scraper into approximate integer counts. tweet_volume should therefore be treated as an approximate platform-provided display value, not as an independently measured exact count.

Timestamp normalization

The original database stored a timezone-naive datetime, but the scraper generated that value using a fixed offset configured for each location. The published field observed_at_utc reverses that transformation:

text
observed_at_utc = observed_at_local − configured UTC offset

This reconstruction preserves the collection instant intended by the original software. However, observed_at_local should not always be interpreted as the location's official civil time because fixed offsets do not account for daylight-saving transitions or multiple time zones within a country. Use observed_at_utc for cross-location temporal analysis.

Data structure

The dataset is distributed as Zstandard-compressed Parquet files, organized by location and year:

text
data/
├── AR/
│   ├── trends-2020.parquet
│   └── trends-2021.parquet
├── CL/
├── CO/
├── EC/
├── ES/
├── GT/
├── MX/
├── PA/
├── PE/
├── PR/
└── VE/

The repository also contains manifest.json, which records source counts, partition counts, file sizes, transformations, configured offsets, and validation results.

Data fields

FieldTypeDescription
source_record_idint64Original auto-increment identifier from MySQL
locationstringTweetDeck trends location
country_codestringISO 3166-1 alpha-2 country or territory code
trend_namestringTrend text exactly as observed
is_hashtagbooleanWhether trend_name begins with #
tweet_volumenullable int64Approximate volume displayed by TweetDeck, or null when unavailable
rankint16Position in the TweetDeck trends column
observed_at_localtimestamp[us]Original timestamp produced using the configured fixed offset
utc_offset_minutesint16Fixed offset used by the scraper, expressed in minutes
observed_at_utctimestamp[us, UTC]Reconstructed UTC observation time
yearint16UTC observation year
monthint8UTC observation month

Example record:

json
{
  "source_record_id": 1,
  "location": "Argentina",
  "country_code": "AR",
  "trend_name": "#MatrimonioIgualitario",
  "is_hashtag": true,
  "tweet_volume": 11400,
  "rank": 1,
  "observed_at_local": "2020-07-15T13:06:08",
  "utc_offset_minutes": -180,
  "observed_at_utc": "2020-07-15T16:06:08Z",
  "year": 2020,
  "month": 7
}

Loading the dataset

Install the Hugging Face datasets library:

bash
python -m pip install datasets

Complete dataset

python
from datasets import load_dataset

dataset = load_dataset(
    "soncco/historical-twitter-trends-spanish-locations-2020-2021",
    "all",
    split="train",
)

print(dataset)
print(dataset[0])

One location

python
from datasets import load_dataset

peru = load_dataset(
    "soncco/historical-twitter-trends-spanish-locations-2020-2021",
    "peru",
    split="train",
)

Available configurations:

text
all, argentina, chile, colombia, ecuador, guatemala, mexico,
panama, peru, puerto_rico, spain, venezuela

Streaming

python
from datasets import load_dataset

stream = load_dataset(
    "soncco/historical-twitter-trends-spanish-locations-2020-2021",
    "all",
    split="train",
    streaming=True,
)

for row in stream.take(5):
    print(row)

Loading one Parquet file with pandas

python
import pandas as pd

url = (
    "hf://datasets/soncco/"
    "historical-twitter-trends-spanish-locations-2020-2021/"
    "data/PE/trends-2021.parquet"
)

df = pd.read_parquet(url)

Potential uses

The dataset may support research involving:

  • —temporal evolution and persistence of trending topics;
  • —comparison of trends across countries and territories;
  • —cross-location diffusion of topics;
  • —event and anomaly detection;
  • —emerging vocabulary and hashtag analysis;
  • —historical studies of social-media attention during 2020–2021;
  • —selection of relevant topics for subsequent news or misinformation research;
  • —construction and evaluation of time-series analysis methods.

Out-of-scope uses

The dataset should not be used by itself to:

  • —classify news as true or false;
  • —analyze the sentiment or content of individual tweets;
  • —identify which users promoted a topic;
  • —estimate the opinions of the general population;
  • —infer causal relationships from trend rankings;
  • —treat tweet_volume as an exact measure of all tweets about a topic.

Limitations and biases

  1. 1.Platform selection bias: Twitter users are not representative of the general population.
  2. 2.Algorithmic opacity: Twitter/TweetDeck determined which topics appeared and in which order; the ranking algorithm is not documented in this dataset.
  3. 3.Incomplete temporal coverage: the scraper was scheduled every 15 minutes, but server, network, browser, authentication, or platform interruptions produced gaps.
  4. 4.Approximate volumes: displayed K and M values were converted to integers and should not be treated as exact counts.
  5. 5.Missing volumes: 13.76% of observations have no reported tweet volume.
  6. 6.Repeated observations: a trend can appear in consecutive snapshots. These are expected longitudinal observations, not necessarily erroneous duplicates.
  7. 7.Unlabeled language: the locations are primarily Spanish-speaking, but individual trend strings may contain other languages, names, abbreviations, or platform-specific expressions.
  8. 8.Fixed timezone offsets: configured offsets do not fully represent daylight-saving changes or all time zones within a country.
  9. 9.No tweet-level context: no posts, users, URLs, interactions, or semantic context are included.
  10. 10.Historical platform behavior: the dataset describes Twitter/TweetDeck during 2020–2021 and should not be assumed to represent the present-day behavior of X.

Ethical considerations

The dataset contains aggregate trending-topic observations and does not include tweet text, account identifiers, profile data, or direct interaction records. Nevertheless, trend names can contain names of people, organizations, political groups, events, or sensitive topics.

Researchers should avoid using aggregate trends to infer sensitive attributes about individuals or populations and should consider the social and political context of each location when interpreting results.

License

This database is made available under the Open Data Commons Attribution License (ODC-By) 1.0.

The license applies to the database compilation, organization, documentation, and derived metadata provided by the dataset curator. No ownership is claimed over third-party trademarks or platform-generated content.

Twitter, TweetDeck, and X are trademarks of their respective owners. This dataset is not affiliated with, authorized by, or endorsed by X Corp. Users are responsible for ensuring that their use complies with applicable laws and relevant third-party terms.

Citation

If you use this dataset, please cite both the dataset and, when relevant, the related publication.

Dataset

bibtex
@dataset{soncco2026historicaltrends,
  author       = {Soncco Pimentel, Braulio Andrés},
  title        = {Historical Twitter Trends Across 11 Spanish-Speaking
                  Locations (2020--2021)},
  year         = {2026},
  publisher    = {Hugging Face},
  version      = {1.0.0},
  url          = {https://huggingface.co/datasets/soncco/historical-twitter-trends-spanish-locations-2020-2021}
}

Related publication

bibtex
@incollection{soncco2020fakenews,
  author    = {Soncco Pimentel, Braulio Andres and
               Portugal, Roxana L. Q.},
  title     = {Fake News in Spanish: Towards the Building of a Corpus
               Based on Twitter},
  booktitle = {Information Management and Big Data},
  pages     = {333--339},
  year      = {2020},
  publisher = {Springer},
  doi       = {10.1007/978-3-030-46140-9_32}
}

Changelog

Version 1.0.0

  • —Initial public release.
  • —5,726,446 validated observations.
  • —11 location configurations.
  • —Parquet files partitioned by location and year.
  • —Missing displayed volumes represented as null.
  • —UTC observation timestamps reconstructed from the original fixed offsets.