CoolFace
Datasetpublic

LadyMia/x_dataset_36129

Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_36129.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes172downloads
Dataset Card

Bittensor Subnet 13 X (Twitter) Dataset

<center> <img src="https://huggingface.co/datasets/macrocosm-os/images/resolve/main/bittensor.png" alt="Data-universe: The finest collection of social media data the web has to offer"> </center>

<center> <img src="https://huggingface.co/datasets/macrocosm-os/images/resolve/main/macrocosmos-black.png" alt="Data-universe: The finest collection of social media data the web has to offer"> </center>

Dataset Description

  • —Repository: LadyMia/xdataset36129
  • —Subnet: Bittensor Subnet 13
  • —Miner Hotkey: 5DSGWeVGSsuwCcZMbVweDG8MST2JMPMuPDHAeDKBeduLtL91

Dataset Summary

This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository.

Supported Tasks

The versatility of this dataset allows researchers and data scientists to explore various aspects of social media dynamics and develop innovative applications. Users are encouraged to leverage this data creatively for their specific research or business needs. For example:

  • —Sentiment Analysis
  • —Trend Detection
  • —Content Analysis
  • —User Behavior Modeling

Languages

Primary language: Datasets are mostly English, but can be multilingual due to decentralized ways of creation.

Dataset Structure

Data Instances

Each instance represents a single tweet with the following fields:

Data Fields

  • —text (string): The main content of the tweet.
  • —label (string): Sentiment or topic category of the tweet.
  • —tweet_hashtags (list): A list of hashtags used in the tweet. May be empty if no hashtags are present.
  • —datetime (string): The date when the tweet was posted.
  • —username_encoded (string): An encoded version of the username to maintain user privacy.
  • —url_encoded (string): An encoded version of any URLs included in the tweet. May be empty if no URLs are present.

Data Splits

This dataset is continuously updated and does not have fixed splits. Users should create their own splits based on their requirements and the data's timestamp.

Dataset Creation

Source Data

Data is collected from public tweets on X (Twitter), adhering to the platform's terms of service and API usage guidelines.

Personal and Sensitive Information

All usernames and URLs are encoded to protect user privacy. The dataset does not intentionally include personal or sensitive information.

Considerations for Using the Data

Social Impact and Biases

Users should be aware of potential biases inherent in X (Twitter) data, including demographic and content biases. This dataset reflects the content and opinions expressed on X and should not be considered a representative sample of the general population.

Limitations

  • —Data quality may vary due to the decentralized nature of collection and preprocessing.
  • —The dataset may contain noise, spam, or irrelevant content typical of social media platforms.
  • —Temporal biases may exist due to real-time collection methods.
  • —The dataset is limited to public tweets and does not include private accounts or direct messages.
  • —Not all tweets contain hashtags or URLs.

Additional Information

Licensing Information

The dataset is released under the MIT license. The use of this dataset is also subject to X Terms of Use.

Citation Information

If you use this dataset in your research, please cite it as follows:

@misc{LadyMia2025datauniversex_dataset_36129,
        title={The Data Universe Datasets: The finest collection of social media data the web has to offer},
        author={LadyMia},
        year={2025},
        url={https://huggingface.co/datasets/LadyMia/x_dataset_36129},
        }

Contributions

To report issues or contribute to the dataset, please contact the miner or use the Bittensor Subnet 13 governance mechanisms.

Dataset Statistics

[This section is automatically updated]

  • —Total Instances: 47602926
  • —Date Range: 2025-01-21T00:00:00Z to 2025-02-13T00:00:00Z
  • —Last Updated: 2025-02-18T19:13:12Z

Data Distribution

  • —Tweets with hashtags: 47.25%
  • —Tweets without hashtags: 52.75%

Top 10 Hashtags

For full statistics, please refer to the stats.json file in the repository.

RankTopicTotal CountPercentage
1NULL2510869552.75%
2#riyadh3755030.79%
3#zelena2569080.54%
4#tiktok2202380.46%
5#bbb251631960.34%
6#jhopeatgaladespiècesjaunes1437780.30%
7#ad1258570.26%
8#trump712640.15%
9#bbmzansi710240.15%
10#granhermano675470.14%

Update History

DateNew InstancesTotal Instances
2025-01-27T06:58:39Z47323514732351
2025-01-30T19:02:22Z950027814232629
2025-02-03T07:05:36Z883699623069625
2025-02-06T19:08:56Z781431730883942
2025-02-10T07:12:46Z726348238147424
2025-02-13T19:16:31Z813759346285017
2025-02-18T04:12:00Z64145846926475
2025-02-18T19:13:12Z67645147602926