CoolFace
Datasetpublic

malicious546/x_dataset_128

Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine… See the full description on the dataset page: https://huggingface.co/datasets/malicious546/x_dataset_128.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes19downloads
Dataset Card

Bittensor Subnet 13 X (Twitter) Dataset

<center> <img src="https://huggingface.co/datasets/macrocosm-os/images/resolve/main/bittensor.png" alt="Data-universe: The finest collection of social media data the web has to offer"> </center>

<center> <img src="https://huggingface.co/datasets/macrocosm-os/images/resolve/main/macrocosmos-black.png" alt="Data-universe: The finest collection of social media data the web has to offer"> </center>

Dataset Description

  • Repository: malicious546/xdataset128
  • Subnet: Bittensor Subnet 13
  • Miner Hotkey: 5EcmufhjLXd3bh2ZCdF8XS3y6hkihtG4yhTvw81ieui45iLi

Miner Data Compliance Agreement

In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.

Dataset Summary

This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository.

Supported Tasks

The versatility of this dataset allows researchers and data scientists to explore various aspects of social media dynamics and develop innovative applications. Users are encouraged to leverage this data creatively for their specific research or business needs. For example:

  • Sentiment Analysis
  • Trend Detection
  • Content Analysis
  • User Behavior Modeling

Languages

Primary language: Datasets are mostly English, but can be multilingual due to decentralized ways of creation.

Dataset Structure

Data Instances

Each instance represents a single tweet with the following fields:

Data Fields

  • text (string): The main content of the tweet.
  • label (string): Sentiment or topic category of the tweet.
  • tweet_hashtags (list): A list of hashtags used in the tweet. May be empty if no hashtags are present.
  • datetime (string): The date when the tweet was posted.
  • username_encoded (string): An encoded version of the username to maintain user privacy.
  • url_encoded (string): An encoded version of any URLs included in the tweet. May be empty if no URLs are present.

Data Splits

This dataset is continuously updated and does not have fixed splits. Users should create their own splits based on their requirements and the data's timestamp.

Dataset Creation

Source Data

Data is collected from public tweets on X (Twitter), adhering to the platform's terms of service and API usage guidelines.

Personal and Sensitive Information

All usernames and URLs are encoded to protect user privacy. The dataset does not intentionally include personal or sensitive information.

Considerations for Using the Data

Social Impact and Biases

Users should be aware of potential biases inherent in X (Twitter) data, including demographic and content biases. This dataset reflects the content and opinions expressed on X and should not be considered a representative sample of the general population.

Limitations

  • Data quality may vary due to the decentralized nature of collection and preprocessing.
  • The dataset may contain noise, spam, or irrelevant content typical of social media platforms.
  • Temporal biases may exist due to real-time collection methods.
  • The dataset is limited to public tweets and does not include private accounts or direct messages.
  • Not all tweets contain hashtags or URLs.

Additional Information

Licensing Information

The dataset is released under the MIT license. The use of this dataset is also subject to X Terms of Use.

Citation Information

If you use this dataset in your research, please cite it as follows:

@misc{malicious5462025datauniversex_dataset_128,
        title={The Data Universe Datasets: The finest collection of social media data the web has to offer},
        author={malicious546},
        year={2025},
        url={https://huggingface.co/datasets/malicious546/x_dataset_128},
        }

Contributions

To report issues or contribute to the dataset, please contact the miner or use the Bittensor Subnet 13 governance mechanisms.

Dataset Statistics

[This section is automatically updated]

  • Total Instances: 1300
  • Date Range: 2025-07-08T00:00:00Z to 2025-07-19T00:00:00Z
  • Last Updated: 2025-07-31T10:17:03Z

Data Distribution

  • Tweets with hashtags: 100.00%
  • Tweets without hashtags: 0.00%

Top 10 Hashtags

For full statistics, please refer to the stats.json file in the repository.

RankTopicTotal CountPercentage
1#bitcoin13410.31%
2#btc836.38%
3#crypto715.46%
4#swapnox574.38%
5#trump513.92%
6#israel272.08%
7#defi272.08%
8#bitcoiner231.77%
9#ukraine231.77%
10#giveaway201.54%

Update History

DateNew InstancesTotal Instances
2025-07-22T09:10:00Z100100
2025-07-23T03:15:22Z100200
2025-07-23T21:31:32Z100300
2025-07-24T15:40:30Z100400
2025-07-25T09:43:54Z100500
2025-07-26T03:48:02Z100600
2025-07-26T21:52:09Z100700
2025-07-27T15:56:17Z100800
2025-07-28T09:59:51Z100900
2025-07-29T04:03:04Z1001000
2025-07-29T22:07:54Z1001100
2025-07-30T16:13:16Z1001200
2025-07-31T10:17:03Z1001300