CoolFace
Datasetpublic

robert-1111/x_dataset_040484

Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_040484.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes246downloads
Dataset Card

Bittensor Subnet 13 X (Twitter) Dataset

<center> <img src="https://huggingface.co/datasets/macrocosm-os/images/resolve/main/bittensor.png" alt="Data-universe: The finest collection of social media data the web has to offer"> </center>

<center> <img src="https://huggingface.co/datasets/macrocosm-os/images/resolve/main/macrocosmos-black.png" alt="Data-universe: The finest collection of social media data the web has to offer"> </center>

Dataset Description

  • Repository: robert-1111/xdataset040484
  • Subnet: Bittensor Subnet 13
  • Miner Hotkey: 5DMEDsCn1rgczUQsz9198S1Sed9MxxAWuC4hkAdHw2ieDuxZ

Miner Data Compliance Agreement

In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.

Dataset Summary

This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository.

Supported Tasks

The versatility of this dataset allows researchers and data scientists to explore various aspects of social media dynamics and develop innovative applications. Users are encouraged to leverage this data creatively for their specific research or business needs. For example:

  • Sentiment Analysis
  • Trend Detection
  • Content Analysis
  • User Behavior Modeling

Languages

Primary language: Datasets are mostly English, but can be multilingual due to decentralized ways of creation.

Dataset Structure

Data Instances

Each instance represents a single tweet with the following fields:

Data Fields

  • text (string): The main content of the tweet.
  • label (string): Sentiment or topic category of the tweet.
  • tweet_hashtags (list): A list of hashtags used in the tweet. May be empty if no hashtags are present.
  • datetime (string): The date when the tweet was posted.
  • username_encoded (string): An encoded version of the username to maintain user privacy.
  • url_encoded (string): An encoded version of any URLs included in the tweet. May be empty if no URLs are present.

Data Splits

This dataset is continuously updated and does not have fixed splits. Users should create their own splits based on their requirements and the data's timestamp.

Dataset Creation

Source Data

Data is collected from public tweets on X (Twitter), adhering to the platform's terms of service and API usage guidelines.

Personal and Sensitive Information

All usernames and URLs are encoded to protect user privacy. The dataset does not intentionally include personal or sensitive information.

Considerations for Using the Data

Social Impact and Biases

Users should be aware of potential biases inherent in X (Twitter) data, including demographic and content biases. This dataset reflects the content and opinions expressed on X and should not be considered a representative sample of the general population.

Limitations

  • Data quality may vary due to the decentralized nature of collection and preprocessing.
  • The dataset may contain noise, spam, or irrelevant content typical of social media platforms.
  • Temporal biases may exist due to real-time collection methods.
  • The dataset is limited to public tweets and does not include private accounts or direct messages.
  • Not all tweets contain hashtags or URLs.

Additional Information

Licensing Information

The dataset is released under the MIT license. The use of this dataset is also subject to X Terms of Use.

Citation Information

If you use this dataset in your research, please cite it as follows:

@misc{robert-11112025datauniversex_dataset_040484,
        title={The Data Universe Datasets: The finest collection of social media data the web has to offer},
        author={robert-1111},
        year={2025},
        url={https://huggingface.co/datasets/robert-1111/x_dataset_040484},
        }

Contributions

To report issues or contribute to the dataset, please contact the miner or use the Bittensor Subnet 13 governance mechanisms.

Dataset Statistics

[This section is automatically updated]

  • Total Instances: 3404052
  • Date Range: 2025-01-02T00:00:00Z to 2025-07-20T00:00:00Z
  • Last Updated: 2025-07-30T01:40:28Z

Data Distribution

  • Tweets with hashtags: 3.06%
  • Tweets without hashtags: 96.94%

Top 10 Hashtags

For full statistics, please refer to the stats.json file in the repository.

RankTopicTotal CountPercentage
1NULL114964691.70%
2#箱根駅伝81470.65%
3#thameposeriesep976050.61%
4#tiktok50320.40%
5#zelena48780.39%
6#smackdown48440.39%
7#कबीरपरमेश्वरनिर्वाण_दिवस48430.39%
8#ad35330.28%
9#delhielectionresults34760.28%
10#箱根駅伝202531640.25%

Update History

DateNew InstancesTotal Instances
2025-01-25T07:10:27Z414446414446
2025-01-25T07:10:56Z414446828892
2025-01-25T07:11:27Z4144461243338
2025-01-25T07:11:56Z4535261696864
2025-01-25T07:12:25Z4535262150390
2025-01-25T07:12:56Z4535262603916
2025-02-18T03:39:28Z4718343075750
2025-07-30T01:40:28Z3283023404052