CoolFace
Datasetpublic

gk4u/reddit_dataset_132

Bittensor Subnet 13 Reddit Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For… See the full description on the dataset page: https://huggingface.co/datasets/gk4u/reddit_dataset_132.

sourceHugging Facemitupdated 1y agoView on Hugging Face
1likes965downloads
Dataset Card

Bittensor Subnet 13 Reddit Dataset

<center> <img src="https://huggingface.co/datasets/macrocosm-os/images/resolve/main/bittensor.png" alt="Data-universe: The finest collection of social media data the web has to offer"> </center>

<center> <img src="https://huggingface.co/datasets/macrocosm-os/images/resolve/main/macrocosmos-black.png" alt="Data-universe: The finest collection of social media data the web has to offer"> </center>

Dataset Description

  • Repository: gk4u/redditdataset132
  • Subnet: Bittensor Subnet 13
  • Miner Hotkey: 5C7s8hkwvCbB374baZrsALvEwBvbC4DEBT36fKRnhsEwM3vy

Miner Data Compliance Agreement

In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.

Dataset Summary

This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository.

Supported Tasks

The versatility of this dataset allows researchers and data scientists to explore various aspects of social media dynamics and develop innovative applications. Users are encouraged to leverage this data creatively for their specific research or business needs. For example:

  • Sentiment Analysis
  • Topic Modeling
  • Community Analysis
  • Content Categorization

Languages

Primary language: Datasets are mostly English, but can be multilingual due to decentralized ways of creation.

Dataset Structure

Data Instances

Each instance represents a single Reddit post or comment with the following fields:

Data Fields

  • text (string): The main content of the Reddit post or comment.
  • label (string): Sentiment or topic category of the content.
  • dataType (string): Indicates whether the entry is a post or a comment.
  • communityName (string): The name of the subreddit where the content was posted.
  • datetime (string): The date when the content was posted or commented.
  • username_encoded (string): An encoded version of the username to maintain user privacy.
  • url_encoded (string): An encoded version of any URLs included in the content.

Data Splits

This dataset is continuously updated and does not have fixed splits. Users should create their own splits based on their requirements and the data's timestamp.

Dataset Creation

Source Data

Data is collected from public posts and comments on Reddit, adhering to the platform's terms of service and API usage guidelines.

Personal and Sensitive Information

All usernames and URLs are encoded to protect user privacy. The dataset does not intentionally include personal or sensitive information.

Considerations for Using the Data

Social Impact and Biases

Users should be aware of potential biases inherent in Reddit data, including demographic and content biases. This dataset reflects the content and opinions expressed on Reddit and should not be considered a representative sample of the general population.

Limitations

  • Data quality may vary due to the nature of media sources.
  • The dataset may contain noise, spam, or irrelevant content typical of social media platforms.
  • Temporal biases may exist due to real-time collection methods.
  • The dataset is limited to public subreddits and does not include private or restricted communities.

Additional Information

Licensing Information

The dataset is released under the MIT license. The use of this dataset is also subject to Reddit Terms of Use.

Citation Information

If you use this dataset in your research, please cite it as follows:

@misc{gk4u2025datauniversereddit_dataset_132,
        title={The Data Universe Datasets: The finest collection of social media data the web has to offer},
        author={gk4u},
        year={2025},
        url={https://huggingface.co/datasets/gk4u/reddit_dataset_132},
        }

Contributions

To report issues or contribute to the dataset, please contact the miner or use the Bittensor Subnet 13 governance mechanisms.

Dataset Statistics

[This section is automatically updated]

  • Total Instances: 510208581
  • Date Range: 2010-01-18T00:00:00Z to 2025-06-07T00:00:00Z
  • Last Updated: 2025-06-09T14:18:19Z

Data Distribution

  • Posts: 5.06%
  • Comments: 94.94%

Top 10 Subreddits

For full statistics, please refer to the stats.json file in the repository.

RankTopicTotal CountPercentage
1r/moviecritic5168880.10%
2r/GenX4506150.09%
3r/videogames4441220.09%
4r/namenerds4319520.08%
5r/Millennials4311700.08%
6r/KinkTown4171140.08%
7r/dirtyr4r3989420.08%
8r/teenagers3939660.08%
9r/Advice3864950.08%
10r/AITAH3848130.08%

Update History

DateNew InstancesTotal Instances
2025-02-15T09:20:32Z224627658224627658
2025-02-19T06:05:22Z21288558245916216
2025-02-19T09:05:50Z127983246044199
2025-02-19T10:36:03Z34255246078454
2025-02-19T12:06:15Z33760246112214
2025-02-23T04:09:26Z20601216266713430
2025-02-23T10:10:22Z457100267170530
2025-02-23T11:40:33Z34294267204824
2025-02-23T13:10:45Z42150267246974
2025-02-23T20:42:28Z943125268190099
2025-02-27T09:58:56Z19707299287897398
2025-02-27T11:29:06Z32681287930079
2025-03-03T00:37:56Z19108585307038664
2025-03-03T02:08:08Z55850307094514
2025-03-06T14:49:14Z18786755325881269
2025-03-10T03:59:48Z19159040345040309
2025-03-13T16:29:43Z18709163363749472
2025-03-17T04:46:24Z18882963382632435
2025-06-09T14:18:19Z127576146510208581