CoolFace
Datasetpublic

smartnuel87/reddit_dataset_239

Bittensor Subnet 13 Reddit Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For… See the full description on the dataset page: https://huggingface.co/datasets/smartnuel87/reddit_dataset_239.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes76downloads
Dataset Card

Bittensor Subnet 13 Reddit Dataset

<center> <img src="https://huggingface.co/datasets/macrocosm-os/images/resolve/main/bittensor.png" alt="Data-universe: The finest collection of social media data the web has to offer"> </center>

<center> <img src="https://huggingface.co/datasets/macrocosm-os/images/resolve/main/macrocosmos-black.png" alt="Data-universe: The finest collection of social media data the web has to offer"> </center>

Dataset Description

  • Repository: smartnuel87/redditdataset239
  • Subnet: Bittensor Subnet 13
  • Miner Hotkey: 5D2qXEaNxxk2j2Bh7cTa5Y8xKZ4p1KAFMTBn6iKWNBpcJyj3

Miner Data Compliance Agreement

In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.

Dataset Summary

This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository.

Supported Tasks

The versatility of this dataset allows researchers and data scientists to explore various aspects of social media dynamics and develop innovative applications. Users are encouraged to leverage this data creatively for their specific research or business needs. For example:

  • Sentiment Analysis
  • Topic Modeling
  • Community Analysis
  • Content Categorization

Languages

Primary language: Datasets are mostly English, but can be multilingual due to decentralized ways of creation.

Dataset Structure

Data Instances

Each instance represents a single Reddit post or comment with the following fields:

Data Fields

  • text (string): The main content of the Reddit post or comment.
  • label (string): Sentiment or topic category of the content.
  • dataType (string): Indicates whether the entry is a post or a comment.
  • communityName (string): The name of the subreddit where the content was posted.
  • datetime (string): The date when the content was posted or commented.
  • username_encoded (string): An encoded version of the username to maintain user privacy.
  • url_encoded (string): An encoded version of any URLs included in the content.

Data Splits

This dataset is continuously updated and does not have fixed splits. Users should create their own splits based on their requirements and the data's timestamp.

Dataset Creation

Source Data

Data is collected from public posts and comments on Reddit, adhering to the platform's terms of service and API usage guidelines.

Personal and Sensitive Information

All usernames and URLs are encoded to protect user privacy. The dataset does not intentionally include personal or sensitive information.

Considerations for Using the Data

Social Impact and Biases

Users should be aware of potential biases inherent in Reddit data, including demographic and content biases. This dataset reflects the content and opinions expressed on Reddit and should not be considered a representative sample of the general population.

Limitations

  • Data quality may vary due to the nature of media sources.
  • The dataset may contain noise, spam, or irrelevant content typical of social media platforms.
  • Temporal biases may exist due to real-time collection methods.
  • The dataset is limited to public subreddits and does not include private or restricted communities.

Additional Information

Licensing Information

The dataset is released under the MIT license. The use of this dataset is also subject to Reddit Terms of Use.

Citation Information

If you use this dataset in your research, please cite it as follows:

@misc{smartnuel872025datauniversereddit_dataset_239,
        title={The Data Universe Datasets: The finest collection of social media data the web has to offer},
        author={smartnuel87},
        year={2025},
        url={https://huggingface.co/datasets/smartnuel87/reddit_dataset_239},
        }

Contributions

To report issues or contribute to the dataset, please contact the miner or use the Bittensor Subnet 13 governance mechanisms.

Dataset Statistics

[This section is automatically updated]

  • Total Instances: 2400
  • Date Range: 2025-06-13T00:00:00Z to 2025-06-29T00:00:00Z
  • Last Updated: 2025-07-30T20:38:05Z

Data Distribution

  • Posts: 5.58%
  • Comments: 94.42%

Top 10 Subreddits

For full statistics, please refer to the stats.json file in the repository.

RankTopicTotal CountPercentage
1r/AskReddit482.00%
2r/AITAH200.83%
3r/mildlyinfuriating200.83%
4r/politics180.75%
5r/BotBouncer180.75%
6r/LoveIslandUSA160.67%
7r/AmIOverreacting140.58%
8r/teenagers140.58%
9r/nba140.58%
10r/wallstreetbets120.50%

Update History

DateNew InstancesTotal Instances
2025-07-15T03:54:55Z100100
2025-07-15T04:11:02Z100200
2025-07-15T05:45:13Z100300
2025-07-15T07:20:56Z100400
2025-07-15T08:55:08Z100500
2025-07-15T10:03:46Z100600
2025-07-16T03:14:31Z100700
2025-07-16T21:18:46Z100800
2025-07-17T15:30:39Z100900
2025-07-18T09:36:54Z1001000
2025-07-19T03:14:46Z1001100
2025-07-19T21:27:02Z1001200
2025-07-22T13:43:38Z1001300
2025-07-23T07:50:49Z1001400
2025-07-24T01:59:12Z1001500
2025-07-24T20:03:34Z1001600
2025-07-25T14:09:40Z1001700
2025-07-26T08:13:28Z1001800
2025-07-27T02:19:02Z1001900
2025-07-27T20:22:21Z1002000
2025-07-28T14:25:47Z1002100
2025-07-29T08:28:43Z1002200
2025-07-30T02:32:05Z1002300
2025-07-30T20:38:05Z1002400