CoolFace
Datasetpublic

immortalizzy/reddit_dataset_197

Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/immortalizzy/reddit_dataset_197.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes543downloads
Dataset Card

Bittensor Subnet 13 Reddit Dataset

<center> <img src="https://huggingface.co/datasets/macrocosm-os/images/resolve/main/bittensor.png" alt="Data-universe: The finest collection of social media data the web has to offer"> </center>

<center> <img src="https://huggingface.co/datasets/macrocosm-os/images/resolve/main/macrocosmos-black.png" alt="Data-universe: The finest collection of social media data the web has to offer"> </center>

Dataset Description

  • Repository: premierinspe/redditdataset197
  • Subnet: Bittensor Subnet 13
  • Miner Hotkey: 5CGUgWdufFyb8PH4zYSVZtUo5nJVBcwVRtUYh6N51tXUqMAC

Dataset Summary

This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository.

Supported Tasks

The versatility of this dataset allows researchers and data scientists to explore various aspects of social media dynamics and develop innovative applications. Users are encouraged to leverage this data creatively for their specific research or business needs. For example:

  • Sentiment Analysis
  • Topic Modeling
  • Community Analysis
  • Content Categorization

Languages

Primary language: Datasets are mostly English, but can be multilingual due to decentralized ways of creation.

Dataset Structure

Data Instances

Each instance represents a single Reddit post or comment with the following fields:

Data Fields

  • text (string): The main content of the Reddit post or comment.
  • label (string): Sentiment or topic category of the content.
  • dataType (string): Indicates whether the entry is a post or a comment.
  • communityName (string): The name of the subreddit where the content was posted.
  • datetime (string): The date when the content was posted or commented.
  • username_encoded (string): An encoded version of the username to maintain user privacy.
  • url_encoded (string): An encoded version of any URLs included in the content.

Data Splits

This dataset is continuously updated and does not have fixed splits. Users should create their own splits based on their requirements and the data's timestamp.

Dataset Creation

Source Data

Data is collected from public posts and comments on Reddit, adhering to the platform's terms of service and API usage guidelines.

Personal and Sensitive Information

All usernames and URLs are encoded to protect user privacy. The dataset does not intentionally include personal or sensitive information.

Considerations for Using the Data

Social Impact and Biases

Users should be aware of potential biases inherent in Reddit data, including demographic and content biases. This dataset reflects the content and opinions expressed on Reddit and should not be considered a representative sample of the general population.

Limitations

  • Data quality may vary due to the nature of media sources.
  • The dataset may contain noise, spam, or irrelevant content typical of social media platforms.
  • Temporal biases may exist due to real-time collection methods.
  • The dataset is limited to public subreddits and does not include private or restricted communities.

Additional Information

Licensing Information

The dataset is released under the MIT license. The use of this dataset is also subject to Reddit Terms of Use.

Citation Information

If you use this dataset in your research, please cite it as follows:

@misc{premierinspe2025datauniversereddit_dataset_197,
        title={The Data Universe Datasets: The finest collection of social media data the web has to offer},
        author={premierinspe},
        year={2025},
        url={https://huggingface.co/datasets/premierinspe/reddit_dataset_197},
        }

Contributions

To report issues or contribute to the dataset, please contact the miner or use the Bittensor Subnet 13 governance mechanisms.

Dataset Statistics

[This section is automatically updated]

  • Total Instances: 18506447
  • Date Range: 2025-03-07T00:00:00Z to 2025-03-17T00:00:00Z
  • Last Updated: 2025-03-28T13:00:20Z

Data Distribution

  • Posts: 6.56%
  • Comments: 93.44%

Top 10 Subreddits

For full statistics, please refer to the stats.json file in the repository.

RankTopicTotal CountPercentage
1r/AskReddit3833822.07%
2r/AITAH2023361.09%
3r/politics1322390.71%
4r/wallstreetbets1254590.68%
5r/mildlyinfuriating1022910.55%
6r/pics906970.49%
7r/CollegeBasketball885620.48%
8r/NoStupidQuestions798280.43%
9r/moviecritic791510.43%
10r/teenagers778310.42%

Update History

DateNew InstancesTotal Instances
2025-03-27T09:52:19Z972559972559
2025-03-27T20:51:54Z9716531944212
2025-03-27T21:03:16Z9737972918009
2025-03-27T22:03:12Z9700363888045
2025-03-27T23:03:15Z9741394862184
2025-03-28T00:03:25Z9729275835111
2025-03-28T00:29:22Z9748006809911
2025-03-28T01:03:17Z9725787782489
2025-03-28T02:03:56Z9748318757320
2025-03-28T03:04:16Z9722329729552
2025-03-28T04:04:24Z97400210703554
2025-03-28T05:03:56Z97193011675484
2025-03-28T06:03:26Z97308312648567
2025-03-28T07:03:40Z96954513618112
2025-03-28T08:03:54Z96933014587442
2025-03-28T09:03:53Z96839715555839
2025-03-28T10:03:45Z97158816527427
2025-03-28T11:03:38Z97031617497743
2025-03-28T12:03:44Z96998318467726
2025-03-28T12:44:14Z1934918487075
2025-03-28T13:00:20Z1937218506447