CoolFace
Datasetpublic

tensorshield/reddit_dataset_85

Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/tensorshield/reddit_dataset_85.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes964downloads
Dataset Card

Bittensor Subnet 13 Reddit Dataset

<center> <img src="https://huggingface.co/datasets/macrocosm-os/images/resolve/main/bittensor.png" alt="Data-universe: The finest collection of social media data the web has to offer"> </center>

<center> <img src="https://huggingface.co/datasets/macrocosm-os/images/resolve/main/macrocosmos-black.png" alt="Data-universe: The finest collection of social media data the web has to offer"> </center>

Dataset Description

  • Repository: tensorshield/redditdataset85
  • Subnet: Bittensor Subnet 13
  • Miner Hotkey: 5FvjTc3UHpdft9Jusu1hbj87czkqHXTrkLvkBakcWLjZKv1X

Dataset Summary

This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository.

Supported Tasks

The versatility of this dataset allows researchers and data scientists to explore various aspects of social media dynamics and develop innovative applications. Users are encouraged to leverage this data creatively for their specific research or business needs. For example:

  • Sentiment Analysis
  • Topic Modeling
  • Community Analysis
  • Content Categorization

Languages

Primary language: Datasets are mostly English, but can be multilingual due to decentralized ways of creation.

Dataset Structure

Data Instances

Each instance represents a single Reddit post or comment with the following fields:

Data Fields

  • text (string): The main content of the Reddit post or comment.
  • label (string): Sentiment or topic category of the content.
  • dataType (string): Indicates whether the entry is a post or a comment.
  • communityName (string): The name of the subreddit where the content was posted.
  • datetime (string): The date when the content was posted or commented.
  • username_encoded (string): An encoded version of the username to maintain user privacy.
  • url_encoded (string): An encoded version of any URLs included in the content.

Data Splits

This dataset is continuously updated and does not have fixed splits. Users should create their own splits based on their requirements and the data's timestamp.

Dataset Creation

Source Data

Data is collected from public posts and comments on Reddit, adhering to the platform's terms of service and API usage guidelines.

Personal and Sensitive Information

All usernames and URLs are encoded to protect user privacy. The dataset does not intentionally include personal or sensitive information.

Considerations for Using the Data

Social Impact and Biases

Users should be aware of potential biases inherent in Reddit data, including demographic and content biases. This dataset reflects the content and opinions expressed on Reddit and should not be considered a representative sample of the general population.

Limitations

  • Data quality may vary due to the nature of media sources.
  • The dataset may contain noise, spam, or irrelevant content typical of social media platforms.
  • Temporal biases may exist due to real-time collection methods.
  • The dataset is limited to public subreddits and does not include private or restricted communities.

Additional Information

Licensing Information

The dataset is released under the MIT license. The use of this dataset is also subject to Reddit Terms of Use.

Citation Information

If you use this dataset in your research, please cite it as follows:

@misc{tensorshield2025datauniversereddit_dataset_85,
        title={The Data Universe Datasets: The finest collection of social media data the web has to offer},
        author={tensorshield},
        year={2025},
        url={https://huggingface.co/datasets/tensorshield/reddit_dataset_85},
        }

Contributions

To report issues or contribute to the dataset, please contact the miner or use the Bittensor Subnet 13 governance mechanisms.

Dataset Statistics

[This section is automatically updated]

  • Total Instances: 150311
  • Date Range: 2025-03-24T00:00:00Z to 2025-03-24T00:00:00Z
  • Last Updated: 2025-03-31T01:48:18Z

Data Distribution

  • Posts: 12.10%
  • Comments: 87.90%

Top 10 Subreddits

For full statistics, please refer to the stats.json file in the repository.

RankTopicTotal CountPercentage
1r/CollegeBasketball57953.86%
2r/AskReddit38052.53%
3r/AITAH19071.27%
4r/mildlyinfuriating11960.80%
5r/90DayFiance8360.56%
6r/denvernuggets8050.54%
7r/tennis7140.48%
8r/politics7100.47%
9r/Advice6870.46%
10r/moviecritic6860.46%

Update History

DateNew InstancesTotal Instances
2025-03-31T00:13:38Z14751475
2025-03-31T00:14:35Z15863061
2025-03-31T00:15:21Z15694630
2025-03-31T00:16:19Z13665996
2025-03-31T00:17:15Z15567552
2025-03-31T00:18:33Z18249376
2025-03-31T00:19:21Z118610562
2025-03-31T00:20:19Z164412206
2025-03-31T00:21:24Z154213748
2025-03-31T00:22:14Z140915157
2025-03-31T00:23:19Z156416721
2025-03-31T00:24:14Z142518146
2025-03-31T00:25:32Z152719673
2025-03-31T00:26:28Z159421267
2025-03-31T00:27:19Z144222709
2025-03-31T00:28:22Z140324112
2025-03-31T00:29:30Z165025762
2025-03-31T00:30:22Z134527107
2025-03-31T00:31:16Z150728614
2025-03-31T00:32:16Z139930013
2025-03-31T00:33:21Z158331596
2025-03-31T00:34:21Z151133107
2025-03-31T00:35:19Z143034537
2025-03-31T00:36:12Z144635983
2025-03-31T00:37:16Z163237615
2025-03-31T00:38:16Z149939114
2025-03-31T00:39:15Z147440588
2025-03-31T00:40:14Z150242090
2025-03-31T00:41:13Z147543565
2025-03-31T00:42:14Z154645111
2025-03-31T00:43:13Z157846689
2025-03-31T00:44:14Z152948218
2025-03-31T01:04:21Z3246980687
2025-03-31T01:05:13Z142082107
2025-03-31T01:06:14Z157083677
2025-03-31T01:07:13Z155985236
2025-03-31T01:08:16Z156786803
2025-03-31T01:09:13Z145088253
2025-03-31T01:10:15Z160689859
2025-03-31T01:11:15Z158491443
2025-03-31T01:12:15Z148892931
2025-03-31T01:13:15Z170294633
2025-03-31T01:14:50Z417398806
2025-03-31T01:16:52Z3554102360
2025-03-31T01:17:43Z1393103753
2025-03-31T01:18:22Z1018104771
2025-03-31T01:19:14Z1623106394
2025-03-31T01:20:16Z1759108153
2025-03-31T01:21:16Z1562109715
2025-03-31T01:22:13Z1418111133
2025-03-31T01:24:24Z3231114364
2025-03-31T01:25:21Z1440115804
2025-03-31T01:26:15Z1306117110
2025-03-31T01:27:12Z1449118559
2025-03-31T01:28:12Z1450120009
2025-03-31T01:29:11Z1475121484
2025-03-31T01:30:11Z1537123021
2025-03-31T01:31:12Z1530124551
2025-03-31T01:32:12Z1497126048
2025-03-31T01:33:11Z1487127535
2025-03-31T01:34:13Z1445128980
2025-03-31T01:35:12Z1492130472
2025-03-31T01:36:25Z1834132306
2025-03-31T01:37:14Z1204133510
2025-03-31T01:38:16Z1563135073
2025-03-31T01:39:14Z1413136486
2025-03-31T01:40:27Z1835138321
2025-03-31T01:41:14Z1240139561
2025-03-31T01:42:19Z1535141096
2025-03-31T01:44:01Z2641143737
2025-03-31T01:44:35Z859144596
2025-03-31T01:45:15Z977145573
2025-03-31T01:46:30Z1728147301
2025-03-31T01:47:15Z1397148698
2025-03-31T01:48:18Z1613150311