datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
douban_movie_reviewgo-mo-dataset
GO-MO, a massive Graph agumented Open urban MObility dataset
This is the official dataset repository for the GO-MO traffic dataset.
The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain).
GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024).
Additionally, the GO-MO dataset introduces two graph… See the full description on the dataset page: https://huggingface.co/datasets/double-blind-anonymous/go-mo-dataset.douban_movie_info该数据集为豆瓣电影信息维表。
更多信息请参考文章《数据获取:豆瓣电影信息爬取》。
dou-brazil-dataset
Dataset Card for Dataset Diário Oficial da União (DOU)
The Diário Oficial da União (DOU) is the official government gazette of Brazil, published by the National Press. It serves as the primary means of communication for federal government acts, including laws, decrees, ordinances, public notices, and other official decisions. The DOU ensures transparency and legal validity for government actions and is divided into three sections:
Section 1: Publishes laws, decrees, and… See the full description on the dataset page: https://huggingface.co/datasets/gerson-vfs/dou-brazil-dataset.spam-douban-movie-review
Description
The Spam Douban Movie Reviews Dataset is a collection of movie reviews scraped from Douban, a popular Chinese social networking platform for movie enthusiasts. This dataset consists of reviews that have been manually classified as either spam or genuine by human reviewers. It contains a total of 1,600 data.
This dataset is created for our project Spam Movie Reviews Detection through Supervised Learning.
MMMLU_subset
About MMMLU subset
This is a subset of MMMLU, specifically, we sampled 10% of the original data to improve evaluation efficiency.
In addition, we categorize the questions into four categories by subject, i.e., STEM, HUMANITIES, SOCIAL SCIENCES, and OTHER, aligned with MMLU.
Multilingual Massive Multitask Language Understanding (MMMLU)
The MMLU is a widely recognized benchmark of general knowledge attained by AI models. It covers a broad range of topics from 57… See the full description on the dataset page: https://huggingface.co/datasets/double7/MMMLU_subset.douban_movie_info该数据集为豆瓣电影信息维表。
更多信息请参考文章《数据获取:豆瓣电影信息爬取》。
sufficient-mistake-2c0e2f
sufficient-mistake-2c0e2f
Synthetic products test data: 50 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/douglasjones/sufficient-mistake-2c0e2f.douyinhuman_proteome_doublets
Dataset Description
Out of 20,577 human proteins (from UniProt human proteome), sequences shorter than 20 amino acids or longer than 512 amino acids were removed, resulting in a set of 12,703 proteins. The uShuffle algorithm (python pacakge) was then used to shuffle these protein sequences while maintaining their doublet distribution. The very few sequences for which uShuffle failed to create a shuffled version were eliminated.
Afterwards, h-CD-HIT algorithm (web server) was used… See the full description on the dataset page: https://huggingface.co/datasets/yarongef/human_proteome_doublets.webvid10m_motiontest-shap-doubletrain0xdf-20-most-recent-posts-8-11-2025single_double_digit_additionchat_0_shot_doubletrendyol-turkish-product-reviews⸻
license: cc-by-nc-4.0
language:
• tr
tags:
• turkish
• nlp
• sentiment-analysis
• ecommerce
• text-classification
• reviews
pretty_name: Trendyol Turkish Product Reviews
⸻
582K+ Turkish E-Commerce Reviews from Trendyol
A large-scale Turkish product reviews dataset collected from Trendyol, designed for Natural Language Processing (NLP), sentiment analysis, and consumer opinion mining. The dataset captures real-world, user-generated Turkish text, including informal… See the full description on the dataset page: https://huggingface.co/datasets/Doukan/trendyol-turkish-product-reviews.Doubt_And_NonDoubtcybersecblogs
