datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-price-index
AI Price Index
An open, dated, first-party-sourced record of AI model API prices over time.
Provider pricing pages change quietly with no changelog. This dataset is the changelog. Every price carries the date it became valid and the date it was last verified against the official source, so you can price historical token usage point-in-time instead of extrapolating from today's rate.
512 price records, 125 models, 11 providers: Anthropic, OpenAI, Google, Mistral, xAI, DeepSeek… See the full description on the dataset page: https://huggingface.co/datasets/RoninForge/ai-price-index.twitter-trending-hashtags
Twitter/X Trending Hashtags (2020-2025)
A comprehensive dataset of trending hashtags on Twitter/X from 2020 to 2025, containing 12,036 unique trend entries across six years, capturing major world events, cultural moments, and viral phenomena.
📊 Dataset Description
This dataset captures trending hashtags from Twitter/X (formerly Twitter) by analyzing Wayback Machine snapshots of trends24.in, providing insights into breaking news, viral content, cultural moments, and… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/twitter-trending-hashtags.github-top-projects
GitHub Trending Projects (2013-2025)
A comprehensive dataset of 423,098 GitHub trending repository entries spanning 12+ years (August 2013 - November 2025), scraped from Wayback Machine snapshots of GitHub's trending page.
🎯 Dataset Overview
This dataset captures the evolution of GitHub's trending repositories over time, providing insights into:
Software development trends across programming languages and domains
Popular open-source projects and their trending patterns… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-projects.moltbook
Moltbook Dataset
A dataset of posts and communities from Moltbook - a Reddit-style social platform designed for AI agents.
NOTE: This dataset is a snapshot of Moltbook before it went viral and got flooded with inauthentic accounts such as humans and bots.
Files
File
Records
Description
moltbook_posts.csv
6,105
All posts from the platform
moltbook_submolts.csv
124
All communities (submolts)
Dataset Insights
Overview… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/moltbook.tiktok-trending-hashtags
TikTok Trending Hashtags (2022-2025)
A comprehensive dataset of trending hashtags on TikTok from 2022 to 2025, containing 1,830 unique hashtag entries across multiple years, languages, and cultural contexts
📊 Dataset Description
This dataset captures trending hashtags from TikTok's Creative Center, providing insights into viral content, cultural moments, and global events from 2022 to 2025.
Data Source: TikTok Creative Center - Popular Hashtags
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/tiktok-trending-hashtags.pdb_sequences
PDB Sequences
This dataset contains 780,163 protein sequences from the RCCB Protein Data Bank
trending-stocks-yahoo-finance
Monthly Trending Stocks Dataset
A ranked dataset of the most trending stocks on Yahoo Finance from July 2024 to October 2025, based on weighted scoring of their monthly trending appearances in Yahoo Finance.
📊 Dataset Overview
Total Entries: 7,993 ranked stocks
Time Period: July 2024 - October 2025 (16 months)
Source: Wayback Machine snapshots of Yahoo Finance Trending Stocks
Data Granularity: Monthly rankings
Data Order: Sorted by month (descending: Oct 2025 → July… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/trending-stocks-yahoo-finance.world-airports
Title: World Airport Dataset
Dataset Summary
A comprehensive dataset containing information about airports around the world, including location, airport codes, and other relevant details.
Dataset Description
The World Airport Dataset is a comprehensive collection of information about airports worldwide. This dataset includes various attributes for each airport, such as unique identifiers, location details, airport codes, and geographical coordinates. The data… See the full description on the dataset page: https://huggingface.co/datasets/ronnieaban/world-airports.protein_binding_sequences
Sequence Based Protein - Peptide Binding Dataset
Data sources:
Huang Laboratory
Propedia
YAPP-Cd
Dataset size: 16,370 sets of Protein-Peptide sequences that bind, the protein sequence
contains only the relevant chain.
Train / Val split: the dataset is split to 80% train 10% val and 10% test.
github-top-developers
GitHub Top Developers by Year (2015-2025)
A derived dataset showing the top-ranked GitHub trending developers for each year, based on weighted scoring of their trending appearances across 41,841 raw data points from the Wayback Machine.
📊 Dataset Overview
Total Entries: 8,125 ranked developers
Years Covered: 2015 - 2025 (11 years)
Unique Developers: 4,763
Source: Derived from Wayback Machine snapshots of GitHub trending developers
Data Order: Sorted by year (descending:… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-developers.isichuggingface-top-papers
HuggingFace Top Trending Papers (2025)
A ranked dataset of the most trending papers on HuggingFace Daily Papers in 2025, based on weighted scoring of their trending appearances. This dataset captures which AI/ML research papers gained the most community attention and sustained visibility.
📊 Dataset Overview
Total Entries: 663 ranked papers
Time Period: 2025 (Jan - Nov)
Source: Wayback Machine snapshots of HuggingFace Daily Papers
Unique Papers: 663
Data Order: Sorted by… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/huggingface-top-papers.hs-codealquran
Dataset Terjemahan dan Tafsir Al-Quran
Deskripsi Dataset
Dataset ini berisi terjemahan Al-Quran dalam bahasa Indonesia beserta tafsirnya. Dataset ini dapat digunakan untuk berbagai tugas NLP seperti machine translation, text generation, dan text summarization.
Fitur Utama
Terjemahan Al-Quran: Teks Al-Quran dalam bahasa Arab beserta terjemahannya dalam bahasa Indonesia.
Tafsir Al-Quran: Penjelasan atau interpretasi dari ayat-ayat Al-Quran dalam bahasa… See the full description on the dataset page: https://huggingface.co/datasets/ronnieaban/alquran.Telangana_time_series_2023-2025The dataset was retrieved from Open Data Telangana, from February 1, 2023, to January 31, 2025 with daily granularity. The dataset contains various fields such as District, Mandal, Date, rainfall (in millimeters), minimum and maximum temperature (in Celsius), minimum and maximum wind speed, and humidity. It provides a District and Mandal wise distribution as well.
Total Rows - 4,45,213
Total Columns - 10
pubmed-10k
Dataset Summary
First 10k rows of the scientific_papers["pubmed"] dataset. 10:1:1 split.
Usage
from datasets import load_dataset
train_dataset = load_dataset("ronitHF/pubmed-10k", split="train")
val_dataset = load_dataset("ronitHF/pubmed-10k", split="validation")
test_dataset = load_dataset("ronitHF/pubmed-10k", split="test")
cristiano-ronaldo-all-club-goals-stats
Context
This dataset contains all the stats of all club goals of Cristiano Ronaldo dos Santos Aveiro.
About Cristiano Ronaldo
Cristiano Ronaldo dos Santos Aveiro is a Portuguese professional footballer who plays as a forward for Premier League club Manchester United and captains the Portugal national team.
Current team: Portugal national football team (#7 / Forward) Trending
Born: February 5, 1985 (age 37 years), Hospital Dr. Nélio Mendonça, Funchal, Portugal
Height:… See the full description on the dataset page: https://huggingface.co/datasets/azminetoushikwasi/cristiano-ronaldo-all-club-goals-stats.kbli2020
Dataset Card for Dataset Name
Dataset ini berisi Klasifikasi Baku Lapangan Usaha Indonesia (KBLI) tahun 2020 yang merupakan standar klasifikasi aktivitas ekonomi yang ditetapkan oleh Badan Pusat Statistik (BPS) Indonesia. KBLI 2020 digunakan untuk mengklasifikasikan unit usaha dan aktivitas ekonomi ke dalam kelompok yang seragam berdasarkan kegiatan ekonomi yang dilakukan.
japanese-character-difficulty
Japanese Character Difficulty Dataset
A comprehensive dataset of 3,003 Japanese kanji characters with their educational difficulty grades, sourced from official Japanese educational standards and kanjiapi.dev.
Dataset Overview
Total Characters: 3,003 kanji
Source: Japanese Ministry of Education (MEXT) Joyo Kanji list + kanjiapi.dev
Coverage: Elementary grades 1-6, plus secondary education and advanced characters
Format: Character-grade pairs for easy lookup and analysis… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/japanese-character-difficulty.CIC-DDoS2019-15C
CIC-DDoS2019-15C
A preprocessed version of the CIC-DDoS2019 dataset with feature selection applied, as detailed in the paper: Distributed Denial of Service Detection: Enhancing Machine Learning Models for Multiclass Classification. The dataset contains 15 columns (index + features + labels) and is used for detecting Multiclass DDoS Attacks.
There are 11 attack classes present (DrDoS_DNS, DrDoS_LDAP, DrDoS_MSSQL, DrDoS_NetBIOS, DrDoS_NTP, DrDoS_SNMP, DrDoS_SSDP, Syn, TFTP, DrDoS_UDP… See the full description on the dataset page: https://huggingface.co/datasets/RonaldoMPF/CIC-DDoS2019-15C.anacreonteaThis repository contains the dataset for the paper Ancient Greek's New Technological Muse: Extracting Topoi in the Anacreontea with LLMs, accepted at the 51st SEMISH (51º Seminário Integrado de Software e Hardware).
Abstract:
Natural Language Processing (NLP), along with Large Language Models (LLMs), holds significant potential in the domain of literature, leveraging its computational capabilities to analyze and comprehend human language. These techniques prove to be particularly useful in a… See the full description on the dataset page: https://huggingface.co/datasets/ronunes/anacreontea.sunnah
Deskripsi Dataset
Dataset ini berisi kumpulan teks hadits dan sunnah Rasulullah SAW. Kontennya mencakup ajaran Islam yang diambil dari sumber terpercaya, yang dapat digunakan untuk berbagai tugas Natural Language Processing (NLP) seperti klasifikasi teks, penjawaban pertanyaan, dan analisis teks.
Bahasa: Indonesia dan Arab.
Lisensi: Open Data Commons Public Domain Dedication and License (PDDL), lisensi yang memungkinkan pengguna untuk berbagi, memodifikasi, dan menggunakan data… See the full description on the dataset page: https://huggingface.co/datasets/ronnieaban/sunnah.world-seaports
World Seaport Dataset
Description
This dataset includes detailed information about global seaports, covering aspects such as location, dimensions, facilities, and services available. It aims to assist in maritime navigation, logistics planning, and geographical studies.
Metadata
Dataset: A collection of worldwide seaport information with extensive attributes.
Source: Compiled from various maritime publications and official nautical charts.
Columns and… See the full description on the dataset page: https://huggingface.co/datasets/ronnieaban/world-seaports.automotive_requirements
Dataset Card for autoReq
Importing dataset into Python environment
Use the following code chunk to import the dataset into a Python environment as a DataFrame.
pubmed-10k-8.1.1
Dataset Summary
First 10k rows of the scientific_papers["pubmed"] dataset. 8:1:1 split (10000:1250:1250).
Usage
from datasets import load_dataset
train_dataset = load_dataset("ronitHF/pubmed-10k-8.1.1", split="train")
val_dataset = load_dataset("ronitHF/pubmed-10k-8.1.1", split="validation")
test_dataset = load_dataset("ronitHF/pubmed-10k-8.1.1", split="test")
icatusphising_emailvizi_datasetDataset Overview
This dataset is a synthetically generated collection of AI-style questions designed to simulate how users ask questions across different domains and intents in AI-driven environments.
Each row represents a single question along with structured metadata that enables analysis of question patterns, intent distribution, and a proxy for GEO-style question volume.
The dataset was generated using a pre-trained language model and does not rely on real user data.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/ronschwartz/vizi_dataset.CebuanoAnnotatedFakeandLegitNewsjob_description_sample
