datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shofo-tiktok-general-small
Shofo TikTok General (Small)
Overview
Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive metadata, transcripts, comments, and engagement metrics. This is a curated subset of Shofo's larger TikTok index, which contains hundreds of millions of indexed videos.
Size: ~50K videos (~500GB)
Modality: Video + Audio + Text (transcripts, comments, captions)
Source: TikTok
Schema
Column
Type
Description
file_name… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-tiktok-general-small.tiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/kwakuobeng/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/blaccastro/tiktok-videos-4b.fineweb-edu-sample-10BT-tiktokenizedtiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/dams2005/tiktok-videos-4b.youtube-tiktok-trends-dataset-2025
🎬 YouTube Shorts & TikTok Trends (2025)
Author: Tarek MasryoLicense: CC BY 4.0
A structured snapshot of short-form video activity across YouTube Shorts and TikTok during 2025 (Jan–Aug).Built for content intelligence, analytics dashboards, and ML baselines (classification/regression).
What’s inside
This repository ships:
Two loadable dataset configs (via datasets.load_dataset):
default → ML-ready table (cleaned + modeling-friendly)
raw → raw video-level table (wider… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/youtube-tiktok-trends-dataset-2025.tiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/KOM-00/tiktok-videos-4b.TikTok-10M
TikTok-10M Dataset
Dataset Description
TikTok-10M is a large-scale dataset containing 10 million short-form posts from TikTok, designed for video understanding, multimodal learning, and social media content analysis. The dataset was curated to bridge the gap between academic video datasets and actual user-generated content, providing researchers with authentic patterns and characteristics of modern short-form video content that dominates social media platforms.… See the full description on the dataset page: https://huggingface.co/datasets/The-data-company/TikTok-10M.tiktok-video-engagement-200k
TikTok Creator and Video Engagement (200K)
This release contains 209,543 TikTok videos from 1,872 creators with daily engagement and follower statistics, covering videos posted from 2024-06-24 to 2024-11-09.
The release contains derived video-content fields. Audio transcripts, screenshots, textual data, and music metadata were used to generate short video summaries; topic labels and emotion scores were then derived from those summaries using machine learning models. These… See the full description on the dataset page: https://huggingface.co/datasets/lingbow/tiktok-video-engagement-200k.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/alex12223322/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/hojj/tiktok-videos-4b.tiktok-videos-users-info
TikTok Data, post + poster (user) info, ~1 Million
Dataset name: EinzzCookie/tiktok-videos-users-info
This dataset contains a large collection of TikTok video records paired with detailed creator/user information, stored in a single Parquet file (tiktok_video_user_data.parquet, ~3.87 GB). It is derived from TikTok’s internal video (“aweme”) data model and includes both post-level metadata/engagement stats and nested author profile data.
Source
Collected by… See the full description on the dataset page: https://huggingface.co/datasets/EinzzCookie/tiktok-videos-users-info.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/mrfakename/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/seanphan/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/kkndlee/tiktok-videos-4b.tiktok-video-engagement-1m
TikTok Creator and Video Engagement (1M)
This release contains 1,035,817 TikTok videos from 4,926 creators with daily engagement and follower statistics, covering videos posted from 2024-06-09 to 2025-03-20.
Github: https://github.com/lingbowzd/tiktok-creator-video-trend-data
Cite this dataset: When does Trend-following Pay off? Evidence from Trending Content and Hashtag use
Uses
This dataset supports research on TikTok creator behavior, content strategy, trend… See the full description on the dataset page: https://huggingface.co/datasets/lingbow/tiktok-video-engagement-1m.tiktok
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/Suthans/tiktok.tiktok-videos-sea
TikTok Videos Southeast Asia (MY / TH / ID / SG)
Country-filtered subset of kuben-developer/tiktok-videos-4b:
every row whose country field is MY (Malaysia), TH (Thailand), ID (Indonesia) or SG (Singapore).
Extracted 2026-09-07 from all 27 source parquet files; schema unchanged.
Sizes (verified against the uploaded files)
Config
Country
Rows
Distinct content_id
Ads (is_ad=1)
With caption
my
Malaysia
11,729,283
11,729,283
967,782
9,787,874
th
Thailand… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/tiktok-videos-sea.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/abdellatifinformation/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/tiktok-videos-4b.tiktok-videos-4b
Mirror of kuben-developer/tiktok-videos-4b, snapshot 2026-09-08. All credit to the original author; same research-use license applies.
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.… See the full description on the dataset page: https://huggingface.co/datasets/merway/tiktok-videos-4b.mya-tiktok-asr-120h
Burmese TikTok ASR (121h)
A weakly-supervised Burmese (Myanmar, my) speech corpus: 168,852 short audio clips / 121.1 hours, segmented from 3,902 public TikTok videos and paired with the Burmese subtitles TikTok generates for those videos.
Intended for pre-training and fine-tuning Burmese ASR models (e.g. Whisper) in a language with very little open speech data.
⚠️ Read this first. The transcripts are machine-generated, not human-verified — see Labels are ASR output. Treat that… See the full description on the dataset page: https://huggingface.co/datasets/t7188409/mya-tiktok-asr-120h.Tiktok-Videos
TikTok Video Analytics Dataset
Sample TikTok video dataset with comprehensive engagement metrics and metadata. Each row represents a single TikTok video with content and detailed analytics.
This is a sample dataset. To access the full version or request any custom dataset tailored to your needs, contact DataHive at contact@datahive.ai.
Files Included
train.csv – TikTok video analytics data
What's included
Video URLs and identifiers
Comprehensive engagement… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/Tiktok-Videos.tiktok-trending-hashtags
TikTok Trending Hashtags (2022-2025)
A comprehensive dataset of trending hashtags on TikTok from 2022 to 2025, containing 1,830 unique hashtag entries across multiple years, languages, and cultural contexts
📊 Dataset Description
This dataset captures trending hashtags from TikTok's Creative Center, providing insights into viral content, cultural moments, and global events from 2022 to 2025.
Data Source: TikTok Creative Center - Popular Hashtags
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/tiktok-trending-hashtags.tiktok-techjam-2026-eval
TikTok TechJam 2026 Eval
Held-out demonstration pair used by Seer:
COCO val2017 photographs plus the WildFake DALL·E Advanced (DALL·E 3) subset.
Do not train on this split.
Contents
label
meaning
count
origin
0 / real
photograph
5,000
COCO val2017
1 / fake
AI-generated
8,843
WildFake DALL·E Advanced
Columns: image, label, source, generator, id.
id is the COCO stem for reals, and {session}_{stem} for fakes so duplicate
WildFake basenames stay… See the full description on the dataset page: https://huggingface.co/datasets/glennwuwu/tiktok-techjam-2026-eval.TikTok-10M
TikTok-10M Dataset
Dataset Description
TikTok-10M is a large-scale dataset containing 10 million short-form posts from TikTok, designed for video understanding, multimodal learning, and social media content analysis. The dataset was curated to bridge the gap between academic video datasets and actual user-generated content, providing researchers with authentic patterns and characteristics of modern short-form video content that dominates social media platforms.… See the full description on the dataset page: https://huggingface.co/datasets/SFYuki/TikTok-10M.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/grimboy/tiktok-videos-4b.tiktok-trending-hashtags-music
TikTok Trending Hashtags and Music (2024 - 2025)
This release contains the top 100 daily trending hashtags and music records from TikTok Creative Center, covering the period from 2024-05-23 to 2025-07-09. It includes 13,399 unique hashtags and 11,157 unique songs.
The data comes from TikTok Creative Center:
https://ads.tiktok.com/business/creativecenter/inspiration/popular/hashtag/pc/en
The trending hashtag and music rankings are not personalized and are updated daily by TikTok.… See the full description on the dataset page: https://huggingface.co/datasets/lingbow/tiktok-trending-hashtags-music.TikTokDresstiktok-top-kol
