Octoparse/tiktok-brand-monitoring-beauty-sample
TikTok Beauty Brand Engagement Dataset 7 videos · 1,811 comments · 34 detected languages · 3.47M combined views A structured, GDPR-compliant sample dataset of TikTok video metadata and user comments across beauty and skincare brand accounts. Built to accelerate research in social media sentiment analysis, brand engagement modeling, and multilingual NLP for the beauty vertical. Produced by Octoparse Managed Data Service — enterprise web data pipelines for brand intelligence… See the full description on the dataset page: https://huggingface.co/datasets/Octoparse/tiktok-brand-monitoring-beauty-sample.
TikTok Beauty Brand Engagement Dataset
7 videos · 1,811 comments · 34 detected languages · 3.47M combined views
A structured, GDPR-compliant sample dataset of TikTok video metadata and user comments across beauty and skincare brand accounts. Built to accelerate research in social media sentiment analysis, brand engagement modeling, and multilingual NLP for the beauty vertical.
Produced by [Octoparse Managed Data Service](https://www.octoparse.com/data-service/social-media-monitoring) — enterprise web data pipelines for brand intelligence teams.
Enterprise Data Pipelines & Production-Grade Delivery
This dataset is a curated, static sample provided by [Octoparse Managed Data Service](https://www.octoparse.com/data-service). If your organization requires: - Real-time automated updates: Daily/Hourly feeds via API, Snowflake, AWS S3, or BigQuery - Custom schema alignment and production-grade anti-bot operations with 99.9% SLA-backed delivery - Compliance-aware data delivery workflows for public or properly authorized data, with support for GDPR/CCPA-aligned requirements [Request a Custom Data Pipeline Workshop on Octoparse Data Service](https://www.octoparse.com/data-service)
Why This Dataset Exists
Most TikTok datasets available publicly are either:
- Too small (a few hundred rows, single account)
- PII-contaminated (raw usernames, trackable comment authors)
- Structurally flat (video metadata and comments collapsed into one table, making joins painful)
This dataset solves all three problems:
- Normalized schema:
videos.parquet(7 rows) +comments.parquet(1,811 rows) joined onvideo_id - GDPR-compliant: commenter display names replaced with salted SHA-256 hashes (
commenter_id_hash), preserving cross-comment behavioral signals without PII exposure - Multilingual: comments span 34 detected languages (ISO 639-1 codes in
comment_lang)
Dataset Structure
tiktok-brand-monitoring-beauty-sample/
├── videos.parquet ← 7 rows, 17 columns (video-level metadata + metrics)
├── comments.parquet ← 1,811 rows, 8 columns (comment-level UGC + engagement)
├── schema.json ← Machine-readable field definitions and types
├── data_dictionary.md ← Business context for every field
├── workflow_stats.json ← Pipeline provenance and processing metadata
└── LICENSE ← CC BY-NC 4.0videos.parquet — Key Fields
comments.parquet — Key Fields
Quick Start
import pandas as pd
videos = pd.read_parquet("hf://datasets/Octoparse/tiktok-brand-monitoring-beauty-sample/videos.parquet")
comments = pd.read_parquet("hf://datasets/Octoparse/tiktok-brand-monitoring-beauty-sample/comments.parquet")
# Join: enrich comments with video context
df = comments.merge(videos[["video_id", "poster", "views_num", "post_date"]], on="video_id")
# Engagement rate per video
videos["engagement_rate"] = (
(videos["like_num"] + videos["comment_num"] + videos["forward_num"])
/ videos["views_num"]
)
# Top comments by influence score
comments["influence_score"] = (comments["comment_likes"] + 1).apply("log1p")
top_comments = comments.nlargest(20, "influence_score")[["comment_text", "comment_likes", "reply_num"]]
# English-only sentiment analysis subset
en_comments = comments[comments["comment_lang"] == "en"]["comment_text"]Dataset Stats
Language Distribution (Top 10)
Note on language detection:langdetectis probabilistic. Short comments (<5 tokens) or emoji-only text may be misclassified. Treatcomment_langas a feature hint, not ground truth. The high diversity of detected languages reflects Elemis's international brand presence (primarily UK/Europe).
Use Cases
1. Sentiment & Brand Perception Analysis
# Feed comment_text to your sentiment model
# Stratify by comment_likes for influence-weighted sentiment2. Engagement Prediction Modeling
Use like_num, views_num, bookmark_num, video_duration, and hashtags as features to predict comment_num or forward_num.
3. Influencer vs. Brand Account Comparison
The dataset includes both a beauty brand account (elemis) and individual content creators, enabling head-to-head engagement pattern analysis.
4. Comment Virality & Thread Analysis
reply_num + comment_likes identify conversation starters. commenter_id_hash is consistent across videos — trace cross-video superfans.
5. Multilingual NLP Benchmarking
Use comment_lang to slice multilingual subsets for cross-lingual transfer learning experiments.
Privacy & Compliance
This dataset was processed to comply with GDPR Article 25 (Privacy by Design):
- Commenter display names → replaced with salted SHA-256 hash (non-reversible). The same commenter maps to the same hash across all rows, preserving behavioral analysis without identity exposure.
- Signed download URLs (
video_download) → removed entirely. - Comment text is retained as-is: all data was publicly posted on TikTok and collected without authentication.
For dataset usage in research publications, we recommend referencing the commenter_id_hash field as a pseudonymized identifier.
Pipeline Provenance
This dataset was built and maintained by [Octoparse Managed Data Service](https://www.octoparse.com/data-service/social-media-monitoring).
Anti-bot systems bypassed during collection:
- TikTok DataDome JS challenge layer
- TikTok device fingerprinting (canvas/WebGL spoofing)
- Adaptive rate-limit detection
Processing stack: Python · pandas · langdetect · Apache Parquet (Snappy)
Full pipeline metadata is in `workflow_stats.json`.
Limitations
- Sample size: 7 videos is sufficient for prototyping; not for statistically robust brand-level conclusions.
- Point-in-time metrics: Engagement numbers (
like_num,views_num, etc.) reflect scrape-time snapshots, not live values. - Comment sampling:
videos.comment_numshows the true TikTok comment count;comments.parquetcontains a scraped sample (sampling rate varies by video: 26%–100%). - Language detection accuracy: Unreliable on comments <5 tokens.
Related Resources
- Service page: Social Media Monitoring Data Service
- Data Service hub + sample library: Octoparse Data Service
- Paired Kaggle dataset: tiktok-brand-monitoring-beauty-sample on Kaggle
Want Production-Scale TikTok Data?
This is a sample dataset. If you need:
- Daily / hourly comment ingestion pipelines
- APAC coverage (Xiaohongshu, Douyin, LINE)
- Competitor monitoring across dozens of brand accounts
- GDPR-compliant delivery to Snowflake, BigQuery, or S3
→ [Talk to Octoparse Managed Data Service](https://www.octoparse.com/data-service/social-media-monitoring)
We build and operate production web data pipelines for enterprise brand intelligence teams — zero ops overhead on your side.
Citation
@dataset{octoparse_tiktok_beauty_2025,
title = {TikTok Beauty Brand Engagement Dataset},
author = {{Octoparse Managed Data Service}},
year = {2025},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/Octoparse/tiktok-brand-monitoring-beauty-sample},
license = {CC BY-NC 4.0},
note = {GDPR-compliant sample of TikTok video metadata and comments, beauty vertical}
}Built by [Octoparse Managed Data Service](https://www.octoparse.com/data-service) · [LinkedIn](https://www.linkedin.com/company/octoparse)
