CoolFace
Datasetpublic

Octoparse/tiktok-brand-monitoring-beauty-sample

TikTok Beauty Brand Engagement Dataset 7 videos · 1,811 comments · 34 detected languages · 3.47M combined views A structured, GDPR-compliant sample dataset of TikTok video metadata and user comments across beauty and skincare brand accounts. Built to accelerate research in social media sentiment analysis, brand engagement modeling, and multilingual NLP for the beauty vertical. Produced by Octoparse Managed Data Service — enterprise web data pipelines for brand intelligence… See the full description on the dataset page: https://huggingface.co/datasets/Octoparse/tiktok-brand-monitoring-beauty-sample.

sourceHugging Facecc-by-nc-4.0updated 4mo agoView on Hugging Face
0likes46downloads
Dataset Card

TikTok Beauty Brand Engagement Dataset

7 videos · 1,811 comments · 34 detected languages · 3.47M combined views

A structured, GDPR-compliant sample dataset of TikTok video metadata and user comments across beauty and skincare brand accounts. Built to accelerate research in social media sentiment analysis, brand engagement modeling, and multilingual NLP for the beauty vertical.

Produced by [Octoparse Managed Data Service](https://www.octoparse.com/data-service/social-media-monitoring) — enterprise web data pipelines for brand intelligence teams.

Enterprise Data Pipelines & Production-Grade Delivery

This dataset is a curated, static sample provided by [Octoparse Managed Data Service](https://www.octoparse.com/data-service). If your organization requires: - Real-time automated updates: Daily/Hourly feeds via API, Snowflake, AWS S3, or BigQuery - Custom schema alignment and production-grade anti-bot operations with 99.9% SLA-backed delivery - Compliance-aware data delivery workflows for public or properly authorized data, with support for GDPR/CCPA-aligned requirements [Request a Custom Data Pipeline Workshop on Octoparse Data Service](https://www.octoparse.com/data-service)

Why This Dataset Exists

Most TikTok datasets available publicly are either:

  • —Too small (a few hundred rows, single account)
  • —PII-contaminated (raw usernames, trackable comment authors)
  • —Structurally flat (video metadata and comments collapsed into one table, making joins painful)

This dataset solves all three problems:

  • —Normalized schema: videos.parquet (7 rows) + comments.parquet (1,811 rows) joined on video_id
  • —GDPR-compliant: commenter display names replaced with salted SHA-256 hashes (commenter_id_hash), preserving cross-comment behavioral signals without PII exposure
  • —Multilingual: comments span 34 detected languages (ISO 639-1 codes in comment_lang)

Dataset Structure

tiktok-brand-monitoring-beauty-sample/
├── videos.parquet          ← 7 rows, 17 columns (video-level metadata + metrics)
├── comments.parquet        ← 1,811 rows, 8 columns (comment-level UGC + engagement)
├── schema.json             ← Machine-readable field definitions and types
├── data_dictionary.md      ← Business context for every field
├── workflow_stats.json     ← Pipeline provenance and processing metadata
└── LICENSE                 ← CC BY-NC 4.0

videos.parquet — Key Fields

FieldTypeDescription
video_idstringTikTok video ID (19-digit, stored as string)
posterstringAccount handle
post_datedatetimeISO 8601 UTC
contentstringVideo caption text
hashtagslist[string]Parsed hashtag tokens
like_numint64Likes at scrape time
views_numint64Play count at scrape time
forward_numint64Share count
bookmark_numint64Save count (purchase intent proxy)
video_durationint64Seconds

comments.parquet — Key Fields

FieldTypeDescription
comment_idint64Sequential row ID
video_idstringFK → videos.video_id
commenter_id_hashstringSalted SHA-256 hash (16 hex chars). Same commenter = same hash.
comment_textstringRaw comment (UTF-8, emoji preserved)
comment_langstringISO 639-1 language code (langdetect)
comment_likesint64Likes on the comment
reply_numint64Replies to the comment
comment_datedatetimeISO 8601 UTC

Quick Start

python
import pandas as pd

videos   = pd.read_parquet("hf://datasets/Octoparse/tiktok-brand-monitoring-beauty-sample/videos.parquet")
comments = pd.read_parquet("hf://datasets/Octoparse/tiktok-brand-monitoring-beauty-sample/comments.parquet")

# Join: enrich comments with video context
df = comments.merge(videos[["video_id", "poster", "views_num", "post_date"]], on="video_id")

# Engagement rate per video
videos["engagement_rate"] = (
    (videos["like_num"] + videos["comment_num"] + videos["forward_num"])
    / videos["views_num"]
)

# Top comments by influence score
comments["influence_score"] = (comments["comment_likes"] + 1).apply("log1p")
top_comments = comments.nlargest(20, "influence_score")[["comment_text", "comment_likes", "reply_num"]]

# English-only sentiment analysis subset
en_comments = comments[comments["comment_lang"] == "en"]["comment_text"]

Dataset Stats

MetricValue
Videos7
Comments1,811
Languages detected34
Top language (en)834 comments (46%)
Combined video views3,473,014
Date range (posts)2022-08-31 → 2025-02-05
Date range (comments)2022-08-31 → 2025-03-23
Avg. comment likes10.9
Max comment likes9,212
Engagement rate range3.2% – 14.6%

Language Distribution (Top 10)

LanguageComments%
English (en)83446.0%
French (fr)553.0%
German (de)543.0%
Afrikaans (af)402.2%
Polish (pl)372.0%
Tagalog (tl)362.0%
Somali (so)362.0%
Welsh (cy)341.9%
Finnish (fi)311.7%
Norwegian (no)291.6%
Note on language detection: langdetect is probabilistic. Short comments (<5 tokens) or emoji-only text may be misclassified. Treat comment_lang as a feature hint, not ground truth. The high diversity of detected languages reflects Elemis's international brand presence (primarily UK/Europe).

Use Cases

1. Sentiment & Brand Perception Analysis

python
# Feed comment_text to your sentiment model
# Stratify by comment_likes for influence-weighted sentiment

2. Engagement Prediction Modeling

Use like_num, views_num, bookmark_num, video_duration, and hashtags as features to predict comment_num or forward_num.

3. Influencer vs. Brand Account Comparison

The dataset includes both a beauty brand account (elemis) and individual content creators, enabling head-to-head engagement pattern analysis.

4. Comment Virality & Thread Analysis

reply_num + comment_likes identify conversation starters. commenter_id_hash is consistent across videos — trace cross-video superfans.

5. Multilingual NLP Benchmarking

Use comment_lang to slice multilingual subsets for cross-lingual transfer learning experiments.


Privacy & Compliance

This dataset was processed to comply with GDPR Article 25 (Privacy by Design):

  • —Commenter display names → replaced with salted SHA-256 hash (non-reversible). The same commenter maps to the same hash across all rows, preserving behavioral analysis without identity exposure.
  • —Signed download URLs (video_download) → removed entirely.
  • —Comment text is retained as-is: all data was publicly posted on TikTok and collected without authentication.

For dataset usage in research publications, we recommend referencing the commenter_id_hash field as a pseudonymized identifier.


Pipeline Provenance

This dataset was built and maintained by [Octoparse Managed Data Service](https://www.octoparse.com/data-service/social-media-monitoring).

Anti-bot systems bypassed during collection:

  • —TikTok DataDome JS challenge layer
  • —TikTok device fingerprinting (canvas/WebGL spoofing)
  • —Adaptive rate-limit detection

Processing stack: Python · pandas · langdetect · Apache Parquet (Snappy)

Full pipeline metadata is in `workflow_stats.json`.


Limitations

  1. 1.Sample size: 7 videos is sufficient for prototyping; not for statistically robust brand-level conclusions.
  2. 2.Point-in-time metrics: Engagement numbers (like_num, views_num, etc.) reflect scrape-time snapshots, not live values.
  3. 3.Comment sampling: videos.comment_num shows the true TikTok comment count; comments.parquet contains a scraped sample (sampling rate varies by video: 26%–100%).
  4. 4.Language detection accuracy: Unreliable on comments <5 tokens.

Related Resources


Want Production-Scale TikTok Data?

This is a sample dataset. If you need:

  • —Daily / hourly comment ingestion pipelines
  • —APAC coverage (Xiaohongshu, Douyin, LINE)
  • —Competitor monitoring across dozens of brand accounts
  • —GDPR-compliant delivery to Snowflake, BigQuery, or S3

→ [Talk to Octoparse Managed Data Service](https://www.octoparse.com/data-service/social-media-monitoring)

We build and operate production web data pipelines for enterprise brand intelligence teams — zero ops overhead on your side.


Citation

bibtex
@dataset{octoparse_tiktok_beauty_2025,
  title        = {TikTok Beauty Brand Engagement Dataset},
  author       = {{Octoparse Managed Data Service}},
  year         = {2025},
  publisher    = {Hugging Face},
  url          = {https://huggingface.co/datasets/Octoparse/tiktok-brand-monitoring-beauty-sample},
  license      = {CC BY-NC 4.0},
  note         = {GDPR-compliant sample of TikTok video metadata and comments, beauty vertical}
}

Built by [Octoparse Managed Data Service](https://www.octoparse.com/data-service) · [LinkedIn](https://www.linkedin.com/company/octoparse)