CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kwakuobeng /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/kwakuobeng/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes1.6k downloads16d agoHugging Face02blaccastro /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/blaccastro/tiktok-videos-4b.tabulartext-classification1B<n<10B2 likes1k downloads18d agoHugging Face03dams2005 /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/dams2005/tiktok-videos-4b.tabulartext-classification1B<n<10B1 likes650 downloads24d agoHugging Face04merway /tiktok-videos-4b Mirror of kuben-developer/tiktok-videos-4b, snapshot 2026-09-08. All credit to the original author; same research-use license applies. TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research.… See the full description on the dataset page: https://huggingface.co/datasets/merway/tiktok-videos-4b.tabulartext-classification1B<n<10B2 likes629 downloads19d agoHugging Face05ftajwar /maxrl_qwen3_4B_base_polaris_rollouts MaxRL Qwen3-4B-Base training rollouts (POLARIS math prompts) Every training rollout from an online RL run, with exact token ids, sampling log-probs, and raw rewards — usable as a replay buffer to study off-policy RL for LLM reasoning completely offline. The run: Qwen3-4B-Base trained with the maxRL advantage estimator (A = (r - mean)/(mean + eps), group mean over 16 rollouts per prompt; maxRL paper) and a pure REINFORCE loss (L = -A * log pi; no importance ratio, no clipping, no… See the full description on the dataset page: https://huggingface.co/datasets/ftajwar/maxrl_qwen3_4B_base_polaris_rollouts.tabulartext-generation1M<n<10M0 likes625 downloads2mo agoHugging Face06TIE-Pilot /qwen3-4b-perfectblend-deepspec-rollout Qwen3-4B PerfectBlend DeepSpec Rollout This dataset contains the complete DeepSpec-aligned Qwen3-4B self-distillation rollout over the filtered PerfectBlend corpus. The seeded 95/5 split is published as separate train and eval splits. Splits Split Conversations Shards Path train 1,349,860 128 data/*.jsonl eval 71,046 64 eval/*.jsonl total 1,420,906 192 Data construction Canonical filtered corpus: 1,420,906 conversations. Split:… See the full description on the dataset page: https://huggingface.co/datasets/TIE-Pilot/qwen3-4b-perfectblend-deepspec-rollout.texttext-generation1M<n<10M0 likes612 downloads5d agoHugging Face07KOM-00 /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/KOM-00/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes592 downloads16d agoHugging Face08JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes585 downloads11d agoHugging Face09alex12223322 /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/alex12223322/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes448 downloads24d agoHugging Face10mrfakename /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/mrfakename/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes401 downloads20d agoHugging Face11seanphan /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/seanphan/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes387 downloads20d agoHugging Face12hojj /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/hojj/tiktok-videos-4b.tabulartext-classification1B<n<10B1 likes378 downloads23d agoHugging Face13kkndlee /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/kkndlee/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes276 downloads23d agoHugging Face14ansulev /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes268 downloads21d agoHugging Face15sk7725 /tv4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/sk7725/tv4b.tabulartext-classification1B<n<10B0 likes241 downloads23d agoHugging Face16CausalNLP /gpt2-training-ar-zh-ko-ja-4b Balanced Arabic-Chinese-Korean-Japanese 4B-token training data Sequential FineWeb-2 documents tokenized with CausalNLP/gpt2-tokenizer-ar-zh-ko-ja. Each language contains at least 1,000,000,000 tokens in complete documents. Total target: 4,000,000,000 tokens. Approximately 100,000,000 tokens per Parquet shard. Splits: arb_Arab, cmn_Hani, kor_Hang, jpn_Jpan. Schema: text: string, input_ids: list<int32>. texttext-generation1M<n<10M0 likes229 downloads3mo agoHugging Face17abdellatifinformation /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/abdellatifinformation/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes203 downloads19d agoHugging Face18JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/4B-reason-only.rule-r-1.0-k-8.L-512.statml-arxiv. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: <last 8 prefix tokens> VALUE: <thought> <|/note|> and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes194 downloads12d agoHugging Face19taufiqdp /Indo4B-hftexttext-generation100M<n<1B3 likes193 downloads3y agoHugging Face20alice1001 /open-perfectblend-qwen3-4b-regen open-perfectblend-qwen3-4b-regen This dataset contains 1,339,649 complete conversations from mlabonne/open-perfectblend with assistant responses regenerated by Qwen3-4B. It is published as one default/train split and does not expose internal source subdivisions. Schema id (string): contiguous dataset row ID from 0. conversations (list): complete ShareGPT messages with from set to human or gpt and the original text in value. source (string): always… See the full description on the dataset page: https://huggingface.co/datasets/alice1001/open-perfectblend-qwen3-4b-regen.texttext-generation1M<n<10M0 likes175 downloads2mo agoHugging Face21marin-community /openthoughts4-code-9168-prompts-qwen3-4b-n16-flattened-logprobs-k16 OpenThoughts-4 Code SDG: Qwen3-4B (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-4B on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model Qwen/Qwen3-4B… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-4b-n16-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes161 downloads5mo agoHugging Face22bobboyms /subset-Itau-Unibanco-aroeira-4B-tokens Subset Corpus Itau-Unibanco/aroeira: 1B tokens (portuguese PT-BR) Subset Corpus Itau-Unibanco/aroeira: 1B tokens (portuguese PT-BR) subset-Itau-Unibanco-aroeira-1B-tokens texttext-generation10M<n<100M1 likes157 downloads1y agoHugging Face23JackHsieh /4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176 — Qwen3-4B-Instruct-2507 after RL against a frozen suffix conditional. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes142 downloads3d agoHugging Face24cds-jb /synthweb-gemma4-26b-a4b Gemma-4-26B-A4B FineWeb Rollouts (~580k docs) Open-ended continuations of FineWeb (sample-10BT) document prefixes, generated by google/gemma-4-26b-a4b (the base, non-it Gemma-4 26B-A4B mixture-of-experts model), then mode-collapse filtered. This is the Gemma-4 analogue of cds-jb/qwen3-8b-fineweb-rollouts-100k: a "synthweb" corpus of natural model-generated documents, intended as the substrate for activation-oracle / interpretability probing (extract a base model's residual… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/synthweb-gemma4-26b-a4b.tabulartext-generation1M<n<10M0 likes141 downloads3mo agoHugging Face25grimboy /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/grimboy/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes138 downloads23d agoHugging Face26cds-jb /cot-gemma4-26b-a4b Gemma-4-26B-A4B-it Chain-of-Thought Oracle Corpus Chain-of-thought rollouts generated with google/gemma-4-26B-A4B-it (MoE, 25.2B total / 3.8B active), in its native thinking mode, across a diverse suite of reasoning tasks. Structure follows ceselder/cot-oracle-corpus-v5 (CoT-only subset of the columns), built for chain-of-thought monitoring / activation-oracle research. 2,121,354 rollouts over 212,161 unique problems (10 sampled thinking rollouts per problem, temperature 0.8).… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-gemma4-26b-a4b.tabulartext-generation1M<n<10M0 likes137 downloads3mo agoHugging Face27Thinking-Space /OpenThought3-Qwen3-4BOpenThought3-Qwen3-4B OpenThought3-Qwen3-4B is a math reasoning supervised fine-tuning dataset in chat-message JSONL format. Data Creation and Cleaning This dataset was generated by Qwen3-4B (Non-thinking) from math-domain prompts selected from OpenThoughts3-1.2M. The generated responses were cleaned through deduplication, removal of degenerate repetition/repeater-style outputs, and template checks on the assistant… See the full description on the dataset page: https://huggingface.co/datasets/Thinking-Space/OpenThought3-Qwen3-4B.texttext-generation100K<n<1M3 likes117 downloads5mo agoHugging Face28twnlp /ChineseErrorCorrector4-4B ChineseErrorCorrect4 Data This dataset is associated with the paper CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards. It contains 340,000 Chain-of-Thought (CoT) reasoning samples designed for Chinese Grammatical Error Correction (CGEC) and Chinese Spelling Correction (CSC). These samples provide explicit error reasoning for diagnostic transparency, helping models internalize linguistic priors and improve edit… See the full description on the dataset page: https://huggingface.co/datasets/twnlp/ChineseErrorCorrector4-4B.texttext-generation100K<n<1M2 likes106 downloads4mo agoHugging Face29YangyiH /qwen3-4b-teacher-rollouts-76k-nonthinking Qwen3-4B Teacher Rollouts 76K Non-Thinking This dataset contains 76,800 fixed teacher trajectories generated for a prompt-aligned reproduction study of on-policy distillation with Qwen3-1.7B. It is an independent research artifact, not an official release from the model or paper authors. Models and generation Teacher: Qwen/Qwen3-4B-Instruct-2507 Tokenizer/chat template: Qwen/Qwen3-1.7B Mode: non-thinking (enable_thinking=False) Temperature: 0.7 Top-p: 1.0 Top-k:… See the full description on the dataset page: https://huggingface.co/datasets/YangyiH/qwen3-4b-teacher-rollouts-76k-nonthinking.tabulartext-generation10K<n<100K0 likes103 downloads2mo agoHugging Face30JackHsieh /4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids, the thoughts written by the prestar-RL policy Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes89 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.