CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01blaccastro /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/blaccastro/tiktok-videos-4b.tabulartext-classification1B<n<10B2 likes891 downloads14d agoHugging Face02kwakuobeng /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/kwakuobeng/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes637 downloads13d agoHugging Face03ftajwar /maxrl_qwen3_4B_base_polaris_rollouts MaxRL Qwen3-4B-Base training rollouts (POLARIS math prompts) Every training rollout from an online RL run, with exact token ids, sampling log-probs, and raw rewards — usable as a replay buffer to study off-policy RL for LLM reasoning completely offline. The run: Qwen3-4B-Base trained with the maxRL advantage estimator (A = (r - mean)/(mean + eps), group mean over 16 rollouts per prompt; maxRL paper) and a pure REINFORCE loss (L = -A * log pi; no importance ratio, no clipping, no… See the full description on the dataset page: https://huggingface.co/datasets/ftajwar/maxrl_qwen3_4B_base_polaris_rollouts.tabulartext-generation1M<n<10M0 likes597 downloads2mo agoHugging Face04dams2005 /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/dams2005/tiktok-videos-4b.tabulartext-classification1B<n<10B1 likes585 downloads21d agoHugging Face05KOM-00 /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/KOM-00/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes537 downloads13d agoHugging Face06JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes518 downloads8d agoHugging Face07alex12223322 /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/alex12223322/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes373 downloads20d agoHugging Face08susun-123 /perfectblend-qwen3-4b-regen perfectblend-qwen3-4b-regen Single-turn SFT-style corpus for training speculative-decoding draft models (DFlash/MTP-style) against Qwen/Qwen3-4B as the target. Prompts come from an open-perfectblend-derived blend; every assistant response was regenerated by Qwen3-4B itself, so the token distribution matches the target model exactly. The sampled output_token_ids are included, letting trainers supervise on the target's own decode without re-tokenization drift.… See the full description on the dataset page: https://huggingface.co/datasets/susun-123/perfectblend-qwen3-4b-regen.text-generation1M<n<10M0 likes371 downloads27d agoHugging Face09hojj /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/hojj/tiktok-videos-4b.tabulartext-classification1B<n<10B1 likes366 downloads20d agoHugging Face10open-llm-leaderboard-old /details_Ramikan-BR__tinyllama_PY-CODER-4bit-lora_4k-v12 Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Ramikan-BR__tinyllama_PY-CODER-4bit-lora_4k-v12.text-generation10K<n<100K0 likes365 downloads2y agoHugging Face11mrfakename /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/mrfakename/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes326 downloads17d agoHugging Face12seanphan /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/seanphan/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes321 downloads17d agoHugging Face13hamishivi /qwen35-4b-drpo-vs0f49th-trainer-logprobs Qwen3.5 4B DRPO trainer logprobs from W&B run vs0f49th This dataset contains the raw trainer-logprob JSONL shards saved by W&B run ai2-llm/open_instruct_internal/vs0f49th (qwen35_4b_drpo__42__1782345587). Contents Source run: https://wandb.ai/ai2-llm/open_instruct_internal/runs/vs0f49th Source path: /weka/oe-adapt-default/allennlp/deletable_rollouts/ Filename pattern: qwen35_4b_drpo__42__1782345587_trainer_logprobs_step*_rank*.jsonl Files: 4320 JSONL shards… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/qwen35-4b-drpo-vs0f49th-trainer-logprobs.tabulartext-generation10K<n<100K0 likes281 downloads3mo agoHugging Face14kkndlee /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/kkndlee/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes274 downloads20d agoHugging Face15JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/4B-reason-only.rule-r-1.0-k-8.L-512.statml-arxiv. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: <last 8 prefix tokens> VALUE: <thought> <|/note|> and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes253 downloads8d agoHugging Face16sk7725 /tv4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/sk7725/tv4b.tabulartext-classification1B<n<10B0 likes237 downloads19d agoHugging Face17CausalNLP /gpt2-training-ar-zh-ko-ja-4b Balanced Arabic-Chinese-Korean-Japanese 4B-token training data Sequential FineWeb-2 documents tokenized with CausalNLP/gpt2-tokenizer-ar-zh-ko-ja. Each language contains at least 1,000,000,000 tokens in complete documents. Total target: 4,000,000,000 tokens. Approximately 100,000,000 tokens per Parquet shard. Splits: arb_Arab, cmn_Hani, kor_Hang, jpn_Jpan. Schema: text: string, input_ids: list<int32>. texttext-generation1M<n<10M0 likes202 downloads2mo agoHugging Face18abdellatifinformation /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/abdellatifinformation/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes200 downloads15d agoHugging Face19merway /tiktok-videos-4b Mirror of kuben-developer/tiktok-videos-4b, snapshot 2026-09-08. All credit to the original author; same research-use license applies. TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research.… See the full description on the dataset page: https://huggingface.co/datasets/merway/tiktok-videos-4b.tabulartext-classification1B<n<10B1 likes195 downloads15d agoHugging Face20ansulev /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes189 downloads18d agoHugging Face21alice1001 /open-perfectblend-qwen3-4b-regen open-perfectblend-qwen3-4b-regen This dataset contains 1,339,649 complete conversations from mlabonne/open-perfectblend with assistant responses regenerated by Qwen3-4B. It is published as one default/train split and does not expose internal source subdivisions. Schema id (string): contiguous dataset row ID from 0. conversations (list): complete ShareGPT messages with from set to human or gpt and the original text in value. source (string): always… See the full description on the dataset page: https://huggingface.co/datasets/alice1001/open-perfectblend-qwen3-4b-regen.texttext-generation1M<n<10M0 likes184 downloads1mo agoHugging Face22taufiqdp /Indo4B-hftexttext-generation100M<n<1B3 likes181 downloads3y agoHugging Face23marin-community /openthoughts4-code-9168-prompts-qwen3-4b-n16-flattened-logprobs-k16 OpenThoughts-4 Code SDG: Qwen3-4B (n=16, top-16 logprobs) Synthetic generations from Qwen/Qwen3-4B on the Marin OpenThoughts-4 code SDG prompt set. Each prompt is sampled n=16 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model Qwen/Qwen3-4B… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-4b-n16-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes168 downloads5mo agoHugging Face24TIE-Pilot /qwen3-4b-perfectblend-deepspec-rollout Qwen3-4B PerfectBlend DeepSpec Rollout This dataset contains the complete DeepSpec-aligned Qwen3-4B self-distillation rollout over the filtered PerfectBlend corpus. The seeded 95/5 split is published as separate train and eval splits. Splits Split Conversations Shards Path train 1,349,860 128 data/*.jsonl eval 71,046 64 eval/*.jsonl total 1,420,906 192 Data construction Canonical filtered corpus: 1,420,906 conversations. Split:… See the full description on the dataset page: https://huggingface.co/datasets/TIE-Pilot/qwen3-4b-perfectblend-deepspec-rollout.texttext-generation1M<n<10M0 likes164 downloads1d agoHugging Face25bobboyms /subset-Itau-Unibanco-aroeira-4B-tokens Subset Corpus Itau-Unibanco/aroeira: 1B tokens (portuguese PT-BR) Subset Corpus Itau-Unibanco/aroeira: 1B tokens (portuguese PT-BR) subset-Itau-Unibanco-aroeira-1B-tokens texttext-generation10M<n<100M1 likes156 downloads1y agoHugging Face26cds-jb /cot-gemma4-26b-a4b Gemma-4-26B-A4B-it Chain-of-Thought Oracle Corpus Chain-of-thought rollouts generated with google/gemma-4-26B-A4B-it (MoE, 25.2B total / 3.8B active), in its native thinking mode, across a diverse suite of reasoning tasks. Structure follows ceselder/cot-oracle-corpus-v5 (CoT-only subset of the columns), built for chain-of-thought monitoring / activation-oracle research. 2,121,354 rollouts over 212,161 unique problems (10 sampled thinking rollouts per problem, temperature 0.8).… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-gemma4-26b-a4b.tabulartext-generation1M<n<10M0 likes155 downloads3mo agoHugging Face27xlr8harder /synthid-qwen3-4b-instruct-2507-wildchat Qwen3-4B SynthID three-arm corpus This export contains aligned unwatermarked, SynthID key-A, and SynthID key-B responses from Qwen/Qwen3-4B-Instruct-2507. Matched splits share prompts and request seeds across configurations; unmatched splits use mutually disjoint prompt pools. Export complete for its source work queue: true. Generation profile Model revision: cdbee75f17c01a7cc42f958dc650907174af0554 Native model dtype: bfloat16 Maximum generated tokens: 4096… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/synthid-qwen3-4b-instruct-2507-wildchat.tabulartext-generation100K<n<1M0 likes154 downloads1mo agoHugging Face28xlr8harder /trellismark-qwen3-4b TrellisMark Qwen3-4B confirmation corpus This is the frozen English confirmation corpus for TrellisMark, an experimental many-user AI-text watermark. It includes exact generated text and token IDs, unwatermarked Qwen controls, public-key detector evidence, the public research key, independent encoder vectors, and the result reports used for the reader-facing curves. The standalone implementation, detector, and reproduction instructions are in the TrellisMark GitHub repository.… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/trellismark-qwen3-4b.imagetext-generation100K<n<1M0 likes153 downloads1mo agoHugging Face29cds-jb /synthweb-gemma4-26b-a4b Gemma-4-26B-A4B FineWeb Rollouts (~580k docs) Open-ended continuations of FineWeb (sample-10BT) document prefixes, generated by google/gemma-4-26b-a4b (the base, non-it Gemma-4 26B-A4B mixture-of-experts model), then mode-collapse filtered. This is the Gemma-4 analogue of cds-jb/qwen3-8b-fineweb-rollouts-100k: a "synthweb" corpus of natural model-generated documents, intended as the substrate for activation-oracle / interpretability probing (extract a base model's residual… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/synthweb-gemma4-26b-a4b.tabulartext-generation1M<n<10M0 likes149 downloads3mo agoHugging Face30marin-community /openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-4B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-4B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes147 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.