CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01juiceb0xc0de /Qwen3.5-4B-Base juiceb0xc0de/Qwen3.5-4B-Base A brain atlas for Qwen/Qwen3.5-4B-Base, a 32-layer hybrid that runs linear attention on 24 layers and full attention on the other 8. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing. This is a base model, before any instruction tuning, so whatever structure shows up here was put there by… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Qwen3.5-4B-Base.imagefeature-extraction1M<n<10M0 likes5.3k downloads11d agoHugging Face02maknee /wikipedia_qwen_4b Vector Database Dataset Generated embeddings dataset for vector database training and evaluation with multiple format support. Dataset Summary This dataset contains 1,000,000 text samples with high-quality vector embeddings generated using Qwen/Qwen3-Embedding-4B from the wikimedia/wikipedia dataset. The dataset is designed for vector database training, similarity search, and retrieval tasks. Dataset Structure Base dataset: 1,000,000 samples with embeddings… See the full description on the dataset page: https://huggingface.co/datasets/maknee/wikipedia_qwen_4b.textfeature-extraction1M<n<10M0 likes3.5k downloads7mo agoHugging Face03kwakuobeng /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/kwakuobeng/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes1.6k downloads16d agoHugging Face04blaccastro /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/blaccastro/tiktok-videos-4b.tabulartext-classification1B<n<10B2 likes1k downloads18d agoHugging Face05bihungba1101 /essay-vocab-range-qwen3.5-4b-trl-completions TRL Completion logs This dataset contains the completions generated during training using trl. Find the trained model at https://huggingface.co/bihungba1101/essay-vocab-range-qwen3.5-4b-grpo. The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument). Each file contains the following columns: step: the step of training prompt: the prompt used to generate the completion completion: the… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-vocab-range-qwen3.5-4b-trl-completions.tabular1K<n<10K0 likes953 downloads4mo agoHugging Face06dams2005 /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/dams2005/tiktok-videos-4b.tabulartext-classification1B<n<10B1 likes650 downloads24d agoHugging Face07bihungba1101 /grammar-accuracy-qwen3.5-4b-trl-completions TRL Completion logs This dataset contains the completions generated during training using trl. Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-grpo. The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument). Each file contains the following columns: step: the step of training prompt: the prompt used to generate the completion completion:… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-completions.tabular10K<n<100K1 likes632 downloads4mo agoHugging Face08merway /tiktok-videos-4b Mirror of kuben-developer/tiktok-videos-4b, snapshot 2026-09-08. All credit to the original author; same research-use license applies. TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research.… See the full description on the dataset page: https://huggingface.co/datasets/merway/tiktok-videos-4b.tabulartext-classification1B<n<10B2 likes629 downloads19d agoHugging Face09ftajwar /maxrl_qwen3_4B_base_polaris_rollouts MaxRL Qwen3-4B-Base training rollouts (POLARIS math prompts) Every training rollout from an online RL run, with exact token ids, sampling log-probs, and raw rewards — usable as a replay buffer to study off-policy RL for LLM reasoning completely offline. The run: Qwen3-4B-Base trained with the maxRL advantage estimator (A = (r - mean)/(mean + eps), group mean over 16 rollouts per prompt; maxRL paper) and a pure REINFORCE loss (L = -A * log pi; no importance ratio, no clipping, no… See the full description on the dataset page: https://huggingface.co/datasets/ftajwar/maxrl_qwen3_4B_base_polaris_rollouts.tabulartext-generation1M<n<10M0 likes625 downloads2mo agoHugging Face10bihungba1101 /essay-grammar-range-qwen3.5-4b-trl-completions TRL Completion logs This dataset contains the completions generated during training using trl. Find the trained model at https://huggingface.co/bihungba1101/essay-grammar-range-qwen3.5-4b-grpo. The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument). Each file contains the following columns: step: the step of training prompt: the prompt used to generate the completion completion: the… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-grammar-range-qwen3.5-4b-trl-completions.tabular1K<n<10K0 likes612 downloads4mo agoHugging Face11Moveo /hebrew_pretrain_v1_4baseData_processedtext1M<n<10M1 likes593 downloads2y agoHugging Face12KOM-00 /tiktok-videos-4b TikTok Videos: 4.5 billion posts dataset Step-by-step guide and access to the scraper code: tiktok-api.seeksocial.io. 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/KOM-00/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes592 downloads16d agoHugging Face13JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes585 downloads11d agoHugging Face14astro-legacy-archive /gwosc-o4b-strain GWOSC O4b 16 kHz strain Upload in progress. Verified span shards are being added while source spans finish downloading and conversion. This dataset contains the public Gravitational Wave Open Science Center O4b strain release at 16,384 Hz for H1, L1, V1. Each detector is represented independently. A contiguous source span produces three Parquet files: Strain, DQmask, and Injmask. Licence and acknowledgement Creative Commons Attribution 4.0 International Data… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gwosc-o4b-strain.time-series-forecasting100B<n<1T0 likes582 downloads6d agoHugging Face15ENSEONG /full-math-private-Qwen3-4B-Instruct-2507-bontabular100K<n<1M0 likes581 downloads7mo agoHugging Face16juiceb0xc0de /gemma-4-e4b-it-atlas juiceb0xc0de/gemma-4-e4b-it-atlas A brain atlas for google/gemma-4-E4B-it, the instruction-tuned E4B member of the Gemma 4 family. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing. If you want to know what sliding-window and full-attention layers actually do differently inside one model, how KV cache sharing splits a… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/gemma-4-e4b-it-atlas.imagefeature-extraction1M<n<10M0 likes573 downloads11d agoHugging Face17afrizalha /Indo4B-CombinedThis is the entire Indo4B dataset, combined into a single file. The original dataset can be found here: https://github.com/IndoNLP/indonlu This is a combination of all the different files in the compressed .tar.xz. The goal is so that anyone who's interested in Indonesian NLP can fairly simply load this dataset from huggingface, already combined in full. Note the original files consists of line-separated strings. This dataset just combines them while removing the available blank lines. text100M<n<1B0 likes569 downloads3y agoHugging Face18juiceb0xc0de /josie-2-4b-oss-atlas juiceb0xc0de/josie-2-4b-oss-atlas image1M<n<10M0 likes511 downloads11d agoHugging Face19JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes496 downloads20d agoHugging Face20Moveo /hebrew_pretrain_v1_4baseDatatext1M<n<10M0 likes476 downloads2y agoHugging Face21alex12223322 /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/alex12223322/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes448 downloads23d agoHugging Face22JackHsieh /4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a few dense, declarative sentences (about two to four) inside a <thought>…</thought> block, focused on the exact state at the cut and what the local grammar… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes413 downloads20d agoHugging Face23mrfakename /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/mrfakename/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes401 downloads20d agoHugging Face24JackHsieh /4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a few dense sentences that the generator wrapped in a literal <thought>…</thought> block (plain text, not special tokens). Only the part after </think> becomes the VALUE, with every <thought> / </thought> tag removed… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes388 downloads20d agoHugging Face25JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a single paragraph of dense reasoning about the immediate continuation. Only the part after </think> — the paragraph — becomes the VALUE. The reasoning inside the think block is dropped. The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes388 downloads20d agoHugging Face26seanphan /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/seanphan/tiktok-videos-4b.tabulartext-classification1B<n<10B0 likes387 downloads20d agoHugging Face27jacobmorrison /preference-test-qwen-30b-3a-4btext1M<n<10M0 likes383 downloads1y agoHugging Face28hojj /tiktok-videos-4b TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the largest public TikTok dataset I am aware of. It is released as-is, for research. What is in it 27 Parquet files, zstd compressed, about 289 GB in total. One row per video. Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/hojj/tiktok-videos-4b.tabulartext-classification1B<n<10B1 likes378 downloads23d agoHugging Face29bihungba1101 /essay-vocab-accuracy-qwen3.5-4b-trl-completions TRL Completion logs This dataset contains the completions generated during training using trl. Find the trained model at https://huggingface.co/bihungba1101/essay-vocab-accuracy-qwen3.5-4b-grpo. The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument). Each file contains the following columns: step: the step of training prompt: the prompt used to generate the completion completion:… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-vocab-accuracy-qwen3.5-4b-trl-completions.tabular1K<n<10K0 likes345 downloads4mo agoHugging Face30juiceb0xc0de /tmax-4b-atlas juiceb0xc0de/tmax-4b-atlas A brain atlas for allenai/tmax-4b, the mid-entry hybrid SSM/Mamba/transformer language model from the tmax family. This is not a chat dataset or a benchmark — it is an internal-mechanics map of the model, built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing. If you want to know where the model stores compliance style, which late-layer directions you can edit without… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/tmax-4b-atlas.imagefeature-extraction1M<n<10M0 likes341 downloads29d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.