CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01asahi417 /seamless-align-enA-viA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes21k downloads2y agoHugging Face02asahi417 /seamless-align-deA-enA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes11k downloads2y agoHugging Face03asahi417 /seamless-align-enA-frA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes10k downloads2y agoHugging Face04asahi417 /seamless-align-enA-esA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes9.8k downloads2y agoHugging Face05asahi417 /seamless-align-enA-zhA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes8.5k downloads2y agoHugging Face06asahi417 /seamless-align-enA-jaA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes5.9k downloads2y agoHugging Face07asahi417 /seamless-align-enA-hiA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes5.9k downloads2y agoHugging Face08asahi417 /seamless-align-enA-koA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes3.5k downloads2y agoHugging Face09laion /relaion2B-en-research-safegatedimage1B<n<10B227 likes2.3k downloads2y agoHugging Face10andropar /relaion2b-natural-embeddings LAION-Natural Embeddings: CLIP ViT-H/14 Features for ~500M Natural Photographs (CCN 2025, Roth & Hebart) LAION-Natural Embeddings provides pre-computed CLIP ViT-H/14 embeddings for ~500 million natural photographs from ReLAION-2B, filtered using the LAION-Natural naturalness classifier (score > 0.7). Also known as: LAION-Natural Embeddings · ReLAION-Natural Embeddings · LAION-2B-Natural Embeddings Part of the LAION-Natural dataset family, introduced in: How to sample the… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural-embeddings.tabularfeature-extraction100M<n<1B1 likes1.5k downloads6mo agoHugging Face11laion /relaion2B-multi-research-safegatedimage1B<n<10B48 likes1.4k downloads2y agoHugging Face12laion /relaion2B-multi-researchgatedimage1B<n<10B12 likes1k downloads2y agoHugging Face13lyakaap /laion2B-japanese-subsetimage100M<n<1B4 likes848 downloads4y agoHugging Face14laion /relaion2B-en-researchgatedimage1B<n<10B45 likes823 downloads2y agoHugging Face15juiceb0xc0de /gemma-4-e2b-atlas image1M<n<10M4 likes800 downloads10d agoHugging Face16juiceb0xc0de /tmax-2b-atlas juiceb0xc0de/tmax-2b-atlas A brain atlas for allenai/tmax-2b, a hybrid SSM/Mamba/transformer language model. This is not a chat dataset or a benchmark — it is an internal-mechanics map of the model, built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing. If you want to know where the model stores compliance style, which late-layer directions you can edit without breaking reasoning, or whether the… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/tmax-2b-atlas.image1M<n<10M0 likes768 downloads10d agoHugging Face17LuciusLan /InfoSeek_emb_qwen3vle_2btabular10M<n<100M0 likes752 downloads5mo agoHugging Face18juiceb0xc0de /Qwen3.5-2B-Base juiceb0xc0de/Qwen3.5-2B-Base A brain atlas for Qwen/Qwen3.5-2B-Base, a 24-layer hybrid that runs linear attention on 18 layers and full attention on the other 6. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing. This is a base model, before any instruction tuning, so whatever structure shows up here was put there by… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Qwen3.5-2B-Base.imagefeature-extraction1M<n<10M0 likes536 downloads10d agoHugging Face19LeeHarrold /gemma-2b-dictionary-embeddings-all-layers Gemma-2B Dictionary Embeddings - All Layers This dataset contains pre-computed embeddings for 77,477 English words from WordNet using the Gemma-2B model across all 27 layers. Dataset Structure metadata.json: Contains dataset metadata (model info, dimensions, word count) embeddings_layer_X.pkl: Pickle files containing embeddings for layer X (0-26) Usage import pickle from huggingface_hub import hf_hub_download # Download a specific layer layer_0_path =… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/gemma-2b-dictionary-embeddings-all-layers.tabularn<1K0 likes527 downloads1y agoHugging Face20Ouroboros-Research /Our1-2b-Datasettabularn<1K0 likes494 downloads5d agoHugging Face21liangyuch /laion2b_seed Dataset Card for "laion2b_seed" This dataset is a subset of laion2B-en-aesthetic, with SEED v1 tokens. image100M<n<1B1 likes436 downloads3y agoHugging Face22laion /laion2B-en-aestheticgatedimage10M<n<100M72 likes393 downloads2y agoHugging Face23makermods /2_blue_cube_orange_box_20260803_104843This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/2_blue_cube_orange_box_20260803_104843.tabularrobotics10K<n<100K0 likes386 downloads2mo agoHugging Face24IDEA-CCNL /laion2B-multi-chinese-subset laion2B-multi-chinese-subset Github: Fengshenbang-LM Docs: Fengshenbang-Docs 简介 Brief Introduction 取自Laion2B多语言多模态数据集中的中文部分,一共143M个图文对。 A subset from Laion2B (a multimodal dataset), around 143M image-text pairs (only Chinese). 数据集信息 Dataset Information 大约一共143M个中文图文对。大约占用19GB空间(仅仅是url等文本信息,不包含图片)。 Homepage: laion-5b Huggingface: laion/laion2B-multi 下载 Download mkdir laion2b_chinese_release && cd laion2b_chinese_release for i in {00000..00012}; do… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-CCNL/laion2B-multi-chinese-subset.imagefeature-extraction10M<n<100M42 likes367 downloads3y agoHugging Face25llm-jp /relaion2B-en-research-safe-japanese-translation relaion2B-en-research-safe-japanese-translation This dataset is the Japanese translation of the English subset of ReLAION-5B (laion/relaion2B-en-research-safe), translated by gemma-2-9b-it. We used text2dataset for translating with open-weight LLMs. By leveraging the fast LLM inference library vLLM, this tool enables the rapid translation of large English datasets into Japanese. Prompt The following is the prompt used for translation with Gemma. You are an… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/relaion2B-en-research-safe-japanese-translation.image1B<n<10B4 likes359 downloads1y agoHugging Face26jasonhuang23 /laion2b_en_sd2.1baseA dataset for SD2.1 base training, which contains metadata filtered from laion2b-en with the following conditions. WIDTH>=512 HEIGHT>=512 punsafe<=0.98 AESTHETIC_SCORE>=4.5 image100M<n<1B9 likes352 downloads3y agoHugging Face27JWei05 /gemma4-e2b-base-topk128-hf-overlay-v128-seed42 Gemma 4 E2B base top-k-128 HF training overlay This is the immutable training-engine overlay used to distill traces from Gemma 4 E2B base into Gemma 4 E4B. It preserves the prompts, responses, and exact response token IDs from JWei05/gemma4-e2b-base-topk128-traces, but replaces the source vLLM top-k targets with targets recomputed by the Hugging Face training engine. This repository is a reproducibility artifact for the corresponding distillation run. It is not a new… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/gemma4-e2b-base-topk128-hf-overlay-v128-seed42.tabulartext-generation10K<n<100K0 likes312 downloads2mo agoHugging Face28lewington /laion2B-multi-joined-translated-to-en-smolimage10M<n<100M2 likes306 downloads2y agoHugging Face29egolimblevskaia /circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations CircuitLens & WeightLens: Transcoder Descriptions and Evaluations This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods. Methods CircuitLens: https://github.com/egolimblevskaia/CircuitLens WeightLens: https://github.com/egolimblevskaia/WeightLens Dataset Structure The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.tabulartext-classification10K<n<100K0 likes296 downloads7mo agoHugging Face30SlayerLab /gollem-corpus-2b-pl GoLLeM Corpus 2B PL Dokładny korpus treningowy polskiego modelu bazowego SlayerLab/GoLLeM-110M-PL-v3 (oraz v2) — ten sam zbiór, po którym model przeszedł dwie epoki. Publikujemy go, aby każdy mógł odtworzyć trening od zera na własnym tokenizerze. Jak powstał ten plik. Korpus odzyskano przez zdekodowanie stokenizowanego checkpointu treningowego (byte-level BPE dynaword-32k, round-trip bezstratny; granice dokumentów = token <|endoftext|>). To jest dokładnie tekst, który model… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/gollem-corpus-2b-pl.tabulartext-generation1M<n<10M1 likes287 downloads25d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.