CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /laions_got_talent LAION's Got Talent: Generated Voice Acting Dataset Overview "LAION's Got Talent" is a generated dataset comprising voice acting samples that exhibit a wide range of emotions, vocal bursts, topics, and content. This dataset is a component of the BUD-E project, spearheaded by LAION with support from Intel. Dataset Composition The dataset includes: Emotional Diversity: Samples portraying various emotions to facilitate research in emotional recognition and… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent.audio100K<n<1M41 likes9.8k downloads2y agoHugging Face02laion /laions_got_talent_rawaudio10K<n<100K7 likes3.5k downloads2y agoHugging Face03laion /laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset Overview "LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants. Updated Composition Voices and Languages English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.audio1M<n<10M6 likes2.2k downloads1y agoHugging Face04laion /laions_got_talent_enhanced_no_metadataaudio10K<n<100K0 likes1.2k downloads2y agoHugging Face05mkrausio /laions_got_talent_embs_only laions_got_talent Whisper Embeddings (Embeddings + Metadata Only) This dataset contains Whisper embeddings (NPY) and metadata (JSON). The original audio files (MP3) are NOT included. Embeddings computed with: mkrausio/EmoWhisper-AnS-Small-v0.1 Includes original audio: No Includes metadata: Yes (JSON) Includes embeddings: Yes (NPY) Creation date: 2025-05-11 text1M<n<10M0 likes979 downloads1y agoHugging Face06laion /laions_got_talent_german_bicodecaudio100K<n<1M0 likes977 downloads2y agoHugging Face07LucasFang /Laion-Aesthetics-High-Resolution-GoT Laion-Aesthetics-High-Resolution-GoT Paper Dataset Description The Laion-Aesthetics-High-Resolution-GoT dataset is a collection of 3.77 million image-text pairs with rich grounding annotations. This dataset extends high-quality images from the LAION-Aesthetics collection with detailed text descriptions and object-level grounding information. Key Features Size: 3.77 million samples Modalities: Image, Text, and Grounding Annotations Image Resolution:… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/Laion-Aesthetics-High-Resolution-GoT.image1M<n<10M12 likes909 downloads2y agoHugging Face08GotThatData /kraken-trading-data 📈 Kraken Trading Data Collection Overview High-frequency cryptocurrency market data from Kraken exchange - perfect for algorithmic trading, time-series forecasting, and market microstructure analysis. This dataset includes real-time price, volume, and order book data for 9 major cryptocurrency trading pairs, collected via WebSocket streaming and REST API polling. 📊 Included Trading Pairs Pair Asset Base Currency Typical Daily Volume XXBTZUSD… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/kraken-trading-data.tabular10K<n<100K6 likes881 downloads8mo agoHugging Face09latent-lab /got-activations-llama3.1-405b-base meta-llama/Llama-3.1-405B — Activation Dataset Cached activations extracted from meta-llama/Llama-3.1-405B (revision unknown). Contents Tensor Layers Dim Pooling Shards Row Bytes hidden_layers 0-125 16384 - 12 - Prompts: 7660 Format version: 1.1 Load with lmprobe from lmprobe import pull_dataset, load_activation_dataset # Option 1: Pull into local cache (enables probe training without re-extraction)… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-llama3.1-405b-base.tabularfeature-extraction1K<n<10K0 likes857 downloads6mo agoHugging Face10latent-lab /got-activations-qwen2.5-0.5b Qwen/Qwen2.5-0.5B — Activation Dataset Cached activations extracted from Qwen/Qwen2.5-0.5B (revision 060db6499f32faf8b98477b0a26969ef7d8b9987). Full-sequence activations (24 layers, 896 dim, float16) and top-100 logits from Qwen/Qwen2.5-0.5B on 7,660 Geometry of Truth statements. Per-layer sharding (v1.2) with independent shard boundaries. Contents Tensor Layers Dim Pooling Shards Row Bytes hidden_layers 0-23 896 - 1 - logits_topk - k=100 last_token 1 1200… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-qwen2.5-0.5b.tabularfeature-extraction1K<n<10K0 likes570 downloads6mo agoHugging Face11laion /laions_got_talent_previewaudio1K<n<10K1 likes380 downloads2y agoHugging Face12LucasFang /OmniEdit-GoTtabular100K<n<1M3 likes350 downloads2y agoHugging Face13alliedtoasters /got-activations-llama3.1-70b-base meta-llama/Llama-3.1-70B — Activation Dataset Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac). Geometry of Truth curated dataset activations for Llama 3.1 70B base Contents Tensor Layers Dim Pooling Shards Row Bytes hidden_layers 0-79 8192 - 4 - Prompts: 7660 Format version: 2.0 Load with lmprobe from lmprobe import load_activations, Probe acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/got-activations-llama3.1-70b-base.tabularfeature-extraction1K<n<10K0 likes265 downloads6mo agoHugging Face14Riksarkivet /goteborgs_poliskammare_fore_1900_linesimage100K<n<1M1 likes220 downloads2y agoHugging Face15laion /laions_got_talent_orpheus_snacSome Laion's Got Talent (https://huggingface.co/datasets/laion/laions_got_talent) voice snippets converted to snac tokens in the format of the Orpheus-TTS https://github.com/canopyai/Orpheus-TTS We converted the data into instructions like format. The snac data is in 7 token frame groups. See the Orpheus blog for more details: https://canopylabs.ai/model-releases We did not create the original dataset and are only providing snac token with minimal text instructions for ease of use. You must be… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_orpheus_snac.text100K<n<1M0 likes179 downloads1y agoHugging Face16Gotech /cad-benchmark-leaderboardtextn<1K0 likes177 downloads15d agoHugging Face17Shuu12121 /go-treesitter-filtered-datasetsV2 Go CodeSearch Dataset (Shuu12121/go-treesitter-filtered-datasetsV2) Dataset Description This dataset contains Go functions and methods paired with their GoDoc comments, extracted from open-source Go repositories on GitHub. It is formatted similarly to the CodeSearchNet challenge dataset. Each entry includes: code: The source code of a go function or method. docstring: The docstring or Javadoc associated with the function/method. func_name: The name of the… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/go-treesitter-filtered-datasetsV2.text1M<n<10M0 likes174 downloads1y agoHugging Face1811-47 /Got_Agentic_AI_5k Got_Agentic_AI_5k A 5,000-example dataset to train LLMs into production-grade agentic assistants (“Angelic Agents”): high-agency, tool-aware, test-driven, and safety-first. This dataset focuses on the kinds of tasks real engineering teams and major AI developers care about: Diff-first coding patches and tests Planner–executor agent architectures Evals, monitoring, and rollback discipline Data engineering transforms with quality checks Incident postmortems and operational… See the full description on the dataset page: https://huggingface.co/datasets/11-47/Got_Agentic_AI_5k.text10K<n<100K5 likes108 downloads9mo agoHugging Face19fangsonglong /gothenburg-price-tagimageimage-to-textn<1K0 likes98 downloads2y agoHugging Face20rakshi719 /GOT10k-I2V GOT-10k Image-to-Video Retrieval MTEB/MOEB representation of the GOT-10k validation split for image-to-video retrieval. Task Given the first frame of a tracking sequence, retrieve the full tracking video. The mapping is one-to-one: each query has exactly one relevant item (the other direction of the same sequence). Contents Queries: 180 first-frame images Corpus: 180 tracking videos Qrels: 180 one-to-one binary relevance judgments… See the full description on the dataset page: https://huggingface.co/datasets/rakshi719/GOT10k-I2V.imagen<1K0 likes91 downloads3d agoHugging Face21rakshi719 /GOT10k-V2I GOT-10k Video-to-Image Retrieval MTEB/MOEB representation of the GOT-10k validation split for video-to-image retrieval. Task Given a tracking video, retrieve its corresponding first frame. The mapping is one-to-one: each query has exactly one relevant item (the other direction of the same sequence). Contents Queries: 180 tracking videos Corpus: 180 first-frame images Qrels: 180 one-to-one binary relevance judgments Source GOT-10k… See the full description on the dataset page: https://huggingface.co/datasets/rakshi719/GOT10k-V2I.imagen<1K0 likes89 downloads3d agoHugging Face22Vano04 /laions-got-talent-enhanced-precomputed-en LAION's Got Talent Enhanced Precomputed English This dataset contains the precomputed embeddings of the LAION's got talent enhanced dataset english split at 16kHz. The audio was preprocessed with TuKoResearch/AuriStream100M_RoPE_librilight and the text transcriptions were preprocessed with google/embeddinggemma-300m. @inproceedings{tuckute2025cochleartokens, title = {Representing Speech Through Autoregressive Prediction of Cochlear Tokens}, author = {Greta Tuckute and… See the full description on the dataset page: https://huggingface.co/datasets/Vano04/laions-got-talent-enhanced-precomputed-en.text1M<n<10M0 likes87 downloads10mo agoHugging Face23Shuu12121 /go-treesitter-dedupe_doc-filtered-dataset Go CodeSearch Dataset (Shuu12121/go-treesitter-dedupe_doc-filtered-dataset) Dataset Description This dataset contains Go functions and methods paired with their GoDoc comments, extracted from open-source Go repositories on GitHub. It is formatted similarly to the CodeSearchNet challenge dataset. Each entry includes: code: The source code of a go function or method. docstring: The docstring or Javadoc associated with the function/method. func_name: The name of the… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/go-treesitter-dedupe_doc-filtered-dataset.text1M<n<10M0 likes85 downloads1y agoHugging Face2411-47 /GOT_Defense_1 GOT_Defense_1 Professional pretraining corpus for defensive security LLMs. Rebuilt 2026-07-14. Records: 186,104 | Avg length: 377 chars | Dedup SHA256 | Split 95/5 This dataset merges 12 Kaggle sources into one high-quality text field optimized for causal LM pretraining: jeffborschowa/malwarebazaar-threat-intelligence-csv oriolakolawole/ransomware-and-goodware-pe-header joebeachcapital/tunadromd-malware-detection atharvasoundankar/global-cybersecurity-threats-2015-2024… See the full description on the dataset page: https://huggingface.co/datasets/11-47/GOT_Defense_1.texttext-generation100K<n<1M0 likes80 downloads2mo agoHugging Face25odoma /gotriple-pretraining-dataset GoTriple Pretraining dataset Summary The GoTriple Pre-training Dataset is a multilingual corpus built from open-access research artefacts harvested via the GoTriple platform. It focuses on Social Sciences and Humanities (SSH) content, addressing their limited presence in standard LLM pre-training corpora. Current release includes History, Sociology, Environmental Sciences, Psychology and Geography texts (~23.14B tokens). Intended Use Continuous… See the full description on the dataset page: https://huggingface.co/datasets/odoma/gotriple-pretraining-dataset.tabular100K<n<1M3 likes70 downloads8mo agoHugging Face26humair025 /laions-got-talent-annotatedtabular10K<n<100K0 likes67 downloads7mo agoHugging Face27gotime /VC-LLM-DatasetDue to the presence of harmful and toxic unsafe content in the fine-tuning data, a portion of the data is displayed. For the full data, please contact ignitesun@163.com. text1M<n<10M0 likes64 downloads2y agoHugging Face28GotThatData /warp_Research Warp Research Dataset Dataset Description Dataset Summary This dataset contains experimental results from warp field research, focusing on the relationship between warp factors, energy efficiency, and field characteristics. Supported Tasks Tabular Regression: Predict energy efficiency based on warp field parameters Time Series Forecasting: Analyze temporal patterns in warp field behavior Optimization: Identify optimal warp factor configurations for… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/warp_Research.imagetabular-regression10K<n<100K1 likes58 downloads8mo agoHugging Face2911-47 /GOT_Defense_2 GOT_Defense_2 Full - Per-file fallback Rebuilt 2026-07-14 with HF_HUB_DISABLE_XET=1, hf_xet removed, per-file skip on 403. Sources Fenrir 99k + Bouquets CVE (CVE-2021..2025, skips broken XET file) + WNT3D (keeps 6 files, skips massive_training if 403) Records 244,761 avg 1746 split 95/5 Usage load_dataset("11-47/GOT_Defense_2") text100K<n<1M0 likes49 downloads2mo agoHugging Face3011-47 /GOT_HQ_Merged_75k HQ-Merged-Dataset A high-quality merged dataset that samples 5,000 examples from each of 16 diverse sources every time it is run, normalises them into a unified schema, and pushes the result to Hugging Face. Quick Stats Metric Value Total examples 75,000 Number of sources 16 Sampling strategy Shuffle(seed=42) + take(5000) per source Generated 2026-05-30 00:12 UTC Category Breakdown Category Count reasoning 25,000… See the full description on the dataset page: https://huggingface.co/datasets/11-47/GOT_HQ_Merged_75k.text10K<n<100K1 likes45 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.