CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sunnypilot /sunnypilot_models_v1text1 likes21k downloads16h agoHugging Face02danish-foundation-models /danish-dynaword 🧨 Danish Dynaword Version 1.2.23 (Changelog) Language dan, dansk, Danish License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 7.40M Number of tokens (Llama 3): 9.81B Average document length in tokens (min, max): 1.33K (2, 19.46M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.imagetext-generation10M<n<100M22 likes11k downloads21d agoHugging Face03Goku-OpenLab /open-models-prompt-datasets 🖼️ Open Models Prompt Dataset 🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.image1K<n<10K1 likes6.4k downloads2mo agoHugging Face04goldfish-models /fish-food Goldfish Datasets These are the training datasets for the Goldfish models, as described in our paper, Goldfish: Monolingual Language Models for 350 Languages (Chang et al., 2026). Citation Along with citing the Goldfish paper, if using this dataset, we encourage researchers to cite the individual datasets listed in our paper. @inproceedings{chang-etal-2026-goldfish, title={Goldfish: Monolingual Language Models for 350 Languages}, author={Chang, Tyler A. and Arnett… See the full description on the dataset page: https://huggingface.co/datasets/goldfish-models/fish-food.text1B<n<10B2 likes5.2k downloads5mo agoHugging Face05TPPIsCriticalFor /colinear_scaling_models Collinear/Non-Collinear Scaling Models Checkpoint repository for scaling law experiments comparing collinear (CO) and non-collinear (NC) experimental designs for the paper Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation under review for NeurIPS 2026. Code Anonymized code repository (reproduces all tables): anonymous.4open.science Directory Structure {dataset}/{design}/N_{param_count}/ Dataset: wikipedia, pes2o, cosmopedia… See the full description on the dataset page: https://huggingface.co/datasets/TPPIsCriticalFor/colinear_scaling_models.tabularn<1K0 likes5k downloads5mo agoHugging Face06AI-C /rvc-modelsCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference imagen<1K1 likes2.8k downloads3y agoHugging Face07LoneResearch /explore-thinking-models-internaldocumentn<1K0 likes2.2k downloads3mo agoHugging Face08danish-foundation-models /norwegian-dynaword 🧨 Norwegian Dynaword Version 0.0.18 (Changelog) Language Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 4.47M Number of tokens (Llama 3): 9.98B Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.imagetext-generation10M<n<100M7 likes2k downloads15d agoHugging Face09burtenshaw /trending-models-top10-2026-03-06 Top 10 Trending Models (2026-03-06) This dataset records the top 10 trending models on the Hugging Face Hub captured on 2026-03-06. Files hf_trending_models_top10_2026-03-06.csv hf_trending_models_top10_2026-03-06.json Collection Method Collected with: hf models ls --sort trending_score --limit 10 Scores are point-in-time values and can change quickly. textn<1K0 likes1.9k downloads7mo agoHugging Face10danish-foundation-models /swedish-dynaword 🧨 Swedish Dynaword Version 0.0.13 (Changelog) Language Swedish (sv, swe) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 547.06M Number of tokens (Llama 3): 36.34B Average document length in tokens (min, max): 66.42 (2, 8.14M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.imagetext-generation1B<n<10B3 likes1.4k downloads14d agoHugging Face11jiangzhuo9357 /sherpa-onnx-tts-modelstextn<1K0 likes1.3k downloads6mo agoHugging Face12bkai-foundation-models /BKAINewsCorpus Dataset Card for "BKAINewsCorpus" The Binhvq News Corpus, a widely used dataset featuring approximately 20 million articles from diverse sources, received its last update in May 2021. To enhance this collection, we gathered an additional 10 million articles up until November 2023. By integrating these newly acquired articles with the existing Binhvq News Corpus, we have created an extensive Vietnamese News Corpus comprising about 32M articles. Subsequent fuzzy deduplication was… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/BKAINewsCorpus.text10M<n<100M14 likes1.1k downloads3y agoHugging Face13LLM-OS-Models /KoHRM-Text-1.4B-prepared-data KoHRM-Text-1.4B Prepared Data This dataset repository contains prepared HRM-Text V1Dataset artifacts for KoHRM-Text-1.4B. The data is intended for continued pretraining and staged training with the project code at: https://github.com/LLM-OS-Models/KoHRM-text https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B https://huggingface.co/LLM-OS-Models/HRM-Text-Ko-Terminal-Tokenizer-131K The upstream architecture and training method are based on: Paper:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-prepared-data.tabulartext-generationn<1K1 likes1k downloads4mo agoHugging Face14burtenshaw /hub-trending-models-2026-03-06tabularn<1K0 likes941 downloads7mo agoHugging Face15danish-foundation-models /dutch-dynaword 🧨 Dutch Dynaword Version 1.0.1 (Changelog) Language nld, Nederlands, Dutch License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 14.45M Number of tokens (Llama 3): 37.89B Average document length in tokens (min, max): 2.62K (2, 5.45M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.imagetext-generation10M<n<100M3 likes835 downloads13d agoHugging Face16danish-foundation-models /icelandic-dynaword 🧨 Icelandic Dynaword Version 0.0.15 (Changelog) Language Icelandic (is, isl) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 39.85M Number of tokens (Llama 3): 2.67B Average document length in tokens (min, max): 66.98 (3, 1.03M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.imagetext-generation100M<n<1B4 likes818 downloads15d agoHugging Face17danish-foundation-models /multilingual-gsm-symbolic Multilingual GSM-Symbolic Multilingual GSM-Symbolic is a benchmark for evaluating arithmetic reasoning in large language models across multiple languages. It extends Apple's GSM-Symbolic approach by providing symbolic templates that generate thousands of structurally equivalent but numerically distinct math problems. Templates and generation are handled by the multilingual-gsm-symbolic package. The dataset lets you test whether a model genuinely understands a problem or merely… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/multilingual-gsm-symbolic.text10K<n<100K3 likes730 downloads2mo agoHugging Face18Paul720810 /gguf-models GGUF Models Collection - 多格式版本 這個倉庫包含多種量化格式的 GGUF 模型檔案。 格式說明 格式 描述 品質 檔案大小 推薦用途 FP16 16位浮點 最高 最大 高精度推理、微調基準 Q8_0 8位量化 高 中等 高品質推理、伺服器部署 Q4_K_M 4位混合量化 良好 最小 本地部署、快速推理 轉換摘要 📊 成功轉換: 3/4 個模型 📈 成功率: 75.0% 🔧 支援格式: FP16, Q8_0, Q4_K_M 🕒 更新時間: 2025-08-29 04:34:43 模型列表 模型名 FP16 Q8_0 Q4_K_M 狀態 deepseek-1.3b-sql-final-t4x2 2569.5MB N/A N/A ✅ 成功 codegemma-2b-sql-coder-finetuned 4786.0MB N/A N/A ✅ 成功… See the full description on the dataset page: https://huggingface.co/datasets/Paul720810/gguf-models.textn<1K0 likes680 downloads1y agoHugging Face19ModelsLab /midashenglm-gen-training-latents ModelsLab/midashenglm-gen-training-latents Precomputed audio latents for fine-tuning mispeech/midashenglm-gen, paired with six-view prompts in the exact format the model was trained on. This is not an audio dataset and not a caption dataset. Each record is the output of the model's frozen DashengTokenizer encoder — 768-dimensional latents at 25 Hz, stored float16 — next to the tagged prompt string built from the source metadata. Why it exists The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.tabulartext-to-audion<1K0 likes634 downloads1mo agoHugging Face20danish-foundation-models /faroese-dynaword 🧨 Faroese Dynaword Version 0.0.7 (Changelog) Language Faroese (fo, fao) License Openly Licensed, See the respective dataset Models Currently there are no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 405.81K Number of tokens (Llama 3): 45.40M Average document length in tokens (min, max): 111.87 (2, 109.50K) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dynaword.imagetext-generation1M<n<10M3 likes588 downloads7d agoHugging Face21bkai-foundation-models /NewsSapoVietnamese NewsSapo Dataset The Vietnamese NewsSapo dataset was constructed to train sentence/passage embeddings. Our dataset is structured in a "title-abstract-contents" format, where each news article is represented by a tuple of (title, abstract, content). The content is the main text body of the article and has been processed to remove images, videos, and other non-textual elements. The dataset contains 31,728,183 triples. To build this dataset, we followed a two-step process: Step 1:… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/NewsSapo.textsummarization1M<n<10M6 likes563 downloads3y agoHugging Face22CGAxis /cgaxis-3d-models-sample CGAxis 3D Models - Free Sample (Furniture / Chairs) A free, licensed sample of human-authored 3D models from CGAxis, a 3D content studio operating since 2008. This sample is a taster of the full CGAxis AI Data corpus (4,390 3D models + 7,794 PBR material sets) available for commercial AI-training licenses. Every model ships as GLB and USDZ, with geometry statistics, real-world scale in centimetres, semantic tags, a natural-language caption, per-file SHA-256 and a… See the full description on the dataset page: https://huggingface.co/datasets/CGAxis/cgaxis-3d-models-sample.3dimage-to-3dn<1K0 likes546 downloads2mo agoHugging Face23Kuperberg /model-storage-v2textn<1K0 likes522 downloads9mo agoHugging Face24bkai-foundation-models /vi-alpaca 🇻🇳 Vietnamese Alpaca Dataset This dataset is especially designed for Vietnamese based on the idea from Stanford Alpaca and Self-Instruct paper. The motivation behind the creation of this dataset stems from the hope to contribute high-quality dataset to Vietnamese commnunity to train language models. To construct this dataset, we follow a two-step process: Step 1: Manually create Vietnamese seed tasks We employ the methodology outlined in the Self-Instruct paper we meticulously… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vi-alpaca.text10K<n<100K25 likes510 downloads3y agoHugging Face25barszot /3d-models-for-isaac-sim-dataset Dataset of 3D models for Isaac Sim (USDZ) 🇬🇧 English Description This dataset contains a collection of 3D models converted to the .usdz format, featuring proper Semantic Labeling. These assets are optimized for generating synthetic training data using NVIDIA Isaac Sim and NVIDIA Replicator. Primary Use Case: Training object detection and segmentation models (e.g., YOLO, RT-DETR, Mask R-CNN). Class List The dataset includes the following 30 semantic… See the full description on the dataset page: https://huggingface.co/datasets/barszot/3d-models-for-isaac-sim-dataset.3dn<1K0 likes505 downloads6mo agoHugging Face26hf-azure-internal /trending-models-analysishttps://github.com/pagezyhf/azure-cron/blob/main/trending_models_analysis.py text10K<n<100K3 likes481 downloads15m agoHugging Face27danish-foundation-models /danish-gigaword Danish Gigaword Corpus Version: 1.0.0 License: See the respective dataset Dataset Summary The Danish Gigaword Corpus contains text spanning several domains and forms. This version does not include the sections containing tweets ("General Discussions" and "Parliament Elections"), "danavis", "Common Crawl" and "OpenSubtitles" due to potential privacy, quality and copyright concerns. Loading the dataset from datasets import load_dataset name =… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-gigaword.texttext-generation100K<n<1M9 likes468 downloads2y agoHugging Face28CCB /cis5300-language-models CIS 5300 Language Models Dataset Dataset for Homework 3 of CIS 5300 (Natural Language Processing) at Penn. Cities config Country-of-origin classification over short city-name strings, drawn from nine countries (Afghanistan, China, Germany, Finland, France, India, Iran, Pakistan, South Africa). from datasets import load_dataset cities = load_dataset("CCB/cis5300-language-models", "cities") Split Rows Has labels? train 12,392 yes validation 1,548 yes test 1… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-language-models.text10K<n<100K0 likes463 downloads4mo agoHugging Face29Arrrlex /models-under-pressure Models Under Pressure This dataset accompanies the paper Detecting High-Stakes Interactions with Activation Probes, presented at the ICML 2025 Workshop on Actionable Interpretability, accepted to NeurIPS 2025. Overview Every sample is a user-facing LLM interaction labelled as high-stakes or low-stakes. The label reflects whether the conversation involves potentially consequential outcomes (medical advice, legal matters, financial decisions, etc.) vs. routine queries. The… See the full description on the dataset page: https://huggingface.co/datasets/Arrrlex/models-under-pressure.tabulartext-classification10K<n<100K0 likes445 downloads8mo agoHugging Face30danish-foundation-models /multi-ifeval MultiIFEval This dataset is an instruction-following dataset for 300+ languages, translated and localised from the English IFEval dataset. Dataset Details Dataset Description All samples come from the English IFEval dataset, and we translate and localise with Gemini-3-flash-preview. When translating and localising samples, we also include a random Wikipedia article in the target language, both to give some context for localisation, but also to… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/multi-ifeval.text100K<n<1M2 likes396 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.