CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01danish-foundation-models /danish-dynaword 🧨 Danish Dynaword Version 1.2.23 (Changelog) Language dan, dansk, Danish License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 7.40M Number of tokens (Llama 3): 9.81B Average document length in tokens (min, max): 1.33K (2, 19.46M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.imagetext-generation10M<n<100M22 likes11k downloads22d agoHugging Face02danish-foundation-models /norwegian-dynaword 🧨 Norwegian Dynaword Version 0.0.18 (Changelog) Language Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 4.47M Number of tokens (Llama 3): 9.98B Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.imagetext-generation10M<n<100M7 likes1.9k downloads17d agoHugging Face03danish-foundation-models /swedish-dynaword 🧨 Swedish Dynaword Version 0.0.13 (Changelog) Language Swedish (sv, swe) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 547.06M Number of tokens (Llama 3): 36.34B Average document length in tokens (min, max): 66.42 (2, 8.14M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.imagetext-generation1B<n<10B3 likes1.4k downloads15d agoHugging Face04LLM-OS-Models /KoHRM-Text-1.4B-prepared-data KoHRM-Text-1.4B Prepared Data This dataset repository contains prepared HRM-Text V1Dataset artifacts for KoHRM-Text-1.4B. The data is intended for continued pretraining and staged training with the project code at: https://github.com/LLM-OS-Models/KoHRM-text https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B https://huggingface.co/LLM-OS-Models/HRM-Text-Ko-Terminal-Tokenizer-131K The upstream architecture and training method are based on: Paper:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-prepared-data.tabulartext-generationn<1K1 likes1k downloads4mo agoHugging Face05danish-foundation-models /dutch-dynaword 🧨 Dutch Dynaword Version 1.0.1 (Changelog) Language nld, Nederlands, Dutch License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 14.45M Number of tokens (Llama 3): 37.89B Average document length in tokens (min, max): 2.62K (2, 5.45M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.imagetext-generation10M<n<100M3 likes841 downloads14d agoHugging Face06danish-foundation-models /icelandic-dynaword 🧨 Icelandic Dynaword Version 0.0.15 (Changelog) Language Icelandic (is, isl) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 39.85M Number of tokens (Llama 3): 2.67B Average document length in tokens (min, max): 66.98 (3, 1.03M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.imagetext-generation100M<n<1B4 likes799 downloads17d agoHugging Face07danish-foundation-models /faroese-dynaword 🧨 Faroese Dynaword Version 0.0.7 (Changelog) Language Faroese (fo, fao) License Openly Licensed, See the respective dataset Models Currently there are no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 405.81K Number of tokens (Llama 3): 45.40M Average document length in tokens (min, max): 111.87 (2, 109.50K) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dynaword.imagetext-generation1M<n<10M3 likes656 downloads8d agoHugging Face08danish-foundation-models /danish-gigaword Danish Gigaword Corpus Version: 1.0.0 License: See the respective dataset Dataset Summary The Danish Gigaword Corpus contains text spanning several domains and forms. This version does not include the sections containing tweets ("General Discussions" and "Parliament Elections"), "danavis", "Common Crawl" and "OpenSubtitles" due to potential privacy, quality and copyright concerns. Loading the dataset from datasets import load_dataset name =… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-gigaword.texttext-generation100K<n<1M9 likes461 downloads2y agoHugging Face09danish-foundation-models /norwegian-dyna-instruct 🧨 Norwegian dyna-instruct Version 0.1.0 (changelog) Languages Norwegian Bokmål (nob), Norwegian Nynorsk (nno), and English (eng) translation input License Mixed open licenses; see the table below Sources Five datasets (source cards) Dataset Description Number of samples: 14.40K Number of tokens (Llama 3): 6.27M Average conversation length in tokens (min, max): 435.63 (4, 8.92K) Average number of turns (min, max): 2.13 (2, 3)… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dyna-instruct.imagequestion-answering10K<n<100K0 likes239 downloads17d agoHugging Face10danish-foundation-models /faroese-dyna-instruct 🧨 Faroese dyna-instruct Version 0.1.0 (Changelog) Language Faroese (fao) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 8.61K Number of tokens (Llama 3): 2.64M Average conversation length in tokens (min, max): 306.67 (98, 1.24K) Average number of… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dyna-instruct.texttext-generation10K<n<100K1 likes223 downloads22d agoHugging Face11danish-foundation-models /ifeval-da IFEval-da This dataset is a translation of the English IFEval dataset, which was published in this paper and contains 541 prompts, each with a combination of one or more of 25 different constraints. The dataset was professionally translated and localised by expert native speakers. Dataset Details Translated by: Rasmus Larsen (rasmus.larsen@alexandra.dk), Nathalie Hau Sørensen (naha@hum.ku.dk) and Kenneth Enevoldsen (kenneth.enevoldsen@cas.au.dk) Funded by: Danish… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ifeval-da.texttext-generationn<1K1 likes208 downloads7mo agoHugging Face12danish-foundation-models /icelandic-dyna-instruct 🧨 Icelandic dyna-instruct Version 0.1.0 (Changelog) Language Icelandic (isl) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 8.11K Number of tokens (Llama 3): 7.09M Average conversation length in tokens (min, max): 874.89 (182, 1.39K) Average number… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dyna-instruct.texttext-generation10K<n<100K1 likes162 downloads22d agoHugging Face13matonski /toy-models-of-sft-data Toy Models of SFT Data This is a public-clean candidate data package for the Toy Models of SFT project. It is built for researcher inspection first. The package answers two questions: What were the models trained on? How did the models actually behave under evaluation? The package includes training data, eval inputs, model rollouts, judge scores, parsed GPQA outputs, aggregate tables, paper figures, frozen plot data, and provenance records. It deliberately includes some… See the full description on the dataset page: https://huggingface.co/datasets/matonski/toy-models-of-sft-data.tabulartext-generation10K<n<100K0 likes156 downloads2mo agoHugging Face14adsgpt /marketing-benchmark-of-more-than-10-ai-models Marketing Benchmark of 10+ AI Models A 5,000-question benchmark for evaluating LLMs across six dimensions of modern marketing — Meta Ads, Google Ads, SEO & Organic, Email & Lifecycle, Critical Thinking, and Action-Based scenarios — graded through 10 distinct marketer personas. Every question is independently authored by the AdsGPT Marketing Bench team. Knowledge MCQs are hand-authored against 2026 platform documentation; open-ended and action-based scenarios are built from… See the full description on the dataset page: https://huggingface.co/datasets/adsgpt/marketing-benchmark-of-more-than-10-ai-models.textquestion-answering1K<n<10K5 likes122 downloads4mo agoHugging Face15beatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K3 likes107 downloads1mo agoHugging Face16ahmedBargady /open-models-benchmark-results ⚡ Local LLM Evaluation Leaderboard Welcome to the official public benchmark leaderboard maintained by @ahmedBargady.This dataset repository hosts benchmark evaluation metrics, accuracy scores, throughput telemetry, and quantization trade-off analyses of open-weights foundation models tested locally on NVIDIA A100 GPUs. 💻 Hardware & System Specifications All evaluations are executed under standardized local cluster environments: Specification Details… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/open-models-benchmark-results.tabulartext-generationn<1K1 likes86 downloads1mo agoHugging Face17danish-foundation-models /nasjonalt-vitenarkiv Nasjonalt vitenarkiv Open-access documents from NVA (Nasjonalt vitenarkiv), the joint national repository where Norwegian research institutions publish their output: master's and PhD theses, journal articles, and technical and research reports. Subjects span the disciplines - marine science, forestry, archaeology, education, public health, engineering - and most documents are recent. Each row is one PDF: the original file exactly as published, the text extracted from it, and the… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/nasjonalt-vitenarkiv.documenttext-generationn<1K1 likes79 downloads2mo agoHugging Face18mondk /12.Models.Ask.Themselves.Q_AThe data files were generated through conversations with various models, notably: Claude Sonnet/Opus/Fable, ChatGPT 5.5, Solar Pro 4, Gemini 3.6 Flash, DeepSeek v4 Flash, MiniMax M3, Kimi k2.5, GLM 4.5, Qwen 3.8 Max, and Inkling Small/Medium. Thank you for reading. Please leave a like. texttext-generationn<1K5 likes75 downloads2mo agoHugging Face19Neura-parse /quantum-machine-learning-models Neura Parse — Quantum Machine Learning Models: Encodings, Kernels, QNNs & Generative/Deep Architectures A hands-on, code-first vertical on quantum models that learn from data. Spans data encodings/feature maps, variational classifiers, quantum kernels/QSVMs, and quantum neural networks through modern generative and deep architectures (quantum GANs, circuit Born machines, quantum Boltzmann machines, QCNNs, quantum autoencoders, quantum RL, and quantum… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-machine-learning-models.tabulartext-generation100K<n<1M1 likes64 downloads3mo agoHugging Face20AI-Sweden-Models /Dolci-Instruct-SFT-translated Dolci-Instruct-SFT-translated (Swedish) This dataset is a Swedish machine translation of the openeurollm/Dolci-Instruct-SFT-translated dataset, originally created as part of the OpenEuroLLM project. Dataset details Examples: 494,841 multi-turn conversations Language: Swedish (sv-SE) Format: Chat/messages format (id, messages) License: Apache 2.0 Translation All English source texts were machine-translated to Swedish using Google Gemma 3 27B-IT (w8a8_fp8… See the full description on the dataset page: https://huggingface.co/datasets/AI-Sweden-Models/Dolci-Instruct-SFT-translated.texttext-generation100K<n<1M0 likes60 downloads6mo agoHugging Face21matCercola18 /quotient-margins-reward-models Quotient Margins for Reward Models — data release Artifacts backing the paper Measure Confidence on Decisions, Not Samples: Quotient Margins for Reward Models. The short version of the paper. Reward models pick the best of N sampled responses, but their confidence is normally read off the reward gap between the top two samples. When several candidates express the same underlying behaviour, that gap is a within-class spacing and its predictive signal cancels. Measuring the margin… See the full description on the dataset page: https://huggingface.co/datasets/matCercola18/quotient-margins-reward-models.texttext-generation0 likes38 downloads9h agoHugging Face22lumen-models /aec-rag-dataset Lumen-Models: AEC-RAG Dataset Lumen-Models is the premier conversational dataset designed to fine-tune LLMs and empower RAG (Retrieval-Augmented Generation) systems within the Architecture, Engineering, and Construction (AEC) sector. This dataset features high-fidelity technical dialogues between a BIM Auditor and a GPT Expert, focused on solving real-world challenges regarding regulatory compliance, complex construction codes, and professional industry standards. Premium… See the full description on the dataset page: https://huggingface.co/datasets/lumen-models/aec-rag-dataset.texttext-generationn<1K1 likes33 downloads3mo agoHugging Face23danish-foundation-models /laerebogengated Lærebogen An instruction-following dataset for Danish. This dataset features 5 million examples of multi-turn conversations in Danish, designed to train instruction-following models, with a commercially usable license. Dataset Structure All examples in the dataset are structured as follows: { "messages": [ { "role": "user", "content": "(...)" }, { "role": "assistant", "content": "(...)" }, { "role": "user", "content": "(...)" }, (...) { "role":… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/laerebogen.texttext-generation1M<n<10M1 likes28 downloads6mo agoHugging Face24exp-models /s1K-1.1-Koreanhttps://huggingface.co/datasets/simplescaling/s1K-1.1 texttext-generation1K<n<10K3 likes26 downloads2y agoHugging Face25fineset-io /protein-language-models-papers Protein Language Models Papers — FineSet A research-paper dataset on Protein Language Models Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Protein Language Models Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/protein-language-models-papers.tabulartext-classificationn<1K0 likes24 downloads3mo agoHugging Face26small-models-for-glam /synthetic-aat-materials Synthetic AAT Materials Dataset Dataset Description This dataset contains 1000 synthetic examples of cultural heritage object descriptions paired with their materials as they would appear in the Getty Art & Architecture Thesaurus (AAT). The data is formatted for training conversational AI models, particularly Qwen3, to identify and extract materials from cultural heritage object descriptions. Dataset Structure Each example contains: messages: Conversation… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/synthetic-aat-materials.texttext-generation1K<n<10K0 likes22 downloads1y agoHugging Face27Pidoxy /Blind_Spots_of_Frontier_Models Qwen3-0.6B-Base — Blind Spots Dataset Model Tested Qwen/Qwen3-0.6B-Base Type: Causal Language Model (base / pretraining only — not instruction-tuned) Parameters: 0.6B (0.44B non-embedding) Released: April–May 2025 by Alibaba Cloud's Qwen Team Context Length: 32,768 tokens How the Model Was Loaded The model was loaded in a Google Colab T4 GPU notebook using HuggingFace transformers >= 4.51.0(required because the qwen3 architecture key was added in… See the full description on the dataset page: https://huggingface.co/datasets/Pidoxy/Blind_Spots_of_Frontier_Models.texttext-generationn<1K0 likes22 downloads7mo agoHugging Face28danish-foundation-models /dfm-dyna-instructgated 🧨 DFM dyna-instruct Version 0.1.3 (Changelog) Language Danish (dan), English (eng), French (fra), German (deu), Italian (ita) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 4.40M Number of tokens (Llama 3): 2.85B Average conversation length in tokens… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dfm-dyna-instruct.imagetext-generation1M<n<10M4 likes22 downloads4mo agoHugging Face29build-small-hackathon /agenda-parser-models-example-agent-traces Agenda Parser — fine-tuned agent models Three Gemma 4 models fine-tuned to drive the Agenda Parser's ReAct agent: at each step the model emits a single JSON action {"thought","tool","args"} over two toolkits — meeting-agenda packets and Michigan local-government law (Open Meetings Act, FOIA, the Michigan Compiled Laws via Cornell LII). This card doubles as the project write-up; the dataset itself (bottom) is a gallery of example traces from the three models. tier base… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-models-example-agent-traces.texttext-generationn<1K0 likes22 downloads4mo agoHugging Face30fineset-io /time-series-foundation-models-papers Time Series Foundation Models Papers — FineSet A research-paper dataset on Time Series Foundation Models Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Time Series Foundation Models Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/time-series-foundation-models-papers.tabulartext-classificationn<1K0 likes22 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.