datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sunnypilot_models_v1danish-dynaword
🧨 Danish Dynaword
Version
1.2.23 (Changelog)
Language
dan, dansk, Danish
License
Openly Licensed, See the respective dataset
Models
For model trained used this data see danish-foundation-models
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 7.40M
Number of tokens (Llama 3): 9.81B
Average document length in tokens (min, max): 1.33K (2, 19.46M)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.open-models-prompt-datasets
🖼️ Open Models Prompt Dataset
🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.fish-food
Goldfish Datasets
These are the training datasets for the Goldfish models, as described in our paper, Goldfish: Monolingual Language Models for 350 Languages (Chang et al., 2026).
Citation
Along with citing the Goldfish paper, if using this dataset, we encourage researchers to cite the individual datasets listed in our paper.
@inproceedings{chang-etal-2026-goldfish,
title={Goldfish: Monolingual Language Models for 350 Languages},
author={Chang, Tyler A. and Arnett… See the full description on the dataset page: https://huggingface.co/datasets/goldfish-models/fish-food.colinear_scaling_models
Collinear/Non-Collinear Scaling Models
Checkpoint repository for scaling law experiments comparing collinear (CO) and non-collinear (NC) experimental designs for the paper Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation under review for NeurIPS 2026.
Code
Anonymized code repository (reproduces all tables): anonymous.4open.science
Directory Structure
{dataset}/{design}/N_{param_count}/
Dataset: wikipedia, pes2o, cosmopedia… See the full description on the dataset page: https://huggingface.co/datasets/TPPIsCriticalFor/colinear_scaling_models.rvc-modelsCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
explore-thinking-models-internalnorwegian-dynaword
🧨 Norwegian Dynaword
Version
0.0.18 (Changelog)
Language
Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 4.47M
Number of tokens (Llama 3): 9.98B
Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.trending-models-top10-2026-03-06
Top 10 Trending Models (2026-03-06)
This dataset records the top 10 trending models on the Hugging Face Hub captured on 2026-03-06.
Files
hf_trending_models_top10_2026-03-06.csv
hf_trending_models_top10_2026-03-06.json
Collection Method
Collected with:
hf models ls --sort trending_score --limit 10
Scores are point-in-time values and can change quickly.
swedish-dynaword
🧨 Swedish Dynaword
Version
0.0.13 (Changelog)
Language
Swedish (sv, swe)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 547.06M
Number of tokens (Llama 3): 36.34B
Average document length in tokens (min, max): 66.42 (2, 8.14M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.sherpa-onnx-tts-modelsBKAINewsCorpus
Dataset Card for "BKAINewsCorpus"
The Binhvq News Corpus, a widely used dataset featuring approximately 20 million articles from diverse sources, received its last update in May 2021. To enhance this collection, we gathered an additional 10 million articles up until November 2023. By integrating these newly acquired articles with the existing Binhvq News Corpus, we have created an extensive Vietnamese News Corpus comprising about 32M articles. Subsequent fuzzy deduplication was… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/BKAINewsCorpus.KoHRM-Text-1.4B-prepared-data
KoHRM-Text-1.4B Prepared Data
This dataset repository contains prepared HRM-Text V1Dataset artifacts for KoHRM-Text-1.4B.
The data is intended for continued pretraining and staged training with the project code at:
https://github.com/LLM-OS-Models/KoHRM-text
https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B
https://huggingface.co/LLM-OS-Models/HRM-Text-Ko-Terminal-Tokenizer-131K
The upstream architecture and training method are based on:
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-prepared-data.hub-trending-models-2026-03-06dutch-dynaword
🧨 Dutch Dynaword
Version
1.0.1 (Changelog)
Language
nld, Nederlands, Dutch
License
Openly Licensed, See the respective dataset
Models
For model trained used this data see danish-foundation-models
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 14.45M
Number of tokens (Llama 3): 37.89B
Average document length in tokens (min, max): 2.62K (2, 5.45M)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.icelandic-dynaword
🧨 Icelandic Dynaword
Version
0.0.15 (Changelog)
Language
Icelandic (is, isl)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 39.85M
Number of tokens (Llama 3): 2.67B
Average document length in tokens (min, max): 66.98 (3, 1.03M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.multilingual-gsm-symbolic
Multilingual GSM-Symbolic
Multilingual GSM-Symbolic is a benchmark for evaluating arithmetic reasoning in large language models across multiple languages. It extends Apple's GSM-Symbolic approach by providing symbolic templates that generate thousands of structurally equivalent but numerically distinct math problems. Templates and generation are handled by the multilingual-gsm-symbolic package.
The dataset lets you test whether a model genuinely understands a problem or merely… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/multilingual-gsm-symbolic.gguf-models
GGUF Models Collection - 多格式版本
這個倉庫包含多種量化格式的 GGUF 模型檔案。
格式說明
格式
描述
品質
檔案大小
推薦用途
FP16
16位浮點
最高
最大
高精度推理、微調基準
Q8_0
8位量化
高
中等
高品質推理、伺服器部署
Q4_K_M
4位混合量化
良好
最小
本地部署、快速推理
轉換摘要
📊 成功轉換: 3/4 個模型
📈 成功率: 75.0%
🔧 支援格式: FP16, Q8_0, Q4_K_M
🕒 更新時間: 2025-08-29 04:34:43
模型列表
模型名
FP16
Q8_0
Q4_K_M
狀態
deepseek-1.3b-sql-final-t4x2
2569.5MB
N/A
N/A
✅ 成功
codegemma-2b-sql-coder-finetuned
4786.0MB
N/A
N/A
✅ 成功… See the full description on the dataset page: https://huggingface.co/datasets/Paul720810/gguf-models.midashenglm-gen-training-latents
ModelsLab/midashenglm-gen-training-latents
Precomputed audio latents for fine-tuning
mispeech/midashenglm-gen,
paired with six-view prompts in the exact format the model was trained on.
This is not an audio dataset and not a caption dataset. Each record is the
output of the model's frozen DashengTokenizer encoder — 768-dimensional latents
at 25 Hz, stored float16 — next to the tagged prompt string built from the
source metadata.
Why it exists
The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.faroese-dynaword
🧨 Faroese Dynaword
Version
0.0.7 (Changelog)
Language
Faroese (fo, fao)
License
Openly Licensed, See the respective dataset
Models
Currently there are no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 405.81K
Number of tokens (Llama 3): 45.40M
Average document length in tokens (min, max): 111.87 (2, 109.50K)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dynaword.NewsSapoVietnamese NewsSapo Dataset
The Vietnamese NewsSapo dataset was constructed to train sentence/passage embeddings. Our dataset is structured in a "title-abstract-contents" format, where each news article is represented by a tuple of (title, abstract, content). The content is the main text body of the article and has been processed to remove images, videos, and other non-textual elements. The dataset contains 31,728,183 triples.
To build this dataset, we followed a two-step process:
Step 1:… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/NewsSapo.cgaxis-3d-models-sample
CGAxis 3D Models - Free Sample (Furniture / Chairs)
A free, licensed sample of human-authored 3D models from CGAxis, a 3D content studio operating since 2008. This sample is a taster of the full CGAxis AI Data corpus (4,390 3D models + 7,794 PBR material sets) available for commercial AI-training licenses.
Every model ships as GLB and USDZ, with geometry statistics, real-world scale in centimetres, semantic tags, a natural-language caption, per-file SHA-256 and a… See the full description on the dataset page: https://huggingface.co/datasets/CGAxis/cgaxis-3d-models-sample.model-storage-v2vi-alpaca
🇻🇳 Vietnamese Alpaca Dataset
This dataset is especially designed for Vietnamese based on the idea from Stanford Alpaca and Self-Instruct paper. The motivation behind the creation of this dataset stems from the hope to contribute high-quality dataset to Vietnamese commnunity to train language models.
To construct this dataset, we follow a two-step process:
Step 1: Manually create Vietnamese seed tasks
We employ the methodology outlined in the Self-Instruct paper we meticulously… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vi-alpaca.3d-models-for-isaac-sim-dataset
Dataset of 3D models for Isaac Sim (USDZ)
🇬🇧 English Description
This dataset contains a collection of 3D models converted to the .usdz format, featuring proper Semantic Labeling. These assets are optimized for generating synthetic training data using NVIDIA Isaac Sim and NVIDIA Replicator.
Primary Use Case: Training object detection and segmentation models (e.g., YOLO, RT-DETR, Mask R-CNN).
Class List
The dataset includes the following 30 semantic… See the full description on the dataset page: https://huggingface.co/datasets/barszot/3d-models-for-isaac-sim-dataset.trending-models-analysishttps://github.com/pagezyhf/azure-cron/blob/main/trending_models_analysis.py
danish-gigaword
Danish Gigaword Corpus
Version: 1.0.0
License: See the respective dataset
Dataset Summary
The Danish Gigaword Corpus contains text spanning several domains and forms. This version does not include the sections containing tweets ("General Discussions" and "Parliament Elections"), "danavis", "Common Crawl" and "OpenSubtitles" due to potential privacy, quality and copyright concerns.
Loading the dataset
from datasets import load_dataset
name =… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-gigaword.cis5300-language-models
CIS 5300 Language Models Dataset
Dataset for Homework 3 of CIS 5300 (Natural Language Processing) at Penn.
Cities config
Country-of-origin classification over short city-name strings, drawn from
nine countries (Afghanistan, China, Germany, Finland, France, India, Iran,
Pakistan, South Africa).
from datasets import load_dataset
cities = load_dataset("CCB/cis5300-language-models", "cities")
Split
Rows
Has labels?
train
12,392
yes
validation
1,548
yes
test
1… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-language-models.models-under-pressure
Models Under Pressure
This dataset accompanies the paper Detecting High-Stakes Interactions with Activation Probes, presented at the ICML 2025 Workshop on Actionable Interpretability, accepted to NeurIPS 2025.
Overview
Every sample is a user-facing LLM interaction labelled as high-stakes or low-stakes. The label reflects whether the conversation involves potentially consequential outcomes (medical advice, legal matters, financial decisions, etc.) vs. routine queries.
The… See the full description on the dataset page: https://huggingface.co/datasets/Arrrlex/models-under-pressure.multi-ifeval
MultiIFEval
This dataset is an instruction-following dataset for 300+ languages, translated and localised from the English IFEval dataset.
Dataset Details
Dataset Description
All samples come from the English IFEval dataset, and we translate and localise with Gemini-3-flash-preview.
When translating and localising samples, we also include a random Wikipedia article in the target language, both to give some context for localisation, but also to… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/multi-ifeval.
