datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sdxl-models
Aisha-AI.com 💜
A NSFW Social Network powered by AI Characters
The models saved in this dataset are currently being used, or have been used at some point, to generate images and videos.
The dataset is public and can be used as a backup or alternative to more unstable servers (like the unfortunate Civitai).
modelsmodels-logocolinear_scaling_models
license: gpl-2.0
Collinear Scaling Models
Checkpoint repository for scaling law experiments comparing collinear (CO) and non-collinear (NC) experimental designs.
Directory Structure
{dataset}/{design}/N_{param_count}/
Dataset: wikipedia, pes2o, cosmopedia, redpajama, c4 (plus _fp16 and _bigtpp variants)
Design: colinear or non_colinear
N: Model parameter count (one of 14 canonical sizes from ~5M to ~70M)
Experimental Designs
Collinear (CO):… See the full description on the dataset page: https://huggingface.co/datasets/leibnitz-lab/colinear_scaling_models.sunnypilot_models_v1mtpnet_image_models
模型训练过程汇总[该仓库只含有image model的训练过程]
本仓库采用扁平化的目录结构和标签系统来组织模型,具体说明如下:
仓库结构
一级目录:直接以模型名称-数据集,例如 ResNet-CIFAR-10、GraphMAE_QM9-Cora 等
二级目录:包含该模型在该数据集下的不同训练任务或变体,例如 normal、noisy、backdoor_invisible 等
训练过程目录结构:每个模型目录下包含:
scripts/:存放模型相关代码和训练脚本
epochs/:存放模型训练过程和权重文件
每个epoch的权重文件(model.pth)和embedding(.npy)
dataset/:模型需要的数据集
仓库结构展示
文件结构展示
danish-dynaword
🧨 Danish Dynaword
Version
1.2.23 (Changelog)
Language
dan, dansk, Danish
License
Openly Licensed, See the respective dataset
Models
For model trained used this data see danish-foundation-models
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 7.40M
Number of tokens (Llama 3): 9.81B
Average document length in tokens (min, max): 1.33K (2, 19.46M)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.wavelet-lstm-camels-models
Wavelet-LSTM CAMELS Streamflow Models
A collection of 61,380 pre-trained LSTM models for daily streamflow forecasting across 620 USGS catchments from the CAMELS dataset.
Each catchment has 99 independently trained models:
33 wavelet filters × 3 lead times (1, 3, 5 days) = 99 wavelet-enhanced models
33 matching baseline models (same architecture, no wavelet transform)
Models are designed to be ensembled across wavelets for robust predictions with uncertainty estimates.… See the full description on the dataset page: https://huggingface.co/datasets/johnswyou/wavelet-lstm-camels-models.hssd-models
HSSD: Habitat Synthetic Scenes Dataset
The Habitat Synthetic Scenes Dataset (HSSD) is a human-authored 3D scene dataset that more closely mirrors real scenes than prior datasets.
Our dataset represents real interiors and contains a diverse set of 211 scenes and more than 18000 models of real-world objects.
KLOM-modelsDataset for the evaluation of data-unlearning techniques using KLOM (KL-divergence of Margins).
How KLOM works:
KLOM works by:
training N models (original models)
Training N fully-retrained models (oracles) on forget set F
unlearning forget set F from the original models
Comparing the outputs of the unlearned models from the retrained models on different points
(specifically, computing the KL divergence between the distribution of margins of oracle models and distribution of… See the full description on the dataset page: https://huggingface.co/datasets/royrin/KLOM-models.goldsrc-models-datasetopen-models-prompt-datasets
🖼️ Open Models Prompt Dataset
🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.p2-etf-samba-modelsfish-food
Goldfish Datasets
These are the training datasets for the Goldfish models, as described in our paper, Goldfish: Monolingual Language Models for 350 Languages (Chang et al., 2026).
Citation
Along with citing the Goldfish paper, if using this dataset, we encourage researchers to cite the individual datasets listed in our paper.
@inproceedings{chang-etal-2026-goldfish,
title={Goldfish: Monolingual Language Models for 350 Languages},
author={Chang, Tyler A. and Arnett… See the full description on the dataset page: https://huggingface.co/datasets/goldfish-models/fish-food.colinear_scaling_models
Collinear/Non-Collinear Scaling Models
Checkpoint repository for scaling law experiments comparing collinear (CO) and non-collinear (NC) experimental designs for the paper Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation under review for NeurIPS 2026.
Code
Anonymized code repository (reproduces all tables): anonymous.4open.science
Directory Structure
{dataset}/{design}/N_{param_count}/
Dataset: wikipedia, pes2o, cosmopedia… See the full description on the dataset page: https://huggingface.co/datasets/TPPIsCriticalFor/colinear_scaling_models.sdxl-pony-models-backupDeepfacelive-DFM-Models
Description
Here you can find files for DeepFaceLab(It's back!) and DeepFaceLive. All sources and active community members are listed below.
Disclaimer
The author of this repository makes no claim to the data uploaded here other than that created by himself. Feel free to open a discussion for me to mention your contacts if I haven't done so.
Quick usage guide
To use the models presented in the repository, you will need installed DeepFaceLive.
.dfm models… See the full description on the dataset page: https://huggingface.co/datasets/dimanchkek/Deepfacelive-DFM-Models.bacpipe_modelscivit-ai-modelsrvc-modelsCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
flux-2-klein-modelsmodelsKoHRM-Text-1.4B-sft-lora-data
KoHRM-Text-1.4B SFT and LoRA Prepared Data
This dataset repo stores curated KoHRM SFT/LoRA subsets in the same tokenized
HRM-Text V1Dataset format used by training. It is intended for quick behavior
alignment experiments after KoHRM pretraining.
Model repo:
https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B
Code repo:
https://github.com/LLM-OS-Models/KoHRM-text
Format
Each folder is a prepared V1Dataset:
<dataset-name>/
metadata.json
tokenizer_info.json… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-sft-lora-data.explore-thinking-models-internalsim_modelsflux-dev-modelsherdsman_modelsnorwegian-dynaword
🧨 Norwegian Dynaword
Version
0.0.18 (Changelog)
Language
Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 4.47M
Number of tokens (Llama 3): 9.98B
Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.aizen-modelstrending-models-top10-2026-03-06
Top 10 Trending Models (2026-03-06)
This dataset records the top 10 trending models on the Hugging Face Hub captured on 2026-03-06.
Files
hf_trending_models_top10_2026-03-06.csv
hf_trending_models_top10_2026-03-06.json
Collection Method
Collected with:
hf models ls --sort trending_score --limit 10
Scores are point-in-time values and can change quickly.
