datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trending-models-top10-2026-03-06
Top 10 Trending Models (2026-03-06)
This dataset records the top 10 trending models on the Hugging Face Hub captured on 2026-03-06.
Files
hf_trending_models_top10_2026-03-06.csv
hf_trending_models_top10_2026-03-06.json
Collection Method
Collected with:
hf models ls --sort trending_score --limit 10
Scores are point-in-time values and can change quickly.
KoHRM-Text-1.4B-prepared-data
KoHRM-Text-1.4B Prepared Data
This dataset repository contains prepared HRM-Text V1Dataset artifacts for KoHRM-Text-1.4B.
The data is intended for continued pretraining and staged training with the project code at:
https://github.com/LLM-OS-Models/KoHRM-text
https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B
https://huggingface.co/LLM-OS-Models/HRM-Text-Ko-Terminal-Tokenizer-131K
The upstream architecture and training method are based on:
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-prepared-data.midashenglm-gen-training-latents
ModelsLab/midashenglm-gen-training-latents
Precomputed audio latents for fine-tuning
mispeech/midashenglm-gen,
paired with six-view prompts in the exact format the model was trained on.
This is not an audio dataset and not a caption dataset. Each record is the
output of the model's frozen DashengTokenizer encoder — 768-dimensional latents
at 25 Hz, stored float16 — next to the tagged prompt string built from the
source metadata.
Why it exists
The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.models-under-pressure
Models Under Pressure
This dataset accompanies the paper Detecting High-Stakes Interactions with Activation Probes, presented at the ICML 2025 Workshop on Actionable Interpretability, accepted to NeurIPS 2025.
Overview
Every sample is a user-facing LLM interaction labelled as high-stakes or low-stakes. The label reflects whether the conversation involves potentially consequential outcomes (medical advice, legal matters, financial decisions, etc.) vs. routine queries.
The… See the full description on the dataset page: https://huggingface.co/datasets/Arrrlex/models-under-pressure.co2_modelsblender-3d-models
Blender 3D Models Database
This repository serves as a database of custom-made 3D models generated programmatically using Blender and uploaded to Hugging Face.
Models Included
⚔️ Stylized Fantasy Sword (stylized_sword.glb) - View Model Page
🛡️ Stylized Fantasy Shield (stylized_shield.glb) - View Model Page
🪄 Stylized Wizard's Staff (stylized_staff.glb) - View Model Page
📦 Stylized Treasure Chest (stylized_chest.glb) - View Model Page
🏹 Stylized Bow… See the full description on the dataset page: https://huggingface.co/datasets/abersbail/blender-3d-models.MMMU-Reasoning-Distill-Validation中文版本
Description
MMMU-Reasoning-Distill-Validation is a Multi-Modal reasoning dataset that contains 839 image descriptions and natural language inference data samples. This dataset is built upon the validation set of MMMU. The construction process begins with using Qwen2.5-VL-72B-Instruct for image understanding and generating detailed image descriptions, followed by generating reasoning conversations using the DeepSeek-R1 model. Its main features are as follows:
Use the… See the full description on the dataset page: https://huggingface.co/datasets/modelscope/MMMU-Reasoning-Distill-Validation.ai-models-2026
AI Models & Releases 2026
AI model releases, benchmarks, capabilities. Updated daily via automated collection pipeline.
Part of the Legion Data Factory — historical AI ecosystem datasets 2026.
Methodology
Automated collection from public sources (HackerNews, RSS feeds, APIs).
Updated daily via cron job. Raw data, minimal processing.
License
CC BY 4.0
🔑 API Access — Updated Daily
Live data via Legion AI API | Documentation
Free: 100… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-models-2026.gspc-own-models-measured
Our own models, measured and excluded from leads
We measure everyone on the board — including our own fine-tunes. This dataset is the
measurement record for the in-house clan-* and council-* models the Council of AI (CSOAI)
trained and then ran through the same GSPC board as every other system. No exception, no
hand-scoring, no hiding the weak rows.
18 in-house fine-tunes, 50 signed measurement cards.
Every card is Ed25519-signed by CSOAI Ltd (UK 16939677) and re-verifiable.… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-own-models-measured.ParserV1-modelsReAPR-Automatic-Program-Repair-via-Retrieval-Augmented-Large-Language-ModelsThis is the Retrieval dataset used in the paper "ReAPR: Automatic Program Repair via Retrieval-Augmented Large Language Models"
toy-models-of-sft-data
Toy Models of SFT Data
This is a public-clean candidate data package for the Toy Models of SFT project.
It is built for researcher inspection first.
The package answers two questions:
What were the models trained on?
How did the models actually behave under evaluation?
The package includes training data, eval inputs, model rollouts, judge scores,
parsed GPQA outputs, aggregate tables, paper figures, frozen plot data, and
provenance records. It deliberately includes some… See the full description on the dataset page: https://huggingface.co/datasets/matonski/toy-models-of-sft-data.self-cognition
介绍(Introduction)
该自我认知数据集由modelsope swift创建, 可以通过将通配符进行替换:{{NAME}}、{{AUTHOER}},来创建属于自己大模型的自我认知数据集,总共108条。
ms-swift github:https://github.com/modelscope/swift/
自我认知微调最佳实践文档:https://github.com/modelscope/swift/blob/main/docs/source/LLM/%E8%87%AA%E6%88%91%E8%AE%A4%E7%9F%A5%E5%BE%AE%E8%B0%83%E6%9C%80%E4%BD%B3%E5%AE%9E%E8%B7%B5.md
This self-cognition dataset was created by modelsope swift and can be customized for your own large model by replacing the placeholders: {{NAME}} and… See the full description on the dataset page: https://huggingface.co/datasets/modelscope/self-cognition.marketing-benchmark-of-more-than-10-ai-models
Marketing Benchmark of 10+ AI Models
A 5,000-question benchmark for evaluating LLMs across six dimensions of modern
marketing — Meta Ads, Google Ads, SEO & Organic, Email & Lifecycle, Critical
Thinking, and Action-Based scenarios — graded through 10 distinct marketer personas.
Every question is independently authored by the AdsGPT Marketing Bench team.
Knowledge MCQs are hand-authored against 2026 platform documentation; open-ended
and action-based scenarios are built from… See the full description on the dataset page: https://huggingface.co/datasets/adsgpt/marketing-benchmark-of-more-than-10-ai-models.korean-embedding-performance-v1-performance-1m
Korean Embedding Performance v1 — 1M
Qwen3-Embedding-8B의 한국어 retrieval data-scale 실험을 위한 정확히
1,000,000-row 연구·비상업 contrastive dataset이다. release_eligible: false, 통합
라이선스 other이며 upstream source 조건을 재허가하지 않는다.
구성
계열
Rows
비율
역할
nlpai-lab/ko-triplet-v1.0
600,254
60.03%
넓은 한국어 QA/retrieval core
F2 Korean QA/instruction
287,000
28.70%
webfaq, mqa, koalpaca, realQA, komagpie
F2 retrieval task train-family
4,146
0.41%
MIRACL, MrTidy, MLDR
F2 PAWS-X… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-performance-1m.filtered_models_swe_smithkorean-embedding-performance-v1-ablation-200k
Korean Embedding Performance v1 — Ablation 200K
Qwen3-Embedding-8B의 한국어 retrieval continued fine-tuning에서 LoRA/DoRA/부분 및
full fine-tuning, loss, hard-negative 전략을 비교하기 위한 200,000-row 연구·비상업
성능 데이터다. release_eligible: false이며 통합 라이선스는 other다. upstream
source별 조건을 재허가하지 않는다.
구성
계열
Rows
역할
nlpai-lab/ko-triplet-v1.0@1f5d72d
100,254
넓은 한국어 QA/retrieval core
F2 Korean QA/instruction
68,000
webfaq, mqa, koalpaca, realQA, komagpie
F2 retrieval task… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-ablation-200k.korean-embedding-performance-v1-pilot-50k
Korean Embedding Performance v1 — Pilot 50K
주의: 이 revision은 공개 benchmark 성능 후보 학습에 사용하면 안 된다.
사후 15-task exact text-hash 감사에서 평가 query 고유 hash 4개가 확인됐다.
파이프라인·최적화 진단과 contamination ablation에만 남기며, 교체본은
ablation-200k이다.
Qwen3-Embedding 계열의 한국어 retrieval 성능 실험을 위한 50,000-row 연구용
contrastive dataset이다. 각 row는 instruction-aware query, positive passage 1개,
hard/easy negative passage 1–7개를 ms-swift embedding message schema로 저장한다.
사용 조건과 공개 범위
이 저장소의 통합 라이선스는 other다.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-pilot-50k.DeepResonance_data_models
DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning
Paper: arxiv
This is a repository of data and models for DeepResonance.
Data
For all the existing datasets, download the multimodal resources according to the original papers.
For Music4way related datasets, first download all the video and music files using the YouTube IDs shown in each dataset. Then referring to M2UGen's pipeline to randomly extract an image from… See the full description on the dataset page: https://huggingface.co/datasets/Sony/DeepResonance_data_models.repro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces
Agent traces
Agent sessions published from a Trackio Logbook.
12.Models.Ask.Themselves.Q_AThe data files were generated through conversations with various models, notably: Claude Sonnet/Opus/Fable, ChatGPT 5.5, Solar Pro 4, Gemini 3.6 Flash, DeepSeek v4 Flash, MiniMax M3, Kimi k2.5, GLM 4.5, Qwen 3.8 Max, and Inkling Small/Medium.
Thank you for reading. Please leave a like.
repro-time-series-saliency-maps-explaining-models-across-multiple-domains-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-how-much-can-language-models-memorize-traces
Agent traces
Agent sessions published from a Trackio Logbook.
zenith-modelsadapt-pre-trained-VL-models-to-text-data-WikipediaThe Wikipedia train data used to train BERT-base baselines and adapt vision-and-language models to text-only tasks in the paper "How to Adapt Pre-trained Vision-and-Language Models to a Text-only Input?".
The data has been created from the "20200501.en" revision of the wikipedia dataset on Huggingface.
korean-legal-retrieval-source-native-250k
Korean Legal Retrieval Source-Native 250K
Legalize-KR의 법령·행정규칙·판례·자치법규 구조에서 query/positive 관계를 추출한
250,000-row 한국어 retrieval dataset이다. release_eligible: false인
target-adapted 연구·비상업 성능 shard이며 통합 라이선스는 other다.
구성과 고정 revision
Source
Revision
Rows
구조 관계
legalize-kr/legalize-kr
db3cd760c14042ee04fd9166e1bdbb662fc999bc
50,000
법령명+조문 → 조문 본문
legalize-kr/admrule-kr
64a5a272909ab5bc077b0ad9519ef31de8febb46
50,000
규칙명+조문 → 조문 본문
legalize-kr/precedent-kr… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-legal-retrieval-source-native-250k.korean-embedding-performance-v1-sionic-retrieval-train-family-4146
Korean Sionic Retrieval Train-Family 4,146
F2LLM-v2가 공개한 Korean MIRACL, MrTidy, MLDR train-family row만 1M
decontaminated curriculum에서 lossless 추출한 target-adaptation dataset이다. 공개
evaluation query는 포함하지 않으며 current-student HN7 mining 전의 source artifact다.
구성과 목적
source
rows
역할
f2_miracl_ko_train
700
MIRACL Korean retrieval train-family
f2_mrtidy_korean_train
1,200
MrTidy Korean train
f2_mldr_ko_train
2,246
MLDR Korean long-document train-family
합계
4… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-retrieval-train-family-4146.korean-embedding-performance-v1-sionic-autorag-100k
Korean Embedding — Sionic AutoRAG domain 100K
AutoRAG의 금융·상거래·법률 domain retrieval을 보강하기 위한 100,000-row
performance dataset이다. F2LLM-v2 collection의 영어 FIQA/Amazon/Banking77과 중국어
e-commerce/legal QA를 query/positive/negative contrastive schema로 묶었다.
사용 조건과 평가 노출
release_eligible: false인 performance/non-commercial 연구용 composite다. 통합
라이선스는 other이며 F2 collection의 Apache-2.0 표기가 개별 upstream 권리를
재허가하지 않는다.
AutoRAG evaluation repository, query, qrel, corpus는 loader 입력으로… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-autorag-100k.language_models_lab_2korean-embedding-performance-v1-sionic-squad-train-60k
Korean Embedding — Sionic SQuAD train-family 60K
KorQuAD v1.0의 원본 train split만 질문→정답 문맥 retrieval 형식으로 변환한
60,000-row target-adaptation 데이터다. Sionic retrieval 9종 중
SQuADKorV1의 train-family 신호를 명시적으로 보강한다.
사용 조건과 점수 공개 방식
release_eligible: false인 performance/non-commercial 실험용 composite다. 이
저장소의 통합 라이선스는 other이며 upstream 권리를 재허가하지 않는다. Hub metadata는
KorQuAD source를 CC-BY-ND-4.0으로 표시하고, upstream dataset card 본문은
CC BY-ND 2.0 KR도 명시한다. 사용자는 원 source 조건을 직접 확인해야 한다.
이… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-squad-train-60k.
