CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01burtenshaw /trending-models-top10-2026-03-06 Top 10 Trending Models (2026-03-06) This dataset records the top 10 trending models on the Hugging Face Hub captured on 2026-03-06. Files hf_trending_models_top10_2026-03-06.csv hf_trending_models_top10_2026-03-06.json Collection Method Collected with: hf models ls --sort trending_score --limit 10 Scores are point-in-time values and can change quickly. textn<1K0 likes1.9k downloads7mo agoHugging Face02LLM-OS-Models /KoHRM-Text-1.4B-prepared-data KoHRM-Text-1.4B Prepared Data This dataset repository contains prepared HRM-Text V1Dataset artifacts for KoHRM-Text-1.4B. The data is intended for continued pretraining and staged training with the project code at: https://github.com/LLM-OS-Models/KoHRM-text https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B https://huggingface.co/LLM-OS-Models/HRM-Text-Ko-Terminal-Tokenizer-131K The upstream architecture and training method are based on: Paper:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-prepared-data.tabulartext-generationn<1K1 likes1k downloads4mo agoHugging Face03ModelsLab /midashenglm-gen-training-latents ModelsLab/midashenglm-gen-training-latents Precomputed audio latents for fine-tuning mispeech/midashenglm-gen, paired with six-view prompts in the exact format the model was trained on. This is not an audio dataset and not a caption dataset. Each record is the output of the model's frozen DashengTokenizer encoder — 768-dimensional latents at 25 Hz, stored float16 — next to the tagged prompt string built from the source metadata. Why it exists The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.tabulartext-to-audion<1K0 likes655 downloads1mo agoHugging Face04Arrrlex /models-under-pressure Models Under Pressure This dataset accompanies the paper Detecting High-Stakes Interactions with Activation Probes, presented at the ICML 2025 Workshop on Actionable Interpretability, accepted to NeurIPS 2025. Overview Every sample is a user-facing LLM interaction labelled as high-stakes or low-stakes. The label reflects whether the conversation involves potentially consequential outcomes (medical advice, legal matters, financial decisions, etc.) vs. routine queries. The… See the full description on the dataset page: https://huggingface.co/datasets/Arrrlex/models-under-pressure.tabulartext-classification10K<n<100K0 likes453 downloads8mo agoHugging Face05sasha /co2_modelstabular1K<n<10K0 likes225 downloads3y agoHugging Face06abersbail /blender-3d-models Blender 3D Models Database This repository serves as a database of custom-made 3D models generated programmatically using Blender and uploaded to Hugging Face. Models Included ⚔️ Stylized Fantasy Sword (stylized_sword.glb) - View Model Page 🛡️ Stylized Fantasy Shield (stylized_shield.glb) - View Model Page 🪄 Stylized Wizard's Staff (stylized_staff.glb) - View Model Page 📦 Stylized Treasure Chest (stylized_chest.glb) - View Model Page 🏹 Stylized Bow… See the full description on the dataset page: https://huggingface.co/datasets/abersbail/blender-3d-models.3dn<1K0 likes220 downloads4mo agoHugging Face07modelscope /MMMU-Reasoning-Distill-Validation中文版本 Description MMMU-Reasoning-Distill-Validation is a Multi-Modal reasoning dataset that contains 839 image descriptions and natural language inference data samples. This dataset is built upon the validation set of MMMU. The construction process begins with using Qwen2.5-VL-72B-Instruct for image understanding and generating detailed image descriptions, followed by generating reasoning conversations using the DeepSeek-R1 model. Its main features are as follows: Use the… See the full description on the dataset page: https://huggingface.co/datasets/modelscope/MMMU-Reasoning-Distill-Validation.imagen<1K2 likes212 downloads2y agoHugging Face08gemmozero /ai-models-2026 AI Models & Releases 2026 AI model releases, benchmarks, capabilities. Updated daily via automated collection pipeline. Part of the Legion Data Factory — historical AI ecosystem datasets 2026. Methodology Automated collection from public sources (HackerNews, RSS feeds, APIs). Updated daily via cron job. Raw data, minimal processing. License CC BY 4.0 🔑 API Access — Updated Daily Live data via Legion AI API | Documentation Free: 100… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-models-2026.text1K<n<10K0 likes186 downloads14h agoHugging Face09csoai /gspc-own-models-measured Our own models, measured and excluded from leads We measure everyone on the board — including our own fine-tunes. This dataset is the measurement record for the in-house clan-* and council-* models the Council of AI (CSOAI) trained and then ran through the same GSPC board as every other system. No exception, no hand-scoring, no hiding the weak rows. 18 in-house fine-tunes, 50 signed measurement cards. Every card is Ed25519-signed by CSOAI Ltd (UK 16939677) and re-verifiable.… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-own-models-measured.textothern<1K0 likes175 downloads8d agoHugging Face10gg676 /ParserV1-modelstextn<1K0 likes165 downloads10mo agoHugging Face11zxliu /ReAPR-Automatic-Program-Repair-via-Retrieval-Augmented-Large-Language-ModelsThis is the Retrieval dataset used in the paper "ReAPR: Automatic Program Repair via Retrieval-Augmented Large Language Models" text100K<n<1M3 likes162 downloads2y agoHugging Face12matonski /toy-models-of-sft-data Toy Models of SFT Data This is a public-clean candidate data package for the Toy Models of SFT project. It is built for researcher inspection first. The package answers two questions: What were the models trained on? How did the models actually behave under evaluation? The package includes training data, eval inputs, model rollouts, judge scores, parsed GPQA outputs, aggregate tables, paper figures, frozen plot data, and provenance records. It deliberately includes some… See the full description on the dataset page: https://huggingface.co/datasets/matonski/toy-models-of-sft-data.tabulartext-generation10K<n<100K0 likes148 downloads2mo agoHugging Face13modelscope /self-cognition 介绍(Introduction) 该自我认知数据集由modelsope swift创建, 可以通过将通配符进行替换:{{NAME}}、{{AUTHOER}},来创建属于自己大模型的自我认知数据集,总共108条。 ms-swift github:https://github.com/modelscope/swift/ 自我认知微调最佳实践文档:https://github.com/modelscope/swift/blob/main/docs/source/LLM/%E8%87%AA%E6%88%91%E8%AE%A4%E7%9F%A5%E5%BE%AE%E8%B0%83%E6%9C%80%E4%BD%B3%E5%AE%9E%E8%B7%B5.md This self-cognition dataset was created by modelsope swift and can be customized for your own large model by replacing the placeholders: {{NAME}} and… See the full description on the dataset page: https://huggingface.co/datasets/modelscope/self-cognition.textn<1K27 likes129 downloads2y agoHugging Face14adsgpt /marketing-benchmark-of-more-than-10-ai-models Marketing Benchmark of 10+ AI Models A 5,000-question benchmark for evaluating LLMs across six dimensions of modern marketing — Meta Ads, Google Ads, SEO & Organic, Email & Lifecycle, Critical Thinking, and Action-Based scenarios — graded through 10 distinct marketer personas. Every question is independently authored by the AdsGPT Marketing Bench team. Knowledge MCQs are hand-authored against 2026 platform documentation; open-ended and action-based scenarios are built from… See the full description on the dataset page: https://huggingface.co/datasets/adsgpt/marketing-benchmark-of-more-than-10-ai-models.textquestion-answering1K<n<10K5 likes122 downloads4mo agoHugging Face15LLM-OS-Models /korean-embedding-performance-v1-performance-1m Korean Embedding Performance v1 — 1M Qwen3-Embedding-8B의 한국어 retrieval data-scale 실험을 위한 정확히 1,000,000-row 연구·비상업 contrastive dataset이다. release_eligible: false, 통합 라이선스 other이며 upstream source 조건을 재허가하지 않는다. 구성 계열 Rows 비율 역할 nlpai-lab/ko-triplet-v1.0 600,254 60.03% 넓은 한국어 QA/retrieval core F2 Korean QA/instruction 287,000 28.70% webfaq, mqa, koalpaca, realQA, komagpie F2 retrieval task train-family 4,146 0.41% MIRACL, MrTidy, MLDR F2 PAWS-X… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-performance-1m.textsentence-similarity1M<n<10M0 likes104 downloads2mo agoHugging Face16Daniel4190 /filtered_models_swe_smithtext1K<n<10K0 likes98 downloads1y agoHugging Face17LLM-OS-Models /korean-embedding-performance-v1-ablation-200k Korean Embedding Performance v1 — Ablation 200K Qwen3-Embedding-8B의 한국어 retrieval continued fine-tuning에서 LoRA/DoRA/부분 및 full fine-tuning, loss, hard-negative 전략을 비교하기 위한 200,000-row 연구·비상업 성능 데이터다. release_eligible: false이며 통합 라이선스는 other다. upstream source별 조건을 재허가하지 않는다. 구성 계열 Rows 역할 nlpai-lab/ko-triplet-v1.0@1f5d72d 100,254 넓은 한국어 QA/retrieval core F2 Korean QA/instruction 68,000 webfaq, mqa, koalpaca, realQA, komagpie F2 retrieval task… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-ablation-200k.textsentence-similarity100K<n<1M0 likes98 downloads2mo agoHugging Face18LLM-OS-Models /korean-embedding-performance-v1-pilot-50k Korean Embedding Performance v1 — Pilot 50K 주의: 이 revision은 공개 benchmark 성능 후보 학습에 사용하면 안 된다. 사후 15-task exact text-hash 감사에서 평가 query 고유 hash 4개가 확인됐다. 파이프라인·최적화 진단과 contamination ablation에만 남기며, 교체본은 ablation-200k이다. Qwen3-Embedding 계열의 한국어 retrieval 성능 실험을 위한 50,000-row 연구용 contrastive dataset이다. 각 row는 instruction-aware query, positive passage 1개, hard/easy negative passage 1–7개를 ms-swift embedding message schema로 저장한다. 사용 조건과 공개 범위 이 저장소의 통합 라이선스는 other다.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-pilot-50k.texttext-retrieval10K<n<100K0 likes95 downloads2mo agoHugging Face19Sony /DeepResonance_data_models DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning Paper: arxiv This is a repository of data and models for DeepResonance. Data For all the existing datasets, download the multimodal resources according to the original papers. For Music4way related datasets, first download all the video and music files using the YouTube IDs shown in each dataset. Then referring to M2UGen's pipeline to randomly extract an image from… See the full description on the dataset page: https://huggingface.co/datasets/Sony/DeepResonance_data_models.text10K<n<100K1 likes84 downloads1y agoHugging Face20Ryukijano /repro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes83 downloads2mo agoHugging Face21mondk /12.Models.Ask.Themselves.Q_AThe data files were generated through conversations with various models, notably: Claude Sonnet/Opus/Fable, ChatGPT 5.5, Solar Pro 4, Gemini 3.6 Flash, DeepSeek v4 Flash, MiniMax M3, Kimi k2.5, GLM 4.5, Qwen 3.8 Max, and Inkling Small/Medium. Thank you for reading. Please leave a like. texttext-generationn<1K5 likes83 downloads2mo agoHugging Face22JUNGU /repro-time-series-saliency-maps-explaining-models-across-multiple-domains-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes77 downloads2mo agoHugging Face23abidlabs /repro-how-much-can-language-models-memorize-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes74 downloads2mo agoHugging Face24alaaaldeen1994 /zenith-modelstabularn<1K0 likes66 downloads3mo agoHugging Face25Lo /adapt-pre-trained-VL-models-to-text-data-WikipediaThe Wikipedia train data used to train BERT-base baselines and adapt vision-and-language models to text-only tasks in the paper "How to Adapt Pre-trained Vision-and-Language Models to a Text-only Input?". The data has been created from the "20200501.en" revision of the wikipedia dataset on Huggingface. text1M<n<10M0 likes63 downloads4y agoHugging Face26LLM-OS-Models /korean-legal-retrieval-source-native-250k Korean Legal Retrieval Source-Native 250K Legalize-KR의 법령·행정규칙·판례·자치법규 구조에서 query/positive 관계를 추출한 250,000-row 한국어 retrieval dataset이다. release_eligible: false인 target-adapted 연구·비상업 성능 shard이며 통합 라이선스는 other다. 구성과 고정 revision Source Revision Rows 구조 관계 legalize-kr/legalize-kr db3cd760c14042ee04fd9166e1bdbb662fc999bc 50,000 법령명+조문 → 조문 본문 legalize-kr/admrule-kr 64a5a272909ab5bc077b0ad9519ef31de8febb46 50,000 규칙명+조문 → 조문 본문 legalize-kr/precedent-kr… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-legal-retrieval-source-native-250k.textsentence-similarity100K<n<1M0 likes63 downloads2mo agoHugging Face27LLM-OS-Models /korean-embedding-performance-v1-sionic-retrieval-train-family-4146 Korean Sionic Retrieval Train-Family 4,146 F2LLM-v2가 공개한 Korean MIRACL, MrTidy, MLDR train-family row만 1M decontaminated curriculum에서 lossless 추출한 target-adaptation dataset이다. 공개 evaluation query는 포함하지 않으며 current-student HN7 mining 전의 source artifact다. 구성과 목적 source rows 역할 f2_miracl_ko_train 700 MIRACL Korean retrieval train-family f2_mrtidy_korean_train 1,200 MrTidy Korean train f2_mldr_ko_train 2,246 MLDR Korean long-document train-family 합계 4… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-retrieval-train-family-4146.texttext-retrieval1K<n<10K0 likes56 downloads2mo agoHugging Face28LLM-OS-Models /korean-embedding-performance-v1-sionic-autorag-100k Korean Embedding — Sionic AutoRAG domain 100K AutoRAG의 금융·상거래·법률 domain retrieval을 보강하기 위한 100,000-row performance dataset이다. F2LLM-v2 collection의 영어 FIQA/Amazon/Banking77과 중국어 e-commerce/legal QA를 query/positive/negative contrastive schema로 묶었다. 사용 조건과 평가 노출 release_eligible: false인 performance/non-commercial 연구용 composite다. 통합 라이선스는 other이며 F2 collection의 Apache-2.0 표기가 개별 upstream 권리를 재허가하지 않는다. AutoRAG evaluation repository, query, qrel, corpus는 loader 입력으로… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-autorag-100k.textsentence-similarity100K<n<1M0 likes50 downloads2mo agoHugging Face29robotamski /language_models_lab_2text1M<n<10M0 likes48 downloads4d agoHugging Face30LLM-OS-Models /korean-embedding-performance-v1-sionic-squad-train-60k Korean Embedding — Sionic SQuAD train-family 60K KorQuAD v1.0의 원본 train split만 질문→정답 문맥 retrieval 형식으로 변환한 60,000-row target-adaptation 데이터다. Sionic retrieval 9종 중 SQuADKorV1의 train-family 신호를 명시적으로 보강한다. 사용 조건과 점수 공개 방식 release_eligible: false인 performance/non-commercial 실험용 composite다. 이 저장소의 통합 라이선스는 other이며 upstream 권리를 재허가하지 않는다. Hub metadata는 KorQuAD source를 CC-BY-ND-4.0으로 표시하고, upstream dataset card 본문은 CC BY-ND 2.0 KR도 명시한다. 사용자는 원 source 조건을 직접 확인해야 한다. 이… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-squad-train-60k.textsentence-similarity10K<n<100K0 likes46 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.