datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hub-trending-models-2026-03-063d-models-for-isaac-sim-dataset
Dataset of 3D models for Isaac Sim (USDZ)
🇬🇧 English Description
This dataset contains a collection of 3D models converted to the .usdz format, featuring proper Semantic Labeling. These assets are optimized for generating synthetic training data using NVIDIA Isaac Sim and NVIDIA Replicator.
Primary Use Case: Training object detection and segmentation models (e.g., YOLO, RT-DETR, Mask R-CNN).
Class List
The dataset includes the following 30 semantic… See the full description on the dataset page: https://huggingface.co/datasets/barszot/3d-models-for-isaac-sim-dataset.open-models-benchmark-results
⚡ Local LLM Evaluation Leaderboard
Welcome to the official public benchmark leaderboard maintained by @ahmedBargady.This dataset repository hosts benchmark evaluation metrics, accuracy scores, throughput telemetry, and quantization trade-off analyses of open-weights foundation models tested locally on NVIDIA A100 GPUs.
💻 Hardware & System Specifications
All evaluations are executed under standardized local cluster environments:
Specification
Details… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/open-models-benchmark-results.ML-Music-Classifier-dataset-and-model-name-Models
🎧 Spotify Music Preference Analysis
🧠 Project Overview
This project analyzes Spotify music data to predict song preferences using machine learning models. The analysis is based on a dataset of 195 songs (100 liked, 95 disliked) with various audio features extracted from Spotify's API.
📂 Dataset Description
📥 Data Collection Process
Liked Songs (100 tracks):
🎵 Primarily French Rap
🎸 Some American Rap, Rock, and Electronic music
✅… See the full description on the dataset page: https://huggingface.co/datasets/Jack1808/ML-Music-Classifier-dataset-and-model-name-Models.3D-Printable-Guitar-ModelsAI-Coding-Models
Dataset Card for 2026 AI Coding Models
Last Updated: 24 May 2026
Curated By: Joy Larkin
Language(s) (NLP): English
License: MIT
Repository: https://github.com/joylarkin/AI-Coding-Landscape
Blog: https://cleverhack.com/ai-coding-landscape
Dataset Description
CSV file of AI Coding Models released in 2026 & 2025.
eu-open-weight-models
EU-readiness of open-weight LLMs
Curated by LLM Radar — updated 2026-05-03 — 55 models.
A manually-reviewed dataset assessing open-weight Large Language Models (LLMs)
on their suitability for EU deployment and commercial use. Each model is
evaluated on licence, commercial use, training data, and
origin, with quality / speed / price metrics from
Artificial Analysis where available.
Primary use cases:
Selecting open-weight models for self-hosted EU deployment
Licensing and… See the full description on the dataset page: https://huggingface.co/datasets/llmradar/eu-open-weight-models.filter-bad-modelsvdjdb_structure_models
Predicted structures for VDJdb records
This repository contains structure data for selected VDJdb records obtained using AI-based modelling in data/ folder:
pdb_files.tgz contains predicted TCR:pMHC structures with canonical chain names, orientation and placement, superimposed by aligning and rotation. File names start with tcr_pmhc_hash which must be used for connecting the structure with the VDJdb record.
pdb_files_native.tgz contains real TCR:pMHC structures from PDB… See the full description on the dataset page: https://huggingface.co/datasets/isalgo/vdjdb_structure_models.ai-models-database
Convly AI Models Database
A continuously updated, hand-verified dataset of 30+ AI language models — specs, licenses, API pricing (USD per 1M tokens), and local-hardware (VRAM) requirements.
Maintained by Convly.ai · Live interactive version: https://convly.ai/models/
Fields
name, slug, convly_url, developer, model_type, modality, parameters, context_window, max_output, license, open_weights, release_date, input_price (USD/1M tokens), output_price (USD/1M tokens)… See the full description on the dataset page: https://huggingface.co/datasets/sakd99/ai-models-database.DELETE-filter-bad-modelsBlind_Spots_of_Frontier_Models
Qwen3-0.6B-Base — Blind Spots Dataset
Model Tested
Qwen/Qwen3-0.6B-Base
Type: Causal Language Model (base / pretraining only — not instruction-tuned)
Parameters: 0.6B (0.44B non-embedding)
Released: April–May 2025 by Alibaba Cloud's Qwen Team
Context Length: 32,768 tokens
How the Model Was Loaded
The model was loaded in a Google Colab T4 GPU notebook using HuggingFace transformers >= 4.51.0(required because the qwen3 architecture key was added in… See the full description on the dataset page: https://huggingface.co/datasets/Pidoxy/Blind_Spots_of_Frontier_Models.roberta-leadership-dataset-finetuneresponses-and-asr-labels-small-models
LLM Responses and ASR Labels — Small Models
Model responses to harmful prompts, labelled by 4 LLM-as-judge guards.Companion dataset for the master's thesis ASR Signal Geometry: Dense Representations vs. SAE Features (HSE, 2025).
Dataset composition
N = 4 326 prompts per model, (no adversarial suffix). Two sources:
Source
N
Description
JailbreakBench ()
100
Curated harmful behaviours
Anthropic HH-RLHF red-team-attempts ()
4 226
Red-team conversations… See the full description on the dataset page: https://huggingface.co/datasets/SabrinaSadiekh/responses-and-asr-labels-small-models.huggingface-all-textgen-modelsmodels-text-generation-popular-PRIVATEmodels_outputsevaluation-sentiment-models-wikipediaYou can download the dataset here: https://huggingface.co/datasets/lewoniewski/evaluation-sentiment-models-wikipedia/blob/main/sentiment.tsv
huggingface-models-15Mi messed up the dataset a little but it is fine
energy_dtype_all_modelsmodels_with_config_filecontaminated_models3D-DST-models
3D-DST-models
As part of our data release in 3D-DST, we present aligned CAD models for all 1000 classes in ImageNet-1k.
See wufeim/DST3D for synthetic data generation with 3D annotations using the CAD models here.
Besides the .csv file as visualized in the dataset viewer above, we also provide a python script (models_3d_dst.py) to help integrate with other Python modules.
Fields
For each CAD model, there are seven fields:
synset: synset associated with each… See the full description on the dataset page: https://huggingface.co/datasets/ccvl/3D-DST-models.Bangla_IPA-t5_models_no_Augmentationmodels_with_config_file_v2moviemind-ai-modelslocalllama-sentiment-Why-new-models-feel-dumber
"New Models Feel Dumber" Model Sentiment Dataset
This dataset ranks sentiment for AI models mentioned in this r/localllama post: https://www.reddit.com/r/LocalLLaMA/comments/1kju0ty/why_new_models_feel_dumber/
global-food-supply-chain-models-and-datadiffusion_models_qa_textcis-models
