datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-models-prompt-datasets
🖼️ Open Models Prompt Dataset
🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.open_model_evolution_data
Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
This dataset, released in conjunction with the paper Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem, provides a rigorous examination of concentration dynamics and evolving characteristics in the open model economy.
It compiles a history of weekly model downloads (February 2025-Present) alongside detailed model metadata from the Hugging Face Model Hub. The… See the full description on the dataset page: https://huggingface.co/datasets/mmpr/open_model_evolution_data.open-models-benchmark-results
⚡ Local LLM Evaluation Leaderboard
Welcome to the official public benchmark leaderboard maintained by @ahmedBargady.This dataset repository hosts benchmark evaluation metrics, accuracy scores, throughput telemetry, and quantization trade-off analyses of open-weights foundation models tested locally on NVIDIA A100 GPUs.
💻 Hardware & System Specifications
All evaluations are executed under standardized local cluster environments:
Specification
Details… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/open-models-benchmark-results.Chinese-Herbal-Medicine-Sentiment
中药情感分析数据集 - 数据说明书
Chinese Herbal Medicine Sentiment Analysis Dataset - Datacard
数据集概述 / Dataset Overview
基本信息 / Basic Information
数据集名称 / Dataset Name: Chinese Herbal Medicine Sentiment Analysis Dataset
版本 / Version: 1.0.0
创建日期 / Created: 2025-08-26
作者 / Author: Xingqiang Chen
许可证 / License: MIT
语言 / Language: 中文 (Chinese)
领域 / Domain: 中药 / 传统中医药 (Traditional Chinese Medicine)
数据规模 / Data Scale
总样本数 / Total Samples: 234,879
唯一产品数 /… See the full description on the dataset page: https://huggingface.co/datasets/OpenModels/Chinese-Herbal-Medicine-Sentiment.open_model_evolution_data
Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
This dataset, released in conjunction with the paper Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem, provides a rigorous examination of concentration dynamics and evolving characteristics in the open model economy.
It compiles a history of weekly model downloads (February 2025-Present) alongside detailed model metadata from the Hugging Face Model Hub. The… See the full description on the dataset page: https://huggingface.co/datasets/Iris4ai/open_model_evolution_data.details_OpenModels4all__gemma-1.1-7b-it
Dataset Card for Evaluation run of OpenModels4all/gemma-1.1-7b-it
Dataset automatically created during the evaluation run of model OpenModels4all/gemma-1.1-7b-it on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_OpenModels4all__gemma-1.1-7b-it.open-model-registry
Open Model Registry
Machine-readable mirror of the GitHub registry for original open and open-weight models, with strict sub-12B primary inclusion and separately labeled 16GB edge-deployment records.
data/models.jsonl contains the primary model registry.
data/deployment_variants.jsonl contains official packed and quantized variants of in-scope models.
data/edge_exceptions.jsonl contains models above 12B that have documented 16GB-class deployment evidence; these records never… See the full description on the dataset page: https://huggingface.co/datasets/DDDDD-433/open-model-registry.open_model_evolution_data
Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
This dataset, released in conjunction with the paper Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem, provides a rigorous examination of concentration dynamics and evolving characteristics in the open model economy.
It compiles a history of weekly model downloads (February 2025-Present) alongside detailed model metadata from the Hugging Face Model Hub. The… See the full description on the dataset page: https://huggingface.co/datasets/PrinceAnX/open_model_evolution_data.open_model_evolution_dataRapidata-Finding_the_Subjective_TruthThis data was provided by Rapidata for open sourcing by the Open Model Initiative
You can learn more about Rapidata's global preferencing & human labeling solutions at https://rapidata.ai/
This folder contains the data behind the paper "Finding the subjective Truth - Collecting 2 million votes for comprehensive gen-ai model evaluation"
The paper can be found here: https://arxiv.org/html/2409.11904
Rapidata-Benchmark_v1.0.tsv: Contains the 282 prompts that were used to generate the images with… See the full description on the dataset page: https://huggingface.co/datasets/openmodelinitiative/Rapidata-Finding_the_Subjective_Truth.initial-test-datasetopenmodelmap-modelsrecord-testphotographydigital-arthdr-or-rawsynthetic-imagesopen-models-hf-replication
