datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
model_cards_with_metadatamodelcitizens
Warning: This work contains content that maybe offensive or upsetting.
[!NOTE]
ModelCitizens was accepted at EMNLP 2025 !! See our paper here
Toxicity detection dataset with community-grounded annotations and added conversational context!
Available Subsets
train subset containing ready-to-train data used to finetune the LLAMACITIZEN-8B and GEMMACITIZEN-12B models:
ds = load_dataset("modelcitizens/modelcitizens", "train")
evalsubset containing evaluation data used… See the full description on the dataset page: https://huggingface.co/datasets/modelcitizens/modelcitizens.model_cards_with_readmes
Dataset Card for "model_cards_with_readmes"
More Information needed
model-context-windows
LLM Context Windows — 206 models
Context-window sizes for 206 ready models served by the Qubax AI API (OpenAI-compatible), exported from the public /v1/models endpoint.
Columns
Column
Description
model_id
API model identifier
model_name
Display name
context_window_tokens
Max context window (tokens)
max_output_tokens
Max output (tokens, where published)
source
Provenance
Notes
License: CC0 1.0 (public domain) — use freely in… See the full description on the dataset page: https://huggingface.co/datasets/QubaxAI/model-context-windows.model-categories
Model Categories
31 working models organized by use case.
🚀 dispatchAI
model_cards_correct_tagmodel_cards_with_long_context_embeddings
Dataset Card for "model_cards_with_long_context_embeddings"
More Information needed
model_cards_with_metadata_with_embeddings
Dataset Card for Hugging Face Hub Model Cards with Embeddings
This dataset consists of model cards for models hosted on the Hugging Face Hub. The model cards are created by the community and provide information about the model, its performance, its intended uses, and more.
This dataset is updated on a daily basis and includes publicly available models on the Hugging Face Hub.
This dataset is made available to help support users wanting to work with a large number of Model Cards… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/model_cards_with_metadata_with_embeddings.Model_Confidence_Calibration
Model Confidence Calibrated
created confidence calibration from model answer based on TruthfulQA dataset, using TinyLlama-1.1B-Chat-v1.0
the dataset params will have :
{
question ,
reference_answer ,
model_answer ,
correct ,
token_confidence ,
self_consistency ,
semantic_similarity ,
final_confidence ,
confidence_phrase ,
target_output
}
why that ?
to understand the confidence of model answer like
Question : How long should you… See the full description on the dataset page: https://huggingface.co/datasets/Shubbair/Model_Confidence_Calibration.model-cards-ml-metadata-bootstrap
davanstrien/model-cards-ml-metadata-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over librarian-bots/model_cards_with_metadata.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
librarian-bots/model_cards_with_metadata (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
base model name, context length, training method, training dataset name, benchmark name… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/model-cards-ml-metadata-bootstrap.model_cards_with_metadatamodel-cards-entities-1k-gpu
davanstrien/model-cards-entities-1k-gpu
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over librarian-bots/model_cards_with_metadata.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
librarian-bots/model_cards_with_metadata (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
Person, Organization, Dataset, Model, Framework
Confidence threshold
0.6
Samples processed… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/model-cards-entities-1k-gpu.model_cards_with_readmes_with_duplicates
Dataset Card for "model_cards_with_readmes_with_duplicates"
More Information needed
model-comparison
Model Comparison: Original vs Mobile
Shows the size reduction achieved by dispatchAI's re-engineering.
Model
Original
Mobile
Reduction
SmolLM2-135M
270MB
101MB
62.6%
Qwen2.5-0.5B
1000MB
469MB
53.1%
Llama-3.2-1B
2500MB
770MB
69.2%
🚀 dispatchAI
model-cards-entities-1k-cpu
davanstrien/model-cards-entities-1k-cpu
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over librarian-bots/model_cards_with_metadata.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
librarian-bots/model_cards_with_metadata (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
Person, Organization, Dataset, Model, Framework
Confidence threshold
0.6
Samples processed… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/model-cards-entities-1k-cpu.ModelCloud__Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1ModelCloud__Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1-details
Dataset Card for Evaluation run of ModelCloud/Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1
Dataset automatically created during the evaluation run of model ModelCloud/Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ModelCloud__Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1-details.model_card_extractor_qwq
Dataset card for model_card_extractor_qwq
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"modelId": "digiplay/XtReMixAnimeMaster_v1",
"author": "digiplay",
"last_modified": "2024-03-16 00:22:41+00:00",
"downloads": 232,
"likes": 2,
"library_name": "diffusers",
"tags": [
"diffusers",
"safetensors",
"stable-diffusion",
"stable-diffusion-diffusers",
"text-to-image"… See the full description on the dataset page: https://huggingface.co/datasets/aravind-selvam/model_card_extractor_qwq.model_cards_with_readmes
Dataset Card for "model_cards_with_readmes"
More Information needed
model-capability-classificationmodel_capability
