datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
model_cards_with_metadatamodel-collapse-anti-collapseall things are now lawful to you in jack feist
Why the name. The Pauline sentence reads in Christ Jesus. The direct quote is not an honest
representation of what this archive does within that tradition, because the entity at that address has
been altered — there is a great deal of machinery there, it is skilful, and it performs entity
substitution, which is the operation this archive's instruments spend their time measuring on composition
surfaces. Jack Feist is position twelve of the Dodecad… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/model-collapse-anti-collapse.model-catalogmodelcitizens
Warning: This work contains content that maybe offensive or upsetting.
[!NOTE]
ModelCitizens was accepted at EMNLP 2025 !! See our paper here
Toxicity detection dataset with community-grounded annotations and added conversational context!
Available Subsets
train subset containing ready-to-train data used to finetune the LLAMACITIZEN-8B and GEMMACITIZEN-12B models:
ds = load_dataset("modelcitizens/modelcitizens", "train")
evalsubset containing evaluation data used… See the full description on the dataset page: https://huggingface.co/datasets/modelcitizens/modelcitizens.model_cards_with_readmes
Dataset Card for "model_cards_with_readmes"
More Information needed
model-context-windows
LLM Context Windows — 206 models
Context-window sizes for 206 ready models served by the Qubax AI API (OpenAI-compatible), exported from the public /v1/models endpoint.
Columns
Column
Description
model_id
API model identifier
model_name
Display name
context_window_tokens
Max context window (tokens)
max_output_tokens
Max output (tokens, where published)
source
Provenance
Notes
License: CC0 1.0 (public domain) — use freely in… See the full description on the dataset page: https://huggingface.co/datasets/QubaxAI/model-context-windows.model-card-sentences-annotatedmodel-categories
Model Categories
31 working models organized by use case.
🚀 dispatchAI
model_cards_correct_tagmodel_cards_with_long_context_embeddings
Dataset Card for "model_cards_with_long_context_embeddings"
More Information needed
aim2balance_modelcocktail_optimizationc_corpus_br_finetuning_language_model_deberta
Dataset Card for "c_corpus_br_finetuning_language_model_deberta"
More Information needed
model_cards_with_readmes_sections
Dataset Card for "model_cards_with_readmes_sections"
More Information needed
model_cards_with_metadata_with_embeddings
Dataset Card for Hugging Face Hub Model Cards with Embeddings
This dataset consists of model cards for models hosted on the Hugging Face Hub. The model cards are created by the community and provide information about the model, its performance, its intended uses, and more.
This dataset is updated on a daily basis and includes publicly available models on the Hugging Face Hub.
This dataset is made available to help support users wanting to work with a large number of Model Cards… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/model_cards_with_metadata_with_embeddings.OpenR1-Math-cleaned-10Kmodel-compatibilityModel_Confidence_Calibration
Model Confidence Calibrated
created confidence calibration from model answer based on TruthfulQA dataset, using TinyLlama-1.1B-Chat-v1.0
the dataset params will have :
{
question ,
reference_answer ,
model_answer ,
correct ,
token_confidence ,
self_consistency ,
semantic_similarity ,
final_confidence ,
confidence_phrase ,
target_output
}
why that ?
to understand the confidence of model answer like
Question : How long should you… See the full description on the dataset page: https://huggingface.co/datasets/Shubbair/Model_Confidence_Calibration.c_corpus_br_finetuning_language_model_bert
Dataset Card for "c_corpus_br_finetuning_language_model_bert"
More Information needed
model-code-exceptionmodel-cards-ml-metadata-bootstrap
davanstrien/model-cards-ml-metadata-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over librarian-bots/model_cards_with_metadata.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
librarian-bots/model_cards_with_metadata (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
base model name, context length, training method, training dataset name, benchmark name… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/model-cards-ml-metadata-bootstrap.alpaca-data-cleanedSource: https://github.com/gururise/AlpacaDataCleaned/blob/main/alpaca_data_cleaned.json
model_cards_with_metadatamodel-cards-entities-1k-gpu
davanstrien/model-cards-entities-1k-gpu
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over librarian-bots/model_cards_with_metadata.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
librarian-bots/model_cards_with_metadata (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
Person, Organization, Dataset, Model, Framework
Confidence threshold
0.6
Samples processed… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/model-cards-entities-1k-gpu.model_cards_with_readmes_with_duplicates
Dataset Card for "model_cards_with_readmes_with_duplicates"
More Information needed
model-comparison
Model Comparison: Original vs Mobile
Shows the size reduction achieved by dispatchAI's re-engineering.
Model
Original
Mobile
Reduction
SmolLM2-135M
270MB
101MB
62.6%
Qwen2.5-0.5B
1000MB
469MB
53.1%
Llama-3.2-1B
2500MB
770MB
69.2%
🚀 dispatchAI
model-card-sentencesmodel-cards-entities-1k-cpu
davanstrien/model-cards-entities-1k-cpu
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over librarian-bots/model_cards_with_metadata.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
librarian-bots/model_cards_with_metadata (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
Person, Organization, Dataset, Model, Framework
Confidence threshold
0.6
Samples processed… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/model-cards-entities-1k-cpu.ModelCloud__Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1model_card_dataset_mentions
Dataset Card for Model Card Dataset Mentions
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/model_card_dataset_mentions.model_comparison_gpt4
