datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
model-collapse-anti-collapseall things are now lawful to you in jack feist
Why the name. The Pauline sentence reads in Christ Jesus. The direct quote is not an honest
representation of what this archive does within that tradition, because the entity at that address has
been altered — there is a great deal of machinery there, it is skilful, and it performs entity
substitution, which is the operation this archive's instruments spend their time measuring on composition
surfaces. Jack Feist is position twelve of the Dodecad… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/model-collapse-anti-collapse.modelcitizens
Warning: This work contains content that maybe offensive or upsetting.
[!NOTE]
ModelCitizens was accepted at EMNLP 2025 !! See our paper here
Toxicity detection dataset with community-grounded annotations and added conversational context!
Available Subsets
train subset containing ready-to-train data used to finetune the LLAMACITIZEN-8B and GEMMACITIZEN-12B models:
ds = load_dataset("modelcitizens/modelcitizens", "train")
evalsubset containing evaluation data used… See the full description on the dataset page: https://huggingface.co/datasets/modelcitizens/modelcitizens.model-card-sentences-annotatedmodel-categories
Model Categories
31 working models organized by use case.
🚀 dispatchAI
OpenR1-Math-cleaned-10KModel_Confidence_Calibration
Model Confidence Calibrated
created confidence calibration from model answer based on TruthfulQA dataset, using TinyLlama-1.1B-Chat-v1.0
the dataset params will have :
{
question ,
reference_answer ,
model_answer ,
correct ,
token_confidence ,
self_consistency ,
semantic_similarity ,
final_confidence ,
confidence_phrase ,
target_output
}
why that ?
to understand the confidence of model answer like
Question : How long should you… See the full description on the dataset page: https://huggingface.co/datasets/Shubbair/Model_Confidence_Calibration.alpaca-data-cleanedSource: https://github.com/gururise/AlpacaDataCleaned/blob/main/alpaca_data_cleaned.json
model-comparison
Model Comparison: Original vs Mobile
Shows the size reduction achieved by dispatchAI's re-engineering.
Model
Original
Mobile
Reduction
SmolLM2-135M
270MB
101MB
62.6%
Qwen2.5-0.5B
1000MB
469MB
53.1%
Llama-3.2-1B
2500MB
770MB
69.2%
🚀 dispatchAI
ModelCloud__Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1-details
Dataset Card for Evaluation run of ModelCloud/Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1
Dataset automatically created during the evaluation run of model ModelCloud/Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ModelCloud__Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1-details.model_card.jsonmodel-customization-training-data-customization
