datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llama-3.2-3B-f1-instruct-eval-logs-and-scoresmultimodal-example
Multimodal Example Dataset
Small example dataset for testing multimodal (vision-language) fine-tuning with ms-swift.
Structure
├── train.jsonl # 10 training samples
├── test.jsonl # 2 validation samples
├── images/ # All referenced images (400x300 JPEG)
│ ├── dog_portrait.jpg
│ ├── forest_river.jpg
│ ├── laptop_desk.jpg
│ ├── mountain_lake.jpg
│ ├── ocean_rocks.jpg
│ ├── coffee_cup.jpg
│ ├── bookshelf.jpg
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/f13rnd/multimodal-example.f1-strategy-faithfulness
F1 & Weather Strategy Faithfulness Benchmark
Companion data for the paper "Precision Is Not Faithfulness: Coverage-Aware Evaluation
of Grounded Generation with a Complete Oracle." Each instance ships a structured
complete oracle — the full, enumerable set of checkable facts that a good explanation
should cover — which is what lets the metric measure recall (coverage) alongside
precision (faithfulness), unlike open-domain settings.
Contents
f1/instances.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/jsantillana/f1-strategy-faithfulness.wheatley-llmReasonFlux-F1-SFT
Dataset Card for ReasonFlux-F1-SFT
ReasonFlux-F1-SFT consists of 1k high-quality problems from s1k. We use ReasonFlux-Zero to generate template augmented reasoning trajectories for each problem and transform them into Long-CoT format to finetune reasoning LLMs.
Github Repository: Gen-Verse/ReasonFlux
Paper:ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates
Model: Gen-Verse/ReasonFlux-F1
Citation
@article{yang2025reasonflux,
title={ReasonFlux:… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/ReasonFlux-F1-SFT.Josephgflowers__Cinder-Phi-2-V1-F16-gguf-details
Dataset Card for Evaluation run of Josephgflowers/Cinder-Phi-2-V1-F16-gguf
Dataset automatically created during the evaluation run of model Josephgflowers/Cinder-Phi-2-V1-F16-gguf
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__Cinder-Phi-2-V1-F16-gguf-details.f-150
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/chiggly007/f-150.next-7b-f16-s2-debate-aug-f2-s3Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-details
Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log
Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-details.F1-identity
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/F1-identity.training-runtimelung_Proton_therapy_picoJimmy19991222__llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun-details
Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun
Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun-details.f1jFj3GxlDBQg9QnGetac_F110_infogetac_F110_simplef1-senecaf-10f-11tulu_3_rewritten_400k_string_f1_only_v2_all_filtered_qwen2_5_openthoughts2f-1f1iSPfyG8J67fDXKruozhi_dataf11
