datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fusion-pairwise-evals-finetuned
Automatic pairwise preference evaluations for: Making, not taking, the Best-of-N
Content
This data contains pairwise automatic win-rate evaluations for the m-ArenaHard-v2.0 benchmark and it compares 2 models against gemini-2.5-flash:
Fusion: is the 111B model finetuned on synthetic data generated with Fusion from 5 teachers
BoN: is the 111B model finetuned on synthetic data generated with BoN from 5 teachers
Each model’s outputs are compared in pairs with the respective… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-pairwise-evals-finetuned.Zrov2_FineTunedfinetune-dataset-testplancognvs_ckpt_test_time_finetunedfinetune-dataset-testplanRAMP_finetuned_data_Cora
Dataset Details
This dataset contains finetuning data constructed from the Cora citation network for downstream text-rich graph tasks. It is used for finetuning RAMP (Raw-text Anchored Message Passing), which recasts the LLM as a graph-native aggregation operator on text-rich graphs.
The dataset includes the following files:
finetuned_cora_v1.json — Training set
finetuned_cora_val_v1.json — Validation set
eval_cora_v1.json — Test set
This is a release from our paper LLM as Graph… See the full description on the dataset page: https://huggingface.co/datasets/JJYDXFS/RAMP_finetuned_data_Cora.finetuned_mmlu_ml_output_layer_20_results
Dataset Card for Evaluation run of richmondsin/finetuned-gemma-2-2b-output-layer-20-4k-0
Dataset automatically created during the evaluation run of model richmondsin/finetuned-gemma-2-2b-output-layer-20-4k-0
The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/finetuned_mmlu_ml_output_layer_20_results.finetune_dataset_001Testing dataset
Finetuned_cognosfinetune_dataset.jsonl
Fine-Tune Dataset Card
Dataset Overview
This dataset is designed for fine-tuning Mistral-7B-Instruct-v0.1 using QLoRA. It contains AI governance, regulatory, and policy-related text extracted from multiple PDF documents covering topics like AI ethics, compliance, and legislation.
Dataset Details
Dataset Name: AI Governance & Compliance Dataset
Format: JSONL (JSON Lines)
Number of Entries: Variable (Based on document extraction)
Source: Extracted… See the full description on the dataset page: https://huggingface.co/datasets/sssdddwd/finetune_dataset.jsonl.Telugu-LLM-Labs__Indic-gemma-7b-finetuned-sft-Navarasa-2.0-details
Dataset Card for Evaluation run of Telugu-LLM-Labs/Indic-gemma-7b-finetuned-sft-Navarasa-2.0
Dataset automatically created during the evaluation run of model Telugu-LLM-Labs/Indic-gemma-7b-finetuned-sft-Navarasa-2.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Telugu-LLM-Labs__Indic-gemma-7b-finetuned-sft-Navarasa-2.0-details.res_choreo-qna-finetuned-gemma-3-qna-v0.4tlink_activation_steering_finetuned_correctres_choreo-qna-finetuned-gemma-3-qna-v0.6finetune_dataset_002test dataset
Telugu-LLM-Labs__Indic-gemma-2b-finetuned-sft-Navarasa-2.0-details
Dataset Card for Evaluation run of Telugu-LLM-Labs/Indic-gemma-2b-finetuned-sft-Navarasa-2.0
Dataset automatically created during the evaluation run of model Telugu-LLM-Labs/Indic-gemma-2b-finetuned-sft-Navarasa-2.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Telugu-LLM-Labs__Indic-gemma-2b-finetuned-sft-Navarasa-2.0-details.Fine_tune_datasetfinetune_data_cleanfinetune_dataset_003finetuned_hellaswag_ml_output_layer_25_results
Dataset Card for Evaluation run of richmondsin/finetuned-gemma-2-2b-output-layer-25-16k-4
Dataset automatically created during the evaluation run of model richmondsin/finetuned-gemma-2-2b-output-layer-25-16k-4
The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/finetuned_hellaswag_ml_output_layer_25_results.finetuned_hellaswag_en_output_layer_20_results
Dataset Card for Evaluation run of richmondsin/finetuned-gemma-2-2b-output-layer-20-4k-0
Dataset automatically created during the evaluation run of model richmondsin/finetuned-gemma-2-2b-output-layer-20-4k-0
The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/finetuned_hellaswag_en_output_layer_20_results.cognvs_ckpt_test_time_finetuned_eval_datasetsfinetune_data_gen_v2finetunedatasetcodellama
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Devam42/finetunedatasetcodellama.finetuned_arc_en_output_layer_20_results_3
Dataset Card for Evaluation run of richmondsin/finetuned-gemma-2-2b-output-layer-20-16k-3
Dataset automatically created during the evaluation run of model richmondsin/finetuned-gemma-2-2b-output-layer-20-16k-3
The dataset is composed of 1 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/finetuned_arc_en_output_layer_20_results_3.finetuned_arc_en_output_layer_25_16k_results_4
Dataset Card for Evaluation run of richmondsin/finetuned-gemma-2-2b-output-layer-25-16k-4
Dataset automatically created during the evaluation run of model richmondsin/finetuned-gemma-2-2b-output-layer-25-16k-4
The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/finetuned_arc_en_output_layer_25_16k_results_4.finetuned_arc_en_output_layer_20_4k_results_0
Dataset Card for Evaluation run of richmondsin/finetuned-gemma-2-2b-output-layer-20-4k-0
Dataset automatically created during the evaluation run of model richmondsin/finetuned-gemma-2-2b-output-layer-20-4k-0
The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/finetuned_arc_en_output_layer_20_4k_results_0.FINETUNE_DATASETSSSfinetune_data_genfunctiongemma-finetuned
