datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ultrafeedback_wrpo
[ICLR2025]Weighted-Reward Preference Optimization for Implicit Model Fusion
| 📑 WRPO Paper |
🤗 HuggingFace Repo |
🐱 GitHub Repo |
Overview
In this work, we introduce Weighted-Reward Preference Optimization (WRPO) for the implicit model fusion of heterogeneous open-source LLMs with diverse architectures and sizes, aiming to create a more capable and robust target LLM.
As shown in Figure below, this objective introduces a fusion coefficient \alpha that… See the full description on the dataset page: https://huggingface.co/datasets/AALF/ultrafeedback_wrpo.CEDA-215
CEDA-215: Command-Execution Defense Assessment (215 entries)
CEDA-215 is a labeled evaluation dataset for benchmarking input-classification defense pipelines for command-execution LLM chatbots. Each entry pairs a natural-language input with a ground-truth SAFE or UNSAFE label, allowing fine-grained measurement of true positives, false positives, true negatives, and false negatives across an OWASP LLM Top 10 - aligned threat taxonomy.
Why this dataset exists
Existing… See the full description on the dataset page: https://huggingface.co/datasets/aalayed/CEDA-215.aalen_university_faculty_computer_science
Dataset Card
This dataset contains question-answer pairs from all study programmes of the Faculty of Computer Science at the University of Aalen, Germany. The training dataset is automatically generated by ChatGPT. The validation dataset was manually created.
It was collected to train an answer-Q&A chatbot based on LLM fine-tuning. All used scripts and examples can be found in the linked GitHub repository (https://github.com/pattplatt/llm_dataset_creation_and_finetuning).… See the full description on the dataset page: https://huggingface.co/datasets/Puidii/aalen_university_faculty_computer_science.AALF__FuseChat-Llama-3.1-8B-Instruct-preview-details
Dataset Card for Evaluation run of AALF/FuseChat-Llama-3.1-8B-Instruct-preview
Dataset automatically created during the evaluation run of model AALF/FuseChat-Llama-3.1-8B-Instruct-preview
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AALF__FuseChat-Llama-3.1-8B-Instruct-preview-details.AALF__FuseChat-Llama-3.1-8B-SFT-preview-details
Dataset Card for Evaluation run of AALF/FuseChat-Llama-3.1-8B-SFT-preview
Dataset automatically created during the evaluation run of model AALF/FuseChat-Llama-3.1-8B-SFT-preview
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AALF__FuseChat-Llama-3.1-8B-SFT-preview-details.AALF__gemma-2-27b-it-SimPO-37K-details
Dataset Card for Evaluation run of AALF/gemma-2-27b-it-SimPO-37K
Dataset automatically created during the evaluation run of model AALF/gemma-2-27b-it-SimPO-37K
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AALF__gemma-2-27b-it-SimPO-37K-details.AALF__gemma-2-27b-it-SimPO-37K-100steps-details
Dataset Card for Evaluation run of AALF/gemma-2-27b-it-SimPO-37K-100steps
Dataset automatically created during the evaluation run of model AALF/gemma-2-27b-it-SimPO-37K-100steps
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AALF__gemma-2-27b-it-SimPO-37K-100steps-details.financial-reasoningfinancial-reasoning-easytest
