UNA
Datasets
All datasets matching “UNA”details_one-man-army__UNA-34Beagles-32K-bf16-v1
Dataset Card for Evaluation run of one-man-army/UNA-34Beagles-32K-bf16-v1
Dataset automatically created during the evaluation run of model one-man-army/UNA-34Beagles-32K-bf16-v1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_one-man-army__UNA-34Beagles-32K-bf16-v1.Videosdetails_one-man-army__una-neural-chat-v3-3-P2-OMA
Dataset Card for Evaluation run of one-man-army/una-neural-chat-v3-3-P2-OMA
Dataset automatically created during the evaluation run of model one-man-army/una-neural-chat-v3-3-P2-OMA on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_one-man-army__una-neural-chat-v3-3-P2-OMA.UnAV-100FaithEval-unanswerable-v1.0
FaithEval
FaithEval is a new and comprehensive benchmark dedicated to evaluating contextual faithfulness in LLMs across three diverse tasks: unanswerable, inconsistent, and counterfactual contexts.
[Paper] FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows", ICLR 2025, https://arxiv.org/abs/2410.03727
[Code and Detailed Instructions] https://github.com/SalesforceAIResearch/FaithEval
Disclaimer and Ethical Considerations… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FaithEval-unanswerable-v1.0.details_fblgit__UNA-SOLAR-10.7B-Instruct-v1.0
Dataset Card for Evaluation run of fblgit/UNA-SOLAR-10.7B-Instruct-v1.0
Dataset automatically created during the evaluation run of model fblgit/UNA-SOLAR-10.7B-Instruct-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_fblgit__UNA-SOLAR-10.7B-Instruct-v1.0.
