datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NexusRaven_API_evaluation
NexusRaven API Evaluation dataset
Please see blog post or NexusRaven Github repo for more information.
License
The evaluation data in this repository consists primarily of our own curated evaluation data that only uses open source commercializable models. However, we include general domain data from the ToolLLM and ToolAlpaca papers. Since the data in the ToolLLM and ToolAlpaca works use OpenAI's GPT models for the generated content, the data is not commercially… See the full description on the dataset page: https://huggingface.co/datasets/Nexusflow/NexusRaven_API_evaluation.details_Nexusflow__Athene-70B
Dataset Card for Evaluation run of Nexusflow/Athene-70B
Dataset automatically created during the evaluation run of model Nexusflow/Athene-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Nexusflow__Athene-70B.Nexusflow_Athene-70B-jdgfct-HarmlessnessNexusflow_Athene-70B-jdgfct-CompletenessFunction_Call_Definitions
Dataset Card for "Nexusflow/Function_Call_Definitions"
More Information needed
VirusTotalBenchmarkllm_logits_Nexusflow_Athene-V2-Chat_493f1bbd561a5a7e3d27c4081d4ee47508bf6831.logits-and-weightsVirusTotalMultiple
Dataset Card for "vt_multiapi_v0"
More Information needed
NVDLibraryBenchmarkClimateAPIBenchmarkCVECPEAPIBenchmarkdetails_Nexusflow__NexusRaven-V2-13B
Dataset Card for Evaluation run of Nexusflow/NexusRaven-V2-13B
Dataset Summary
Dataset automatically created during the evaluation run of model Nexusflow/NexusRaven-V2-13B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Nexusflow__NexusRaven-V2-13B.details_Nexusflow__Starling-LM-7B-beta
Dataset Card for Evaluation run of Nexusflow/Starling-LM-7B-beta
Dataset automatically created during the evaluation run of model Nexusflow/Starling-LM-7B-beta on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Nexusflow__Starling-LM-7B-beta.Nexusflow_Athene-70B-jdgfct-ConcisenessPlacesAPIBenchmarkOTXAPIBenchmarkNexusflow__NexusRaven-V2-13B-details
Dataset Card for Evaluation run of Nexusflow/NexusRaven-V2-13B
Dataset automatically created during the evaluation run of model Nexusflow/NexusRaven-V2-13B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nexusflow__NexusRaven-V2-13B-details.MultiverseMathHardVT_MultiAPIs
Dataset Card for "new_vt_apis"
More Information needed
LangChainMathBenchmarkITType0BenchmarkTicketTrackingBenchmarkITType1BenchmarkVirusTotalAgenticHallucinationTMIBenchmarkLangChainRelationalLangChainMultitoolTypeWriterHardowlet_queries_20250105_agreement_with_gpt_4o_final_trajectoryNexusflow_Athene-70B-jdgfct-Factualitymy_MW_sorted_prompts
