datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bench-automationbench
AutomationBench public task payloads
Materialized public task inputs for Zapier AutomationBench,
pinned to upstream revision 4a8e1061254004d9dac807054eed33fad7d1ff14.
This repository contains six Parquet files (100 tasks per public domain) generated from the
upstream get_<domain>_dataset() functions. It exists so the solar-system evaluation integration
can fetch immutable task payloads without committing multi-megabyte generated Python modules.
Source and license: Zapier… See the full description on the dataset page: https://huggingface.co/datasets/hyeonseop-upstage/bench-automationbench.details_upstage__SOLAR-10.7B-Instruct-v1.0
Dataset Card for Evaluation run of upstage/SOLAR-10.7B-Instruct-v1.0
Dataset automatically created during the evaluation run of model upstage/SOLAR-10.7B-Instruct-v1.0.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_upstage__SOLAR-10.7B-Instruct-v1.0.automationBench
AutomationBench public task payloads
Materialized public task inputs for Zapier AutomationBench,
pinned to upstream revision 4a8e1061254004d9dac807054eed33fad7d1ff14.
This repository contains six Parquet files (100 tasks per public domain) generated from the
upstream get_<domain>_dataset() functions. It exists so the solar-system evaluation integration
can fetch immutable task payloads without committing multi-megabyte generated Python modules.
Source and license: Zapier… See the full description on the dataset page: https://huggingface.co/datasets/UpstageShareSpace/automationBench.details_upstage__SOLAR-10.7B-v1.0
Dataset Card for Evaluation run of upstage/SOLAR-10.7B-v1.0
Dataset automatically created during the evaluation run of model upstage/SOLAR-10.7B-v1.0.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_upstage__SOLAR-10.7B-v1.0.upstage-arc-cot_50kCReSt
Dataset for CReSt
This repository contains the dataset used in the paper "CReSt: A Comprehensive Benchmark for Retrieval-Augmented Generation with Complex Reasoning over Structured Documents".
📂 Dataset Overview
This dataset consists of the following two subsets:
refusal: The subset that contains only refusal cases
non_refusal: The subset that contains only non-refusal cases
🚀 How to Use
You can access detailed usage instructions in the CReSt GitHub… See the full description on the dataset page: https://huggingface.co/datasets/upstage/CReSt.Pretraining_Datasetupstage-arc-cot_20kupstage__SOLAR-10.7B-v1.0-details
Dataset Card for Evaluation run of upstage/SOLAR-10.7B-v1.0
Dataset automatically created during the evaluation run of model upstage/SOLAR-10.7B-v1.0
The dataset is composed of 83 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 23 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/upstage__SOLAR-10.7B-v1.0-details.upstage__SOLAR-10.7B-Instruct-v1.0-details
Dataset Card for Evaluation run of upstage/SOLAR-10.7B-Instruct-v1.0
Dataset automatically created during the evaluation run of model upstage/SOLAR-10.7B-Instruct-v1.0
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/upstage__SOLAR-10.7B-Instruct-v1.0-details.upstage_preprocessed_datasetupstage__solar-pro-preview-instruct-details
Dataset Card for Evaluation run of upstage/solar-pro-preview-instruct
Dataset automatically created during the evaluation run of model upstage/solar-pro-preview-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/upstage__solar-pro-preview-instruct-details.upstage-arc-cot_5k_512_1024upstage-arc-mc-20k-aeupstage__SOLAR-10.7B-Instruct-v1.0upstage__solar-pro-preview-instructwmlu-koarenahard-0-1upstage-arc-cot_20k_4096upstage-arc-cot_168kupstage-arc-cot_10k_512_1024upstage-arc-mc-1k-at-arconlyupstage-arc-cot_50k_4096upstage-arc-mc-3k-at-arconlyupstage-arc-mc-20kupstage_lora_finetuning_practicewmlu-arenahard-2-0upstage-arc-mc-20k-at
