upstage
Datasets
All datasets matching “upstage”dp-bench
DP-Bench: Document Parsing Benchmark
Document parsing refers to the process of converting complex documents, such as PDFs and scanned images, into structured text formats like HTML and Markdown.
It is especially useful as a preprocessor for RAG systems, as it preserves key structural information from visually rich documents.
While various parsers are available on the market, there is currently no standard evaluation metric to assess their performance.
To address this gap, we… See the full description on the dataset page: https://huggingface.co/datasets/upstage/dp-bench.details_upstage__SOLAR-10.7B-v1.0
Dataset Card for Evaluation run of upstage/SOLAR-10.7B-v1.0
Dataset automatically created during the evaluation run of model upstage/SOLAR-10.7B-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_upstage__SOLAR-10.7B-v1.0.scenarios
ISD-Bench-Ko (v2.0)
교수설계(Instructional Systems Design) AI 에이전트 평가용 한국어 벤치마크 데이터셋입니다.
📊 데이터셋 개요
항목
내용
총 시나리오
25,795개
Train
24,593개 (95.3%)
Test
1,202개 (4.7%)
언어
한국어
라이선스
Apache 2.0
🆕 v2.0 업데이트 내역
Train/Test 분리: 층화 추출(Stratified Sampling)으로 9개 축 분포 유지
품질 개선:
context_variant 논리적 일관성 개선 (#97)
difficulty 균형화 (33%:34%:33%) (#99)
전체 데이터 품질 재검증 (#100)
폴더 구조 변경: train/part1~4로 분리 (HF 파일 수 제한 대응)
📁 데이터 구조
scenarios/
├──… See the full description on the dataset page: https://huggingface.co/datasets/upstage-isd-agent/scenarios.bench-automationbench
AutomationBench public task payloads
Materialized public task inputs for Zapier AutomationBench,
pinned to upstream revision 4a8e1061254004d9dac807054eed33fad7d1ff14.
This repository contains six Parquet files (100 tasks per public domain) generated from the
upstream get_<domain>_dataset() functions. It exists so the solar-system evaluation integration
can fetch immutable task payloads without committing multi-megabyte generated Python modules.
Source and license: Zapier… See the full description on the dataset page: https://huggingface.co/datasets/hyeonseop-upstage/bench-automationbench.details_upstage__SOLAR-0-70b-16bit
Dataset Card for Evaluation run of upstage/SOLAR-0-70b-16bit
Dataset Summary
Dataset automatically created during the evaluation run of model upstage/SOLAR-0-70b-16bit on the Open LLM Leaderboard.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_upstage__SOLAR-0-70b-16bit.details_upstage__SOLAR-10.7B-Instruct-v1.0
Dataset Card for Evaluation run of upstage/SOLAR-10.7B-Instruct-v1.0
Dataset automatically created during the evaluation run of model upstage/SOLAR-10.7B-Instruct-v1.0.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_upstage__SOLAR-10.7B-Instruct-v1.0.
