datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clinc150
clinc150
This is a text classification dataset. It is intended for machine learning research and experimentation.
This dataset is obtained via formatting another publicly available data to be compatible with our AutoIntent Library.
Usage
It is intended to be used with our AutoIntent Library:
from autointent import Dataset
banking77 = Dataset.from_hub("AutoIntent/clinc150")
Source
This dataset is taken from cmaldona/All-Generalization-OOD-CLINC150 and formatted… See the full description on the dataset page: https://huggingface.co/datasets/DeepPavlov/clinc150.clinc150clinc150_subsetovos-intent-template-bench-clinc150
OVOS intent_template bench — clinc150
Per-sample predictions of the template-paradigm intent league fighters of the
OVOS Plugin Arena over
clinc/clinc_oos.
One dedicated repo per benchmark modality; one dataset split per language;
one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl.
Rows follow the arena §3.2 contract (pinned dataset_revision,
plugin_version, fired pipeline stage, exact_match with correct-OOD
semantics). Produced by the reproducible… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-template-bench-clinc150.ovos-intent-bench-clinc150
OVOS intent bench — clinc150
Per-sample predictions of the open intent league (mixed-paradigm pipeline fusions) fighters of the
OVOS Plugin Arena over
clinc/clinc_oos.
One dedicated repo per benchmark modality; one dataset split per language;
one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl.
Rows follow the arena §3.2 contract (pinned dataset_revision,
plugin_version, fired pipeline stage, exact_match with correct-OOD
semantics). Produced by the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-clinc150.clinc150-ko
CLINC150-ko (한국어 번역 버전)
개요
이 데이터셋은 DeepPavlov/clinc150 데이터셋의 한국어 번역 버전입니다.
CLINC150은 150개의 인텐트(의도) 분류와 Out-of-Scope(범위 외) 탐지를 위한 대규모 벤치마크 데이터셋입니다. 다양한 도메인(banking, travel, kitchen, work, auto 등)의 사용자 질의를 포함하며, 대화형 AI 시스템의 의도 분류 및 범위 외 탐지 성능 평가에 널리 사용됩니다.
이 번역 버전은 한국어 텍스트 분류, 인텐트 분류 모델의 학습 및 평가, 한국어 챗봇 개발 등에 활용할 수 있도록 제작되었습니다.
데이터셋 정보
항목
내용
원본 데이터셋
DeepPavlov/clinc150
원본 출처
An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction… See the full description on the dataset page: https://huggingface.co/datasets/neuralfoundry-coder/clinc150-ko.clinc_150Generalization-MultiClass-CLINC150-ROSTDThis dataset merge 3 datasets and have two setup for experiments in generalisation for multi-class clasificacitino task.
ID, near-OOD, covariate-shitf: CLINC150
ID, near-OOD, covariate-shitf: ROSTD+OOD (fbreleasecoarse version)
far-OOD Validation: SST2
far-OOD Test: News Category (v3)
clinc150_ru
Russian clinc150
This is a text classification dataset. It is intended for machine learning research and experimentation.
This dataset is obtained via formatting another publicly available data to be compatible with our AutoIntent Library.
Usage
It is intended to be used with our AutoIntent Library:
from autointent import Dataset
clinc150_ru = Dataset.from_hub("AutoIntent/clinc150_ru")
Source
This dataset is taken from private github repository… See the full description on the dataset page: https://huggingface.co/datasets/DeepPavlov/clinc150_ru.All-Generalization-OOD-CLINC150Datasets structure.
Attributes:
data: text
labels: class (str)
domain: parent class (str) - This attribute signifies the parent class in the hierarchy and may be absent in some datasets.
generalisation: type of OOD
Splits:
Train:
ID: Clinc150
near-OOD: Clinc150
far-OOD: Yelp
Validation:
ID: Clinc150
near-OOD: Clinc150
far-OOD: SST2
Test:
ID: Clinc150
near-OOD: Clinc150
far-OOD: NewCategoryV3
cov-shift: ROSTD+
ovos-intent-bench-clinc150-train
ovos-intent-bench-clinc150-train
An OVOS Plugin Arena benchmark repository. The arena's prediction runner
publishes each plugin's raw output on the clinc150-train dataset here, one JSON-lines
file per plugin under predictions/<lang>/, and where a sample-set manifest
governs scoring it lives under sample_sets/. The public leaderboard at
https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows,
and anyone can recompute an entry from them without access to the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-clinc150-train.cls_clinc150_TextVsDomain__BaseDefaultcls_clinc150_TextVsIntent__BaseDefaultclinc150_esclinc150_frclinc150_aug_qwen2.5-7b-awq
