datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
preprocessing_rebuttal
CORA preprocessing Test 500 - no, general, CORA specific
데이터셋 개요
이 데이터셋은 CORA test split에서 500개 target-audio row를 균형 샘플링한 text-only subset입니다. 각 row는 하나의 target audio item에 대응하며, CORA에서 제공하는 5개 query type을 같은 row 안에 유지합니다.
각 원본 query에는 두 가지 전처리 방식의 rewrite 결과가 함께 제공됩니다.
general: query type 정보를 사용하지 않는 일반적인 audio retrieval query rewriting
cora_aware: CORA query type 정보를 사용해 caption-style audio description으로 정규화하는 rewriting
이 데이터셋에는 audio file이 포함되어 있지 않습니다. CORA… See the full description on the dataset page: https://huggingface.co/datasets/msnowchanj/preprocessing_rebuttal.mcmd_race_sample_no_preprocessingtest_mcmd_race_sample_no_preprocessing
