datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dsd-llm-datasetKorean-YouTube-Comment-Sentiment-Dataset
Korean YouTube Comment Sentiment Dataset
Data Overview
Summary
본 데이터셋은 유튜브에서 수집된 한국어 댓글 5,482개와 이에 대응하는 감정 레이블(긍정, 부정, 중립, 불명확)로 구성된 감정 분류용 데이터셋입니다.
주요 레이블: 긍정, 부정, 중립, 불명확
Features
수집 대상: 요리, 뷰티, 게임, 여행, 쇼핑 등 분야의 10만 명 이상 구독자를 보유한 유튜브 채널
형식: JSON (id, text, label)
검수: 한국인 검수자에 의한 수작업 라벨링 및 교차 검토
본 데이터셋은 구어체, 이모지, 줄임말 등 실제 사용자 표현이 반영되어 있습니다.
Dataset Structure
Dataset Fields
Field
Type
Description
id
string
각 댓글의… See the full description on the dataset page: https://huggingface.co/datasets/LLM-SocialMedia/Korean-YouTube-Comment-Sentiment-Dataset.Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset
📌 Dataset Contents
Each sample includes:
category: The evaluation domain
prompt: The question given to the LLM
temperature: Environmental temperature input
humidity: Environmental humidity input
context: A scenario label (e.g., cool_humid, hot_dry, average_day)
reference: Expert-crafted expected output
All data is provided in a single JSON file.
🧪 Intended Use
This dataset supports research on:
LLM evaluation methods (semantic similarity, contextual… See the full description on the dataset page: https://huggingface.co/datasets/wayne-redemption/Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset.DPO-Dataset
AMALIA DPO Dataset
This is the DPO (preference optimization) dataset used to train AMALIA-DPO.
It is a mix of preference pairs, mainly in European Portuguese and English, covering general conversation, instruction following, math, and safety. These pairs come from different sources, including prompts from the SFT mix, responses generated by different models, including an early version of the model, and some public datasets.
This dataset is provided as part of the AMALIA project… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/DPO-Dataset.
