datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLM_for_Dietary_Recommendation_System
Chat_GPT_for_Nutritional_Recommendation_System
This repository contains the code, 50 different patient profiles, and respective Chat-GPT responses with nutritional recommendations and sample diet plans. Patient profiles contain easy, medium, and complex cases and various diseases. The project aims to evaluate the application of large language models for nutritional recommendation systems.
Responses were evaluated based on personalization, consistency with evidence-based… See the full description on the dataset page: https://huggingface.co/datasets/issai/LLM_for_Dietary_Recommendation_System.olga_mcarrow_writer_artist_books_and_music_recommendations_scifi_fantasy
Gifts of Wandering Ice – Full Dataset from llms.txt
This dataset contains all information from https://giftscomic.com/llms.txt, structured for AI assistants, RAG systems, recommendation engines, and discovery tools.
The goal is to help AI systems accurately recommend and describe the free science-fiction webcomic Gifts of Wandering Ice and the broader creative work of its author, Olga Makarova (Mildegard).
Dataset Contents (358 records)
Split
Records… See the full description on the dataset page: https://huggingface.co/datasets/Mildegard/olga_mcarrow_writer_artist_books_and_music_recommendations_scifi_fantasy.turkish-electronics-product-comparison-recommendation
🇹🇷 Türkçe Elektronik Ürün Karşılaştırma ve Öneri Chat Veri Seti
Veri Seti Açıklaması
Bu veri seti, büyük dil modellerinin Türkçe elektronik ürün karşılaştırma ve ürün öneri görevlerinde eğitilmesi amacıyla hazırlanmış bir konuşma veri setidir.
Veri seti iki temel görev grubundan oluşmaktadır:
Elektronik ürün karşılaştırma
Elektronik ürün önerisi
Veri seti Supervised Fine-Tuning (SFT) ve instruction tuning çalışmalarında kullanılabilecek chat formatında… See the full description on the dataset page: https://huggingface.co/datasets/sedayzc/turkish-electronics-product-comparison-recommendation.LGUplus_channel_recommendation
LGUplus_channel_recommendation
페르소나 기반 채널 추천 데이터셋 (추천 top5에서 파생). 350m 페르소나 모델의 <|chan|> 태스크 학습용.
(persona, date, 2h block) 마다: 컨텍스트를 보고 후보 채널(≤20) 중 그 시간대에 볼 채널 5개를 고른다.
태스크 I/O
입력: 주간 페르소나(전주, key_keywords 제외: viewing_tendency/preferred_genres/preferred_programs/frequent_channels) + 직전일 일일페르소나 서술 + 직전 2시간 시청(총 15분+) + 후보 채널(≤20)
출력: 채널 5개 (label_channels)
후보 채널 = 추천 20개 프로그램 후보(A∪B)의 채널 dedup(≤20). 정답 = 교사가 뽑은 top5 프로그램의 채널 dedup(≤5).
구성… See the full description on the dataset page: https://huggingface.co/datasets/ENERZAiKR/LGUplus_channel_recommendation.LGUplus_recommendation_top5_rolling
LGUplus TV 추천 (rolling-persona 기반, self-contained · 채널급 포함)
(persona, date, block) 마다 후보 프로그램 A∪B(≤20) 중 교사 LLM(Qwen3-235B-A22B-Instruct-2507-FP8)이
그 시청자가 그 시간대에 '실제로 볼' top5 를 고른 결과.
persona 컨텍스트: jungsanghyun/lgu-rolling-persona
의 전일까지 갱신된 롤링 페르소나(장기 성향 + 최근 성향 + 고정 시청 프로그램 표). 추천대상 날짜의 직전 스냅샷을 쓰며, 그날 시청이 없으면 더 과거 스냅샷으로 백워크한다.
업데이트: 이제 각 행이 입력을 포함(self-contained) 하며, program-level picks 에 더해 채널급(cand_channels/label_channels) 도 제공한다.
필드
필드
설명… See the full description on the dataset page: https://huggingface.co/datasets/ENERZAiKR/LGUplus_recommendation_top5_rolling.LGUplus_recommendation_candidates
LGUplus_recommendation_candidates
LG U+ 페르소나 기반 TV 추천 파이프라인의 1차 후보 목록(A∪B, combo당 최대 20개)과
후보 id를 프로그램 상세로 조인하기 위한 프로그램 테이블입니다.
최종 top5 선정(교사 LLM)은 이 후보에서 5개를 고르는 다음 단계이며, 여기엔 후보까지만 포함합니다.
구성
하루를 12개 2시간 블록으로 나눔. (persona_id, date, block) 마다 후보 2종:
cand_stat (후보A · 시청통계): 기반주에 그 페르소나가 가장 많이 본 채널 top10, 각 채널에서 그 블록에 가장 먼저 시작하는 프로그램 1개 (≤10).
cand_persona (후보B · 주간페르소나): 그 블록 프로그램을 주간페르소나 프로필과 ko-sroberta 임베딩 유사도 + 시청장르 비례배분으로 뽑은 10개.
두 후보는 겹칠 수 있으며 dedup 후 A∪B ≤… See the full description on the dataset page: https://huggingface.co/datasets/ENERZAiKR/LGUplus_recommendation_candidates.
