datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenOrca-Top5percent🐋 The OpenOrca-Top5Percent Dataset! 🐋
We are excited to introduce the OpenOrca-Top5Percent dataset, a refined version of the original OpenOrca dataset. This dataset contains only those entries which utilize the top 5% most frequently used words in the OpenOrca dataset, aiming to focus on high-frequency vocabulary for various NLP tasks.
Dataset Summary
The OpenOrca-Top5Percent dataset is a curated subset of the augmented FLAN Collection data, focusing specifically on entries that… See the full description on the dataset page: https://huggingface.co/datasets/dynopii/OpenOrca-Top5percent.domain-agnostic-reasoning-traces-balanced-top50-v1
BOTCOIN Balanced Top-50 Reasoning Traces
This public dataset contains enriched BOTCOIN reasoning-trace attempts selected
from canonical dataset/v2 research-ready objects.
Selection policy:
Source only attempts/research-ready objects.
Rank each domain by trace_quality.reasoning_trace_quality_score.
Keep each domain's top 50 percent.
Equalize domains to the smallest top-half count.
The rows are self-contained and intentionally rich: prompt/messages… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-reasoning-traces-balanced-top50-v1.Bluemoon_Top50MB_Sorted_Fixed_ja
Bluemoon_Top50MB_Sorted_Fixed_ja
SicariusSicariiStuff/Bluemoon_Top50MB_Sorted_Fixedを、GENIAC-Team-Ozaki/karakuri-lm-8x7b-chat-v0.1-awqを用いて日本語に翻訳したロールプレイ学習用データセットです。
LLMの推論にはDeepInfraというサービスを使いました。
翻訳の詳細
3-shots promptingでの翻訳
mistralのtokenizerで出力が8000トークンを超えるまで翻訳
元データセットにある非常に長い対話は上記条件で途中のターンで翻訳を終了しています。
LLM特有の同じ出力が繰り返される現象に遭遇した場合、その時点で該当レコードの翻訳を終了
この結果1ターン未満となったレコード(157件)を削除… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Bluemoon_Top50MB_Sorted_Fixed_ja.LGUplus_recommendation_top5_rolling
LGUplus TV 추천 (rolling-persona 기반, self-contained · 채널급 포함)
(persona, date, block) 마다 후보 프로그램 A∪B(≤20) 중 교사 LLM(Qwen3-235B-A22B-Instruct-2507-FP8)이
그 시청자가 그 시간대에 '실제로 볼' top5 를 고른 결과.
persona 컨텍스트: jungsanghyun/lgu-rolling-persona
의 전일까지 갱신된 롤링 페르소나(장기 성향 + 최근 성향 + 고정 시청 프로그램 표). 추천대상 날짜의 직전 스냅샷을 쓰며, 그날 시청이 없으면 더 과거 스냅샷으로 백워크한다.
업데이트: 이제 각 행이 입력을 포함(self-contained) 하며, program-level picks 에 더해 채널급(cand_channels/label_channels) 도 제공한다.
필드
필드
설명… See the full description on the dataset page: https://huggingface.co/datasets/ENERZAiKR/LGUplus_recommendation_top5_rolling.
