CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenOneRec /Explorer_LLM_Rec_Competition Explorer_LLM_Rec_Competition This dataset contains user historical behaviors and content metadata, constructed from real user interaction histories. It covers behavior sequences across multiple domains for a single user and supports cross-domain recommendation, semantic-ID retrieval / generation, content understanding, and related tasks. Files File Description OneReason_UserProfile/ Per-user multi-domain behavior records (~500k rows)… See the full description on the dataset page: https://huggingface.co/datasets/OpenOneRec/Explorer_LLM_Rec_Competition.21 likes2.2k downloads3mo agoHugging Face02akk666 /Explorer_LLM_Rec_Competition Explorer_LLM_Rec_Competition This dataset contains user historical behaviors and content metadata, constructed from real user interaction histories. It covers behavior sequences across multiple domains for a single user and supports cross-domain recommendation, semantic-ID retrieval / generation, content understanding, and related tasks. Files File Description OneReason_UserProfile/ Per-user multi-domain behavior records (~500k rows)… See the full description on the dataset page: https://huggingface.co/datasets/akk666/Explorer_LLM_Rec_Competition.0 likes234 downloads3mo agoHugging Face03Breakfeeling /Explorer_LLM_Rec_Competition Explorer_LLM_Rec_Competition This dataset contains user historical behaviors and content metadata, constructed from real user interaction histories. It covers behavior sequences across multiple domains for a single user and supports cross-domain recommendation, semantic-ID retrieval / generation, content understanding, and related tasks. Files File Description OneReason_UserProfile/ Per-user multi-domain behavior records (~500k rows)… See the full description on the dataset page: https://huggingface.co/datasets/Breakfeeling/Explorer_LLM_Rec_Competition.0 likes136 downloads3mo agoHugging Face04Richardzzl /Explorer_LLM_Rec_Competition Explorer_LLM_Rec_Competition This dataset contains user historical behaviors and content metadata, constructed from real user interaction histories. It covers behavior sequences across multiple domains for a single user and supports cross-domain recommendation, semantic-ID retrieval / generation, content understanding, and related tasks. Files File Description OneReason_UserProfile/ Per-user multi-domain behavior records (~500k rows)… See the full description on the dataset page: https://huggingface.co/datasets/Richardzzl/Explorer_LLM_Rec_Competition.0 likes114 downloads3mo agoHugging Face05bolobolooo /Explorer_LLM_Rec_Competition Explorer_LLM_Rec_Competition This dataset contains user historical behaviors and content metadata, constructed from real user interaction histories. It covers behavior sequences across multiple domains for a single user and supports cross-domain recommendation, semantic-ID retrieval / generation, content understanding, and related tasks. Files File Description OneReason_UserProfile/ Per-user multi-domain behavior records (~500k rows)… See the full description on the dataset page: https://huggingface.co/datasets/bolobolooo/Explorer_LLM_Rec_Competition.0 likes105 downloads3mo agoHugging Face06akk666 /Explorer_LLM_Rec_Competition2 Explorer_LLM_Rec_Competition This dataset contains user historical behaviors and content metadata, constructed from real user interaction histories. It covers behavior sequences across multiple domains for a single user and supports cross-domain recommendation, semantic-ID retrieval / generation, content understanding, and related tasks. Files File Description OneReason_UserProfile/ Per-user multi-domain behavior records (~500k rows)… See the full description on the dataset page: https://huggingface.co/datasets/akk666/Explorer_LLM_Rec_Competition2.0 likes88 downloads3mo agoHugging Face07YangXiangWu /Explorer_LLM_Rec_Competition Explorer_LLM_Rec_Competition This dataset contains user historical behaviors and content metadata, constructed from real user interaction histories. It covers behavior sequences across multiple domains for a single user and supports cross-domain recommendation, semantic-ID retrieval / generation, content understanding, and related tasks. Files File Description OneReason_UserProfile/ Per-user multi-domain behavior records (~500k rows)… See the full description on the dataset page: https://huggingface.co/datasets/YangXiangWu/Explorer_LLM_Rec_Competition.0 likes78 downloads3mo agoHugging Face08williamhkml /Explorer_LLM_Rec_Competition Explorer_LLM_Rec_Competition This dataset contains user historical behaviors and content metadata, constructed from real user interaction histories. It covers behavior sequences across multiple domains for a single user and supports cross-domain recommendation, semantic-ID retrieval / generation, content understanding, and related tasks. Files File Description OneReason_UserProfile/ Per-user multi-domain behavior records (~500k rows)… See the full description on the dataset page: https://huggingface.co/datasets/williamhkml/Explorer_LLM_Rec_Competition.0 likes67 downloads2mo agoHugging Face09weblab-llm-competition-2025-bridge /team-watanabe-MegaSciencetext100K<n<1M1 likes35 downloads1y agoHugging Face10weblab-llm-competition-2025-bridge /neko-prelim-HLE_SFT_OlymMATHtextn<1K0 likes34 downloads11mo agoHugging Face11weblab-llm-competition-2025-bridge /neko-prelim-dna_dpo_hh-rlhf neko-prelim-dna_dpo_hh-rlhf データセットの説明 このデータセットは、以下の分割(split)ごとに整理された処理済みデータを含みます。 train: 1 JSON files, 1 Parquet files データセット構成 各 split は JSON 形式と Parquet 形式の両方で利用可能です: JSONファイル: 各 split 用サブフォルダ内の元データ(train/) Parquetファイル: split名をプレフィックスとした最適化データ(data/train_*.parquet) 各 JSON ファイルには、同名の split プレフィックス付き Parquet ファイルが対応しており、大規模データセットの効率的な処理が可能です。 使い方 from datasets import load_dataset # 特定の split を読み込む train_data =… See the full description on the dataset page: https://huggingface.co/datasets/weblab-llm-competition-2025-bridge/neko-prelim-dna_dpo_hh-rlhf.text10K<n<100K0 likes32 downloads11mo agoHugging Face12weblab-llm-competition-2025-bridge /RAMEN-phase1text10K<n<100K0 likes31 downloads1y agoHugging Face13weblab-llm-competition-2025-bridge /neko-prelim-HLE_SFT_OlympiadBenchtextn<1K0 likes26 downloads11mo agoHugging Face14weblab-llm-competition-2025-bridge /team-truthowl-mixed-reasoning-dataset Team P11 Mixed Reasoning Dataset 📊 Dataset description HLE(Humanity's Last Exam)向けに作成した、数学中心+科学MCの混合推論データセットです。 推論過程(Chain-of-Thought)を保持し、最終解答の正規化を行っています。 対象モデルは DeepSeek-R1-Distill-Qwen-32B、学習はQLoRAを想定しています。 🎯 Purpose Competition: 松尾研LLMコンペ 2025 Target Model: DeepSeek-R1-Distill-Qwen-32B Training Method: QLoRA Fine-tuning(4bit NF4, double quant) 📦 Composition Math Hard(MATH Level≥3, HARDMath) Math Mid(GSM8K, MetaMathQA) Science(GPQA… See the full description on the dataset page: https://huggingface.co/datasets/weblab-llm-competition-2025-bridge/team-truthowl-mixed-reasoning-dataset.texttext-generation10K<n<100K0 likes24 downloads11mo agoHugging Face15weblab-llm-competition-2025-bridge /neko-prelim-dna_vanilla_harmful_v2 neko-prelim-dna_vanilla_harmful_v2 データセットの説明 このデータセットは、以下の分割(split)ごとに整理された処理済みデータを含みます。 train: 1 JSON files, 1 Parquet files データセット構成 各 split は JSON 形式と Parquet 形式の両方で利用可能です: JSONファイル: 各 split 用サブフォルダ内の元データ(train/) Parquetファイル: split名をプレフィックスとした最適化データ(data/train_*.parquet) 各 JSON ファイルには、同名の split プレフィックス付き Parquet ファイルが対応しており、大規模データセットの効率的な処理が可能です。 使い方 from datasets import load_dataset # 特定の split を読み込む train_data =… See the full description on the dataset page: https://huggingface.co/datasets/weblab-llm-competition-2025-bridge/neko-prelim-dna_vanilla_harmful_v2.text1K<n<10K0 likes23 downloads11mo agoHugging Face16weblab-llm-competition-2025-bridge /neko-prelim-HLE_SFT_LIMOtextn<1K0 likes23 downloads11mo agoHugging Face17weblab-llm-competition-2025-bridge /neko-prelim-HLE_SFT_PhysReasontextn<1K0 likes22 downloads11mo agoHugging Face18weblab-llm-competition-2025-bridge /neko-prelim-HLE_SFT_OpenMathReasoningtext1K<n<10K0 likes21 downloads11mo agoHugging Face19weblab-llm-competition-2025-bridge /neko-prelim-HLE_SFT_MixtureOfThoughts0 likes20 downloads11mo agoHugging Face20weblab-llm-competition-2025-bridge /neko-prelim-HLE_SFT_GPQA_Diamondtextn<1K0 likes20 downloads11mo agoHugging Face21weblab-llm-competition-2025-bridge /MedMCQA MedMCQA-CoT: 医学多肢選択問題with Chain-of-Thought推論 データセット概要 MedMCQA-CoTは、MedMCQAデータセットの拡張版で、各医学多肢選択問題に高品質なChain-of-Thought(CoT)推論を追加したデータセットです。医学的な推論プロセスを説明できるAIシステムの開発を支援することを目的としています。 主な特徴 2,020件の医学MCQ問題 - 元のMedMCQAデータセットから抽出 Chain-of-Thought推論 - DeepSeek-R1モデルで生成 95.5%の回答精度 - 生成されたCoTが正解に導く割合 0.952の平均品質スコア - 医学用語密度と推論品質に基づく評価 包括的なメタデータ - 品質スコア、医学専門分野、生成統計を含む データセット詳細 各レコードの構成: question: MedMCQAからの元の医学問題 answer: 正解の選択肢(A, B, C, D) cot:… See the full description on the dataset page: https://huggingface.co/datasets/weblab-llm-competition-2025-bridge/MedMCQA.textquestion-answering1K<n<10K0 likes19 downloads1y agoHugging Face22weblab-llm-competition-2025-bridge /neko-prelim-HLE_SFT_OpenThoughts-114ktext1K<n<10K0 likes19 downloads11mo agoHugging Face23weblab-llm-competition-2025-bridge /team-camino-Omni-MATH_difficulty5plus_qatext1K<n<10K0 likes19 downloads1y agoHugging Face24weblab-llm-competition-2025-bridge /neko-prelim-HLE_SFT_PHYBenchtextn<1K0 likes18 downloads11mo agoHugging Face25weblab-llm-competition-2025-bridge /team-pont-neuf-sft-dataset-25101 likes18 downloads10mo agoHugging Face26weblab-llm-competition-2025-bridge /team-watanabe-AoPs-Instructtext10K<n<100K1 likes17 downloads1y agoHugging Face27weblab-llm-competition-2025-bridge /neko-prelim-HLE_RL_Olympiadbench-v2textn<1K0 likes17 downloads11mo agoHugging Face28weblab-llm-competition-2025-bridge /neko-prelim-HLE_SFT_LIMO-v2textn<1K0 likes16 downloads11mo agoHugging Face29weblab-llm-competition-2025-bridge /neko-prelim-wj-vanilla_benign_v2 neko-prelim-wj-vanilla_benign_v2 データセットの説明 このデータセットは、以下の分割(split)ごとに整理された処理済みデータを含みます。 train: 1 JSON files, 1 Parquet files データセット構成 各 split は JSON 形式と Parquet 形式の両方で利用可能です: JSONファイル: 各 split 用サブフォルダ内の元データ(train/) Parquetファイル: split名をプレフィックスとした最適化データ(data/train_*.parquet) 各 JSON ファイルには、同名の split プレフィックス付き Parquet ファイルが対応しており、大規模データセットの効率的な処理が可能です。 使い方 from datasets import load_dataset # 特定の split を読み込む train_data =… See the full description on the dataset page: https://huggingface.co/datasets/weblab-llm-competition-2025-bridge/neko-prelim-wj-vanilla_benign_v2.text1K<n<10K0 likes15 downloads11mo agoHugging Face30weblab-llm-competition-2025-bridge /team-watanabe-UGPhysicstext10K<n<100K1 likes13 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.