datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMRC_Real_World_Conversation
MMRC - Multi-Modal Open-Ended Conversation Dataset
Overview:
MMRC is a benchmark dataset designed for evaluating Multi-Modal Large Language Models (MLLMs) in open-ended, multi-turn conversations. It provides diverse, real-world conversational data that integrates both textual and visual modalities, aiming to push the boundaries of MLLM performance in practical settings.
Dataset Details:
The MMRC dataset is composed of multi-turn conversations with integrated… See the full description on the dataset page: https://huggingface.co/datasets/WUUE/MMRC_Real_World_Conversation.bmvs_sparse_dtu2my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw - Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/wuuski/my-personal-codex-data.GUI-CIDERSaliTraptasks_list_think
基本說明
資料集內容:思考能力的 tasks list, 以及評測思考能力的題目
資料集版本: 20250225
本資料集運用 Gemini 2.0 Flash Thinking Experimental 01-21 生成
思考要如何建構推理模型,於是先跟 LLM 討論何謂推理,有哪些樣態,之後根據可能的類別請 LLM 生成 tasks list , 結果就是檔案 taskslist.csv
根據每一個 task 描述,再去生成不同難度的題目,就是 validation.csv
taskslist.csv
目前分六大類,共 85 子類
ID 編碼說明
都是 TK 開頭
L2.2 就是 TK22
目前後面兩碼是序號
group 說明
代號
說明
L1.1
基礎推理類型
L2.1
KR&R 推理
L3.1
深度
L4.1
領域
L5.1
推理呈現能力
L6.1
元推理
validation.csv
id:對應的 task id, 定義在… See the full description on the dataset page: https://huggingface.co/datasets/wuulong/tasks_list_think.purchasing_exam_questions
資料來源:採購法規題庫
資料產生日期:114/03/07
大部分項目內本沒有 「依據法源」欄位,為求統一所以有欄位,為空值
順手用 colab 觀察資料內容:採購網題庫1.ipynb
highresdtutest
