datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ogiri-bokete
読み込み方
from datasets import load_dataset
dataset = load_dataset("YANS-official/ogiri-bokete", split="train")
概要
大喜利投稿サイトBoketeのクロールデータです。元データは CLoT-Oogiri-Go [Zhang+ CVPR2024]というデータの一部です。
詳細はCVPRのプロジェクトページをご確認ください。
このデータは以下の3タスクが含まれます。
text_to_text: テキストでお題が渡され、それに対する回答を返します。
image_to_text: いわゆる「画像で一言」です。画像のみが渡されて、テキストによる回答を返します。
text_image_to_text: 画像中にテキストが書かれています。テキストの一部が空欄になっているので、そこに穴埋めする形で回答を返します。
それぞれの量は以下の通りです。(8/30現在。ハッカソン当日までに増やす可能性があります。)
タスク… See the full description on the dataset page: https://huggingface.co/datasets/YANS-official/ogiri-bokete.refute
Can AI read new science honestly?
Models can sound convincing while misreading a result or expressing more confidence than the evidence deserves. That matters when people use them to summarize papers, compare studies, or decide what to investigate next.
REFUTE tests whether a model knows the finding, spots quiet flaws, names what would overturn a claim, and matches its confidence to the evidence.
Truth Score is the main result. It combines factual accuracy, flaw… See the full description on the dataset page: https://huggingface.co/datasets/BGPT-OFFICIAL/refute.ontario-hansard-official
Ontario Hansard (Official) — speaker-identity-resolved
~4.3 million paragraph-level records from the official Hansard of the Legislative
Assembly of Ontario, parliaments 29–44 (1974–2025), with each paragraph's
speaker resolved to a stable member identity and tagged with the confidence tier
of that resolution.
Overview
4,316,254 records, one per Hansard paragraph.
Parliaments 29–44, sessions spanning 1974–2025.
51 years of floor proceedings: debates, question… See the full description on the dataset page: https://huggingface.co/datasets/agoulah/ontario-hansard-official.aime24-official
AIME 2024 — official wording, figures retained
All 30 problems from the 2024 American Invitational Mathematics Examination (AIME I and AIME II),
transcribed from the official exam text with every figure retained as Asymptote source.
This exists because the circulating text-only versions of AIME 2024 are not faithful to the
official problems, and at least one problem in them cannot be solved as written.
Why this dataset exists
While evaluating a reasoning model on… See the full description on the dataset page: https://huggingface.co/datasets/YichengWangCA/aime24-official.ctext
说明
ctext - all - slice:自主收集的文言文繁体语料库,主要爬取自维基文库、漢川草廬等处
ctext - 副本 - 副本:ctext - all - slice的先秦部分
ctext - 白话:自主收集的白话及现代汉语繁体语料库,主要爬取自维基百科、BWIKI、维基文库、知乎、繁體中文書庫等处
项目链接
deepscaler-teacher-sft-vllm-official-40k
DeepScaleR teacher SFT vLLM official 40k
Generated run: exp_003_vllm_official_brainlab_2gpu.
Summary
{
"num_examples": 40300,
"sft_dir": "data/processed/deepscaler/teacher_sft/exp_003_vllm_official_brainlab_2gpu",
"parse_rate": 0.9999751861042183,
"correct_rate": 0.5728039702233251,
"format_rate": 0.005955334987593052,
"mean_reward": 0.42432258064534184,
"deepscaler_mean_reward": 0.6266997518610422,
"deepscaler_match_mean_reward":… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.dataagentbench-derived-official54-altimate-prompt-shell
DataAgentBench-Derived Official54 Altimate Prompt Shell
This public dataset contains 54 DataAgentBench-derived task prompt-shell rows compiled for the Altimate-centered DAB sandbox runtime. It is a derived compile/export artifact, not the official raw DataAgentBench release.
It is intended as a portable task/prompt/manifest source for downstream Altimate teacher rollouts, SFT construction, or RL data compilation. It is not a completed rollout dataset: the rows do not contain… See the full description on the dataset page: https://huggingface.co/datasets/forseasons/dataagentbench-derived-official54-altimate-prompt-shell.deepscaler-teacher-sft-vllm-official-40k-clean-v2
DeepScaleR Teacher40k Clean v2
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
minimum official reward: 1.0
maximum text tokens: 8192
maximum response chars: 65000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem dedupe
Counts
raw examples: 40300
kept examples: 21727
train examples: 21292
val… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v2.deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter
deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
reward filter enabled: False
minimum official reward: 1.0
scoring errors rejected: False
maximum text tokens: 8192
maximum response chars: 65000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter.deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual
deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
reward filter enabled: False
minimum official reward: 1.0
scoring errors rejected: False
maximum text tokens: 32768
maximum response chars: 200000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem dedupe… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual.Reasoning-MiniGPT-brazilian-portugueseAMBILE_Shah_Jo_Risalo_Labeled
AMBILE Shah Jo Risalo
Developed by:Abdul Majid Bhurgri Institute of Language Engineering (AMBILE), HyderabadUnder the administrative control of the Culture, Tourism, Antiquities & Archives Department, Government of Sindh
Dataset Overview
The "Shah Jo Risalo" dataset serves as a comprehensive linguistic and literary resource, encompassing 4,767 Sindhi poetic verses drawn from the 30 traditional Surs (sections) of the esteemed magnum opus of Shah Abdul Latif Bhittai. Each… See the full description on the dataset page: https://huggingface.co/datasets/ambile-official/AMBILE_Shah_Jo_Risalo_Labeled.python-optimization-dpo-sampleAdvanced_Dataset_SampleThis is a high-fidelity Direct Preference Optimization (DPO) dataset curated by OptiRefine. It is designed to train Large Language Models (LLMs) to act as helpful, honest, and thoughtful assistants across complex domains.
While our core datasets focus on code refactoring, this dataset provides preference trajectories for broader system architecture, computer science fundamentals, logic, and professional communication.
Curated by: OptiRefine
Language: English
License: Apache-2.0
Format: JSONL… See the full description on the dataset page: https://huggingface.co/datasets/OptiRefine-Official/Advanced_Dataset_Sample.aba-official-curriculum-sft
ABA Official Curriculum SFT
Structured supervision dataset derived from official QABA curriculum sources for:
ABAT
QASP-S
QBA
Files
official_lessons.jsonl
official_qa.jsonl
official_mcq.jsonl
official_curriculum_sft.jsonl
official_curriculum_train.jsonl
official_curriculum_eval.jsonl
manifest.json
Intended use
This dataset is intended for:
instruction tuning on official ABA curriculum content
grounded lesson planning
grounded question answering
grounded… See the full description on the dataset page: https://huggingface.co/datasets/nopoh44/aba-official-curriculum-sft.templeos-official-holyc-dataset
TempleOS Official HolyC Dataset
Summary
This dataset is built only from the official cia-foundation/TempleOS repository.
It contains two subsets:
raw: raw HolyC/code corpus for continued pretraining or domain adaptation
sft_strict: instruction/chat dataset for supervised fine-tuning
Source
Repository: cia-foundation/TempleOS
URL: https://github.com/cia-foundation/TempleOS
Source status: official public mirror/final snapshot
License label used in this… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/templeos-official-holyc-dataset.Moon-1-DataMoon-2-Data
