CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01YANS-official /ogiri-bokete 読み込み方 from datasets import load_dataset dataset = load_dataset("YANS-official/ogiri-bokete", split="train") 概要 大喜利投稿サイトBoketeのクロールデータです。元データは CLoT-Oogiri-Go [Zhang+ CVPR2024]というデータの一部です。 詳細はCVPRのプロジェクトページをご確認ください。 このデータは以下の3タスクが含まれます。 text_to_text: テキストでお題が渡され、それに対する回答を返します。 image_to_text: いわゆる「画像で一言」です。画像のみが渡されて、テキストによる回答を返します。 text_image_to_text: 画像中にテキストが書かれています。テキストの一部が空欄になっているので、そこに穴埋めする形で回答を返します。 それぞれの量は以下の通りです。(8/30現在。ハッカソン当日までに増やす可能性があります。) タスク… See the full description on the dataset page: https://huggingface.co/datasets/YANS-official/ogiri-bokete.imagetext-generationn<1K4 likes7.8k downloads2y agoHugging Face02BGPT-OFFICIAL /refute Can AI read new science honestly? Models can sound convincing while misreading a result or expressing more confidence than the evidence deserves. That matters when people use them to summarize papers, compare studies, or decide what to investigate next. REFUTE tests whether a model knows the finding, spots quiet flaws, names what would overturn a claim, and matches its confidence to the evidence. Truth Score is the main result. It combines factual accuracy, flaw… See the full description on the dataset page: https://huggingface.co/datasets/BGPT-OFFICIAL/refute.imagetext-generationn<1K2 likes1.4k downloads2mo agoHugging Face03agoulah /ontario-hansard-official Ontario Hansard (Official) — speaker-identity-resolved ~4.3 million paragraph-level records from the official Hansard of the Legislative Assembly of Ontario, parliaments 29–44 (1974–2025), with each paragraph's speaker resolved to a stable member identity and tagged with the confidence tier of that resolution. Overview 4,316,254 records, one per Hansard paragraph. Parliaments 29–44, sessions spanning 1974–2025. 51 years of floor proceedings: debates, question… See the full description on the dataset page: https://huggingface.co/datasets/agoulah/ontario-hansard-official.tabulartext-classification1M<n<10M0 likes149 downloads3mo agoHugging Face04YichengWangCA /aime24-official AIME 2024 — official wording, figures retained All 30 problems from the 2024 American Invitational Mathematics Examination (AIME I and AIME II), transcribed from the official exam text with every figure retained as Asymptote source. This exists because the circulating text-only versions of AIME 2024 are not faithful to the official problems, and at least one problem in them cannot be solved as written. Why this dataset exists While evaluating a reasoning model on… See the full description on the dataset page: https://huggingface.co/datasets/YichengWangCA/aime24-official.textquestion-answeringn<1K0 likes74 downloads23d agoHugging Face05Nulll-Official /ctext 说明 ctext - all - slice:自主收集的文言文繁体语料库,主要爬取自维基文库、漢川草廬等处 ctext - 副本 - 副本:ctext - all - slice的先秦部分 ctext - 白话:自主收集的白话及现代汉语繁体语料库,主要爬取自维基百科、BWIKI、维基文库、知乎、繁體中文書庫等处 项目链接 texttext-generation1K<n<10K0 likes53 downloads7mo agoHugging Face06ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k DeepScaleR teacher SFT vLLM official 40k Generated run: exp_003_vllm_official_brainlab_2gpu. Summary { "num_examples": 40300, "sft_dir": "data/processed/deepscaler/teacher_sft/exp_003_vllm_official_brainlab_2gpu", "parse_rate": 0.9999751861042183, "correct_rate": 0.5728039702233251, "format_rate": 0.005955334987593052, "mean_reward": 0.42432258064534184, "deepscaler_mean_reward": 0.6266997518610422, "deepscaler_match_mean_reward":… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.texttext-generation10K<n<100K0 likes50 downloads4mo agoHugging Face07forseasons /dataagentbench-derived-official54-altimate-prompt-shell DataAgentBench-Derived Official54 Altimate Prompt Shell This public dataset contains 54 DataAgentBench-derived task prompt-shell rows compiled for the Altimate-centered DAB sandbox runtime. It is a derived compile/export artifact, not the official raw DataAgentBench release. It is intended as a portable task/prompt/manifest source for downstream Altimate teacher rollouts, SFT construction, or RL data compilation. It is not a completed rollout dataset: the rows do not contain… See the full description on the dataset page: https://huggingface.co/datasets/forseasons/dataagentbench-derived-official54-altimate-prompt-shell.text-generationn<1K0 likes40 downloads3mo agoHugging Face08ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k-clean-v2 DeepScaleR Teacher40k Clean v2 Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k. Filtering minimum official reward: 1.0 maximum text tokens: 8192 maximum response chars: 65000 near-duplicate SimHash hamming threshold: 4 required <think>...</think> and final boxed answer after reasoning exact text/problem/response dedupe and near problem dedupe Counts raw examples: 40300 kept examples: 21727 train examples: 21292 val… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v2.tabulartext-generation10K<n<100K0 likes29 downloads4mo agoHugging Face09ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k. Filtering reward filter enabled: False minimum official reward: 1.0 scoring errors rejected: False maximum text tokens: 8192 maximum response chars: 65000 near-duplicate SimHash hamming threshold: 4 required <think>...</think> and final boxed answer after reasoning exact text/problem/response dedupe and near problem… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter.tabulartext-generation10K<n<100K0 likes25 downloads4mo agoHugging Face10ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k. Filtering reward filter enabled: False minimum official reward: 1.0 scoring errors rejected: False maximum text tokens: 32768 maximum response chars: 200000 near-duplicate SimHash hamming threshold: 4 required <think>...</think> and final boxed answer after reasoning exact text/problem/response dedupe and near problem dedupe… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual.tabulartext-generation10K<n<100K1 likes25 downloads4mo agoHugging Face11AxionLab-official /Reasoning-MiniGPT-brazilian-portuguesetexttext-generationn<1K0 likes17 downloads9mo agoHugging Face12ambile-official /AMBILE_Shah_Jo_Risalo_Labeled AMBILE Shah Jo Risalo Developed by:Abdul Majid Bhurgri Institute of Language Engineering (AMBILE), HyderabadUnder the administrative control of the Culture, Tourism, Antiquities & Archives Department, Government of Sindh Dataset Overview The "Shah Jo Risalo" dataset serves as a comprehensive linguistic and literary resource, encompassing 4,767 Sindhi poetic verses drawn from the 30 traditional Surs (sections) of the esteemed magnum opus of Shah Abdul Latif Bhittai. Each… See the full description on the dataset page: https://huggingface.co/datasets/ambile-official/AMBILE_Shah_Jo_Risalo_Labeled.tabulartext-classification1K<n<10K0 likes12 downloads1y agoHugging Face13OptiRefine-Official /python-optimization-dpo-sampletexttext-generationn<1K1 likes11 downloads6mo agoHugging Face14OptiRefine-Official /Advanced_Dataset_SampleThis is a high-fidelity Direct Preference Optimization (DPO) dataset curated by OptiRefine. It is designed to train Large Language Models (LLMs) to act as helpful, honest, and thoughtful assistants across complex domains. While our core datasets focus on code refactoring, this dataset provides preference trajectories for broader system architecture, computer science fundamentals, logic, and professional communication. Curated by: OptiRefine Language: English License: Apache-2.0 Format: JSONL… See the full description on the dataset page: https://huggingface.co/datasets/OptiRefine-Official/Advanced_Dataset_Sample.texttext-generationn<1K0 likes10 downloads5mo agoHugging Face15nopoh44 /aba-official-curriculum-sft ABA Official Curriculum SFT Structured supervision dataset derived from official QABA curriculum sources for: ABAT QASP-S QBA Files official_lessons.jsonl official_qa.jsonl official_mcq.jsonl official_curriculum_sft.jsonl official_curriculum_train.jsonl official_curriculum_eval.jsonl manifest.json Intended use This dataset is intended for: instruction tuning on official ABA curriculum content grounded lesson planning grounded question answering grounded… See the full description on the dataset page: https://huggingface.co/datasets/nopoh44/aba-official-curriculum-sft.texttext-generationn<1K0 likes6 downloads6mo agoHugging Face16NickIBrody /templeos-official-holyc-dataset TempleOS Official HolyC Dataset Summary This dataset is built only from the official cia-foundation/TempleOS repository. It contains two subsets: raw: raw HolyC/code corpus for continued pretraining or domain adaptation sft_strict: instruction/chat dataset for supervised fine-tuning Source Repository: cia-foundation/TempleOS URL: https://github.com/cia-foundation/TempleOS Source status: official public mirror/final snapshot License label used in this… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/templeos-official-holyc-dataset.texttext-generationn<1K0 likes4 downloads5mo agoHugging Face17NeoAI-Official /Moon-1-Datatexttext-generationn<1K0 likes3 downloads6mo agoHugging Face18NeoAI-Official /Moon-2-Datatexttext-generation1K<n<10K0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.