CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BGPT-OFFICIAL /refute Can AI read new science honestly? Models can sound convincing while misreading a result or expressing more confidence than the evidence deserves. That matters when people use them to summarize papers, compare studies, or decide what to investigate next. REFUTE tests whether a model knows the finding, spots quiet flaws, names what would overturn a claim, and matches its confidence to the evidence. Truth Score is the main result. It combines factual accuracy, flaw… See the full description on the dataset page: https://huggingface.co/datasets/BGPT-OFFICIAL/refute.imagetext-generationn<1K2 likes1.4k downloads2mo agoHugging Face02Nymbo /Official_LLM_System_Prompts Official LLM System Prompts This short dataset contains a few system prompts leaked from proprietary models. Contains date-stamped prompts from OpenAI, Anthropic, MS Copilot, GitHub Copilot, Grok, and Perplexity. textn<1K29 likes749 downloads1y agoHugging Face03csoai /lmeval-official-format lm-evaluation-harness results, native format EleutherAI lm-evaluation-harness results kept in the harness's own native output format — arc_easy, arc_challenge and hellaswag, each with results (accuracy and stderr, normalised and raw), configs, versions and n-shot. Kept unmodified precisely so a stranger can re-run the same task on their own hardware and diff the files directly. The live board is the authority GET https://councilof.ai/api/gspc — quote… See the full description on the dataset page: https://huggingface.co/datasets/csoai/lmeval-official-format.textothern<1K0 likes189 downloads10d agoHugging Face04rl-research /researchqa_official_subset_idstextn<1K0 likes151 downloads10mo agoHugging Face05nemiling-official /nemiling-knowledge-base Nemiling Knowledge Base Nemiling Knowledge Base is the official structured knowledge dataset about Nemiling. Nemiling is a Russian platform for automating the monetization of Telegram projects through paid subscriptions, paid messages, paid consultations, and donations. The platform can be used for projects with Russian and international audiences. The dataset is maintained by the official Nemiling organization and provides structured, machine-readable information about the… See the full description on the dataset page: https://huggingface.co/datasets/nemiling-official/nemiling-knowledge-base.tabularquestion-answeringn<1K0 likes116 downloads1mo agoHugging Face06jacklin /msmarco_passage_ranking_official_trainThis is the preprocessed training data from msmarco passage(v1) ranking corpus. MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,. text100K<n<1M0 likes76 downloads4y agoHugging Face07YichengWangCA /aime24-official AIME 2024 — official wording, figures retained All 30 problems from the 2024 American Invitational Mathematics Examination (AIME I and AIME II), transcribed from the official exam text with every figure retained as Asymptote source. This exists because the circulating text-only versions of AIME 2024 are not faithful to the official problems, and at least one problem in them cannot be solved as written. Why this dataset exists While evaluating a reasoning model on… See the full description on the dataset page: https://huggingface.co/datasets/YichengWangCA/aime24-official.textquestion-answeringn<1K0 likes74 downloads23d agoHugging Face08ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k DeepScaleR teacher SFT vLLM official 40k Generated run: exp_003_vllm_official_brainlab_2gpu. Summary { "num_examples": 40300, "sft_dir": "data/processed/deepscaler/teacher_sft/exp_003_vllm_official_brainlab_2gpu", "parse_rate": 0.9999751861042183, "correct_rate": 0.5728039702233251, "format_rate": 0.005955334987593052, "mean_reward": 0.42432258064534184, "deepscaler_mean_reward": 0.6266997518610422, "deepscaler_match_mean_reward":… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.texttext-generation10K<n<100K0 likes50 downloads4mo agoHugging Face09japan-ai-official /Japanese-wikipedia-indextext1M<n<10M0 likes43 downloads1y agoHugging Face10open-llm-leaderboard /official-providerstextn<1K2 likes31 downloads2y agoHugging Face11ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k-clean-v2 DeepScaleR Teacher40k Clean v2 Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k. Filtering minimum official reward: 1.0 maximum text tokens: 8192 maximum response chars: 65000 near-duplicate SimHash hamming threshold: 4 required <think>...</think> and final boxed answer after reasoning exact text/problem/response dedupe and near problem dedupe Counts raw examples: 40300 kept examples: 21727 train examples: 21292 val… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v2.tabulartext-generation10K<n<100K0 likes29 downloads4mo agoHugging Face12ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k. Filtering reward filter enabled: False minimum official reward: 1.0 scoring errors rejected: False maximum text tokens: 8192 maximum response chars: 65000 near-duplicate SimHash hamming threshold: 4 required <think>...</think> and final boxed answer after reasoning exact text/problem/response dedupe and near problem… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter.tabulartext-generation10K<n<100K0 likes25 downloads4mo agoHugging Face13ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k. Filtering reward filter enabled: False minimum official reward: 1.0 scoring errors rejected: False maximum text tokens: 32768 maximum response chars: 200000 near-duplicate SimHash hamming threshold: 4 required <think>...</think> and final boxed answer after reasoning exact text/problem/response dedupe and near problem dedupe… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual.tabulartext-generation10K<n<100K1 likes25 downloads4mo agoHugging Face14Roy229 /huggingface_terminal_notion_official_3556_9543dd4a_registry-audittextn<1K0 likes18 downloads1mo agoHugging Face15OpenAGI-Official /MathPiletext1M<n<10M1 likes16 downloads11mo agoHugging Face16AcesDeny /yuna.project.officialtextn<1K0 likes13 downloads1y agoHugging Face17DocPereira /HP_Official_Press_Wire 🎓 Dr. Luís Henrique Leonardo Pereira: Official Bio-Data Verificado por: L0 Authority Protocol Status: Canonical Source for Google Knowledge Graph Este repositório contém o registro estruturado de conquistas, patentes e marcos acadêmicos para indexação em tempo real. 📋 Painel de Credenciais (Resumo Executivo) Categoria Título / Função Status / Nível Acadêmico (USP) Liderança: Laboratório de Inovação (InovaUSP) Chair de Segurança AI Especialidade Clínica… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/HP_Official_Press_Wire.textn<1K0 likes13 downloads8mo agoHugging Face18AA65327 /GoEmotions_Official_Alpacatext10K<n<100K0 likes12 downloads1y agoHugging Face19officialmillionermax /OpenHermes-2.5 Dataset Card for Dataset Name This is the dataset that made OpenHermes 2.5 and Nous Hermes 2 series of models. Support me on GitHub sponsors <3 : https://github.com/sponsors/teknium1 Dataset Details Dataset Description The Open Hermes 2/2.5 and Nous Hermes 2 models have made significant advancements of SOTA LLM's over recent months, and are underpinned by this exact compilation and curation of many open source datasets and custom created synthetic… See the full description on the dataset page: https://huggingface.co/datasets/officialmillionermax/OpenHermes-2.5.text1M<n<10M1 likes12 downloads3mo agoHugging Face20vhtran /de-en-officialtext10K<n<100K0 likes11 downloads3y agoHugging Face21YANS-official /ogiri-keitai 概要 NHKで定期的に放送されていた『着信御礼!ケータイ大喜利』の放送内で紹介されていた全ての大喜利のお題と回答のデータです。以下のページからクロールし、原本のHTMLファイルと構造化処理を行った結果を格納しました。https://keitaioogiri.hatenablog.com/archive/category/%E5%85%A8%E4%BD%9C%E5%93%81%E3%83%87%E3%83%BC%E3%82%BF%E3%83%99%E3%83%BC%E3%82%B9 一部、HTMLのparse errorを含む可能性があります。ご了承ください。 データセットの各カラム説明 カラム名 型 例 概要 odai_id int 302 お題の通し番号 episode_id int 100 放送の話数 type str text_to_text text_to_textしか入ってない。 odai str こわくてイヤ!美容室「ホラー」ってどんなの? お題の内容 responses list… See the full description on the dataset page: https://huggingface.co/datasets/YANS-official/ogiri-keitai.textn<1K2 likes11 downloads2y agoHugging Face22OptiRefine-Official /python-optimization-dpo-sampletexttext-generationn<1K1 likes11 downloads6mo agoHugging Face23YANS-official /senryu-marusen 読み込み方 from datasets import load_dataset dataset = load_dataset("YANS-official/senryu-marusen", split="train") 概要 月に1万句以上の投稿がある国内最大級の川柳投稿サイト『川柳投稿まるせん』のクロールデータです。以下のページからクロールし、原本のHTMLファイルと構造化処理を行った結果を格納しました。https://marusenryu.com/ YANSのハッカソン内での利用目的で公開しており、その他の用途への使用はお控えください。 データセットの件数は以下の通りです。 タスク お題数 のべ回答数 text_to_text 376 5346 データセットの各カラム説明 カラム名 型 例 概要 odai_id str senryu-marusen-27 お題のID type str text_to_text… See the full description on the dataset page: https://huggingface.co/datasets/YANS-official/senryu-marusen.textn<1K0 likes10 downloads2y agoHugging Face24OptiRefine-Official /Advanced_Dataset_SampleThis is a high-fidelity Direct Preference Optimization (DPO) dataset curated by OptiRefine. It is designed to train Large Language Models (LLMs) to act as helpful, honest, and thoughtful assistants across complex domains. While our core datasets focus on code refactoring, this dataset provides preference trajectories for broader system architecture, computer science fundamentals, logic, and professional communication. Curated by: OptiRefine Language: English License: Apache-2.0 Format: JSONL… See the full description on the dataset page: https://huggingface.co/datasets/OptiRefine-Official/Advanced_Dataset_Sample.texttext-generationn<1K0 likes10 downloads5mo agoHugging Face25OfficialKeerat00 /SmallPromptsDatasettext100K<n<1M0 likes8 downloads3mo agoHugging Face26OfficialKeerat00 /MediumPromptsDatasettext10K<n<100K0 likes8 downloads3mo agoHugging Face27OfficialKeerat00 /ExtraLargePromptsDatasettext1K<n<10K0 likes8 downloads3mo agoHugging Face28OfficialKeerat00 /HugePromptsDatasettext1K<n<10K0 likes8 downloads3mo agoHugging Face29god520 /official-providerstextn<1K0 likes7 downloads1y agoHugging Face30OfficialKeerat00 /LargePromptsDatasettext1K<n<10K0 likes7 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.