CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aimosprite /prompt-swap-mixed12-5xlr-e1-mxfp4-mergedtabularn<1K0 likes530 downloads6mo agoHugging Face02aimosprite /prompt-swap-mixed12-5xlr-e2-mxfp4-mergedtabularn<1K0 likes515 downloads6mo agoHugging Face03aimosprite /prompt-swap-medium12-e2-mxfp4-mergedtabularn<1K0 likes343 downloads6mo agoHugging Face04gemmozero /ai-models-2026 AI Models & Releases 2026 AI model releases, benchmarks, capabilities. Updated daily via automated collection pipeline. Part of the Legion Data Factory — historical AI ecosystem datasets 2026. Methodology Automated collection from public sources (HackerNews, RSS feeds, APIs). Updated daily via cron job. Raw data, minimal processing. License CC BY 4.0 🔑 API Access — Updated Daily Live data via Legion AI API | Documentation Free: 100… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-models-2026.text1K<n<10K0 likes191 downloads1h agoHugging Face05snsm13 /aimo3-train-datatext1K<n<10K0 likes109 downloads6mo agoHugging Face06aphoticshaman /aimo3-math-dataset AIMO3 Math Dataset Training data for AI Mathematical Olympiad Progress Prize 3. Files train_cot.jsonl - Chain-of-Thought examples train_tir.jsonl - Tool-Integrated Reasoning examples Author Ryan J Cardwell (Archer Phoenix) - AIMO3 Competitor texttext-generationn<1K0 likes101 downloads10mo agoHugging Face07batow133 /aimodel-sft-v1text1K<n<10K0 likes101 downloads9d agoHugging Face08AI-Mock-Interviewer /Train_datatextquestion-answering1K<n<10K0 likes73 downloads1y agoHugging Face09ChuGyouk /AI-MO-NuminaMath-TIR-korean-240918 IMPORTANT NOTE This data is part of the progress. Current translation progress: 24.85% (2024-09-18 01:32 KST) I'm taking a short break due to personal reasons. I'll be back in a month. TODO-LIST Finish translation Translation I used gemini-1.5-pro-exp-0827. The prompt used for translation will be disclosed at the end. Dataset Card for NuminaMath CoT Dataset Summary Tool-integrated reasoning (TIR) plays a crucial role in this… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/AI-MO-NuminaMath-TIR-korean-240918.texttext-generation10K<n<100K5 likes72 downloads2y agoHugging Face10AI-MO /GeometryLeanBenchtextn<1K1 likes55 downloads1y agoHugging Face11aimosprite /prompt-swap-medium12-e1-mxfp4-mergedtabularn<1K0 likes53 downloads6mo agoHugging Face12referencesource /ai-model-deprecation-and-retirement AI model deprecation and retirement dates by provider Canonical, always-current version: https://referencesource.org/ai-model-deprecation-and-retirement/ Machine-readable: https://referencesource.org/ai-model-deprecation-and-retirement/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-12 Stale after: 2026-09-11 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 347 Which AI API models are deprecated… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/ai-model-deprecation-and-retirement.textn<1K0 likes34 downloads1mo agoHugging Face13open-llm-leaderboard /AI-MO__NuminaMath-7B-TIR-detailsgated Dataset Card for Evaluation run of AI-MO/NuminaMath-7B-TIR Dataset automatically created during the evaluation run of model AI-MO/NuminaMath-7B-TIR The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AI-MO__NuminaMath-7B-TIR-details.tabular10K<n<100K0 likes24 downloads2y agoHugging Face14open-llm-leaderboard /AI-MO__NuminaMath-7B-CoT-detailsgated Dataset Card for Evaluation run of AI-MO/NuminaMath-7B-CoT Dataset automatically created during the evaluation run of model AI-MO/NuminaMath-7B-CoT The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AI-MO__NuminaMath-7B-CoT-details.tabular10K<n<100K0 likes23 downloads2y agoHugging Face15alexlyzhov-aimon /dataframer_text_to_sqltext1K<n<10K0 likes23 downloads9mo agoHugging Face16aimosprite /training-data-oss120btabular1K<n<10K0 likes17 downloads6mo agoHugging Face17UR-xiaoyang /AIMO3_to_Hardtextn<1K0 likes15 downloads9mo agoHugging Face18aimosprite /brian-rollouts-311-training-turns brian-rollouts-311-training-turns-v1 Lean per-turn GPT-OSS training export derived from brian-rollouts-311-full-v2. Contents: rollouts.training_turns.jsonl: one row per assistant call with full prompt token IDs and completion token IDs manifest.json: export metadata and row counts Semantics: prompt_token_ids are the full context shown to the model for that assistant call. completion_token_ids are the tokens generated by the model on that call. prompt_token_ids include system… See the full description on the dataset page: https://huggingface.co/datasets/aimosprite/brian-rollouts-311-training-turns.tabular10K<n<100K0 likes11 downloads6mo agoHugging Face19firatmihci /ai-model-accent-corpus AI Model Accent Corpus A paired corpus for studying the prose "accent" of large language models: 5 models × the same 102 open-ended prompts = 510 plain-prose passages, generated June 2026 at a fixed decoding temperature, constrained to plain paragraphs (no lists or headings) so the data isolates prose style rather than formatting choices. Models: OpenAI GPT-4o, GPT-4o-mini, GPT-3.5-turbo; Anthropic Claude Sonnet 4.5, Claude Haiku 4.5. Released with the study "Every Model Has an… See the full description on the dataset page: https://huggingface.co/datasets/firatmihci/ai-model-accent-corpus.textn<1K0 likes11 downloads3mo agoHugging Face20maikahj /kaggle-aimo2text10K<n<100K0 likes10 downloads1y agoHugging Face21yiboowang /aimo-validation-amc-repeated3textn<1K0 likes10 downloads1y agoHugging Face22aimosprite /moh_8_fake_rollouts MOH-8 Fake Rollouts 480 math competition problems, each with 8 candidate solution rollouts from OSS 120B. A controlled number of rollouts per problem are correct — use this to train/test a verifier model that must identify which solutions are right. Source Problems and rollouts sampled from aimosprite/training-data-oss120b (the oss128-fixed-FINAL.jsonl file). Only polymath-source problems in the 2/8–4/8 pass rate range (32–64 correct out of 128 attempts). 4 problems… See the full description on the dataset page: https://huggingface.co/datasets/aimosprite/moh_8_fake_rollouts.tabulartext-generationn<1K0 likes9 downloads6mo agoHugging Face23AI-Mock-Interviewer /Test_Datatextn<1K0 likes8 downloads1y agoHugging Face24PraMamba /AIMO-2_ReferenceThis CSV file is reference.csv in Kaggle's AI Mathematical Olympiad - Progress Prize 2. textquestion-answeringn<1K0 likes7 downloads2y agoHugging Face25aimosprite /training-settabular1K<n<10K0 likes6 downloads6mo agoHugging Face26aimosprite /test-largedate: mar 02 textn<1K0 likes5 downloads6mo agoHugging Face27aimosprite /gpt-oss-120b-marina-cleaned-tok-idtext10K<n<100K0 likes5 downloads6mo agoHugging Face28aimosprite /chinese-translatedtabularn<1K0 likes3 downloads6mo agoHugging Face29aimosprite /brian-benchtabularn<1K0 likes2 downloads7mo agoHugging Face30dvyomkesh /aimo-phase1-teacher-corpus AIMO Phase 1 Teacher Corpus Private export of the current Phase 1 teacher-style corpus used for overnight data preparation. source file: runtime-teacher-phase1.jsonl text1K<n<10K0 likes2 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.