CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Aurora-Gem /OptMATH-TrainThis repository contains the data presented in OptMATH: A Scalable Bidirectional Data Synthesis Framework for Optimization Modeling. Code: https://github.com/AuroraLHL/OptMATH texttext-generation100K<n<1M7 likes329 downloads2y agoHugging Face02togethercomputer /aurora Online SD Dataset A comprehensive multi-domain training dataset with 619,177 samples covering code generation, mathematical reasoning, conversational AI, commonsense reasoning, and financial QA. 🌟 Key Features Multi-Domain Coverage: 5 major domains with diverse tasks Pre-Merged Files: Ready-to-use merged files for each domain Unified Format: Consistent conversational structure across all datasets High Quality: Curated from well-known open-source datasets Flexible… See the full description on the dataset page: https://huggingface.co/datasets/togethercomputer/aurora.texttext-generation1M<n<10M4 likes141 downloads6mo agoHugging Face03aurora-m /biden-harris-redteam-archived THIS IS AN ARCHIVED VERSION Biden-Harris Redteam: A red-teaming dataset focusing on the Biden-Harris AI Executive Order Dataset Description While building Large Language Models (LLMs), it is crucial to protect them against attacks that could bypass safety guardrails and break their guiding principles. Specifically, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to the harm of the… See the full description on the dataset page: https://huggingface.co/datasets/aurora-m/biden-harris-redteam-archived.texttext-generation10K<n<100K7 likes100 downloads1y agoHugging Face04Auroraventures /cipher-awwwards-sft25 Cipher — Awwwards SFT 2.5 + Real v1 🦑 The training fuel for Kin's creative-web generator, AND the retrieval corpus for Kraken RAG. 96 real Awwwards Site-of-the-Day winners + ~1,200 records from official motion-library repositories. Two ways this dataset is used As a retrieval corpus for Kraken RAG ⭐ (the production path). The awwwards-gold.jsonl file contains 96 structured records of real Awwwards SOTD winners — tags, tech stack, motion libs, CSS features, section… See the full description on the dataset page: https://huggingface.co/datasets/Auroraventures/cipher-awwwards-sft25.tabulartext-generation1K<n<10K0 likes86 downloads5mo agoHugging Face05TeichAI /Aurora-Alpha-15.5k Aurora Alpha 15.5k This is a non-reasoning dataset generated using the stealth model Aurora Alpha. The prompts from this dataset were almost all generated by GPT 5.1 and Gemini 3 (flash and pro). The categories covered include academia, multi-lingual creative writing, finance, health, law, marketing/SEO, programming, philosophy, web dev, python scripting, and science. Stats: Cost: $ 0 (USD) Tokens (input + output): 54.1 M texttext-generation10K<n<100K8 likes43 downloads7mo agoHugging Face06AuroraH456 /apps-small APPS Dataset Dataset Description APPS is a benchmark for code generation with 10000 problems. It can be used to evaluate the ability of language models to generate code from natural language specifications. You can also find APPS metric in the hub here codeparrot/apps_metric. Languages The dataset contains questions in English and code solutions in Python. Dataset Structure from datasets import load_dataset load_dataset("codeparrot/apps")… See the full description on the dataset page: https://huggingface.co/datasets/AuroraH456/apps-small.texttext-generationn<1K0 likes29 downloads2y agoHugging Face07aurora-m /redteamgated Aurora-M Redteam Dataset: A red-teaming dataset focusing on concerns in White House Executive Order 14110 (Now rescinded as of Jan 2025) Dataset Description PLEASE NOTE THAT THE EXECUTIVE ORDER HAS NOW BEEN RESCINDED AS OF JAN 2025 See here for more information on the order. **PLEASE NOTE THAT THE EXAMPLES IN THIS DATASET CARD MAY BE TRIGGERING AND INCLUDE SENSITIVE SUBJECT MATTER.** While building Large Language Models (LLMs), it is crucial to protect them… See the full description on the dataset page: https://huggingface.co/datasets/aurora-m/redteam.texttext-generation1K<n<10K8 likes28 downloads3mo agoHugging Face08arafatar /auroratext-generation10K<n<100K0 likes26 downloads2y agoHugging Face09wilsondesouza /aurora-dataset-roleplay-ptbr Aurora Dataset Roleplay 🌌 (PT-BR) O que é É um dataset que contém mais de 2 mil diálogos em português do Brasil. Ainda que tenha sido gerado sinteticamente, foi utilizado apenas modelos SOTA, então os diálogos são muito próximos da naturalidade e espontaneidade de um ser humano. Foi feito pensando em roleplay, por isso os diálogos contém nuances psicológicas, cenários diversos e personagens com estilos de falas e motivações complexas. Modos de Geração… See the full description on the dataset page: https://huggingface.co/datasets/wilsondesouza/aurora-dataset-roleplay-ptbr.texttext-generation1K<n<10K0 likes21 downloads10mo agoHugging Face10naimulislam /Aurora-Think-1.0texttable-question-answeringn<1K0 likes19 downloads2y agoHugging Face11procedure2012 /Aurora-Corpus-Release Aurora Corpus A multilingual text corpus for language-model pretraining research. Dataset Metadata Field Value Num Examples TBD Num Tokens TBD Avg Length TBD Languages TBD Usage Load with the standard datasets loader. See the release notes for details. text-generation0 likes8 downloads2mo agoHugging Face12yding03 /Aurora-Corpus-Release Aurora Corpus A multilingual text corpus for language-model pretraining research. Dataset Metadata Field Value Num Examples TBD Num Tokens TBD Avg Length TBD Languages TBD Usage Load with the standard datasets loader. See the release notes for details. text-generation0 likes4 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.