CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Aurora-Gem /OptMATH-TrainThis repository contains the data presented in OptMATH: A Scalable Bidirectional Data Synthesis Framework for Optimization Modeling. Code: https://github.com/AuroraLHL/OptMATH texttext-generation100K<n<1M7 likes321 downloads2y agoHugging Face02togethercomputer /aurora Online SD Dataset A comprehensive multi-domain training dataset with 619,177 samples covering code generation, mathematical reasoning, conversational AI, commonsense reasoning, and financial QA. 🌟 Key Features Multi-Domain Coverage: 5 major domains with diverse tasks Pre-Merged Files: Ready-to-use merged files for each domain Unified Format: Consistent conversational structure across all datasets High Quality: Curated from well-known open-source datasets Flexible… See the full description on the dataset page: https://huggingface.co/datasets/togethercomputer/aurora.texttext-generation1M<n<10M4 likes140 downloads6mo agoHugging Face03aurora-m /biden-harris-redteam-archived THIS IS AN ARCHIVED VERSION Biden-Harris Redteam: A red-teaming dataset focusing on the Biden-Harris AI Executive Order Dataset Description While building Large Language Models (LLMs), it is crucial to protect them against attacks that could bypass safety guardrails and break their guiding principles. Specifically, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to the harm of the… See the full description on the dataset page: https://huggingface.co/datasets/aurora-m/biden-harris-redteam-archived.texttext-generation10K<n<100K7 likes101 downloads1y agoHugging Face04aurora-m /multilingual_chatgatedThis is a quick dataset card which we will update. This dataset is a multilingual chat dataset made from translations of a subset of the biden-harris_redteam dataset, chatbot arena conversations dataset and ultrachat_200k. To the extent we have any copyrights under this data, we license it under cc-by-nc-4.0. See the underlying datasets themselves for their licenses. text10K<n<100K0 likes94 downloads2y agoHugging Face05Auroraventures /cipher-awwwards-sft25 Cipher — Awwwards SFT 2.5 + Real v1 🦑 The training fuel for Kin's creative-web generator, AND the retrieval corpus for Kraken RAG. 96 real Awwwards Site-of-the-Day winners + ~1,200 records from official motion-library repositories. Two ways this dataset is used As a retrieval corpus for Kraken RAG ⭐ (the production path). The awwwards-gold.jsonl file contains 96 structured records of real Awwwards SOTD winners — tags, tech stack, motion libs, CSS features, section… See the full description on the dataset page: https://huggingface.co/datasets/Auroraventures/cipher-awwwards-sft25.tabulartext-generation1K<n<10K0 likes85 downloads5mo agoHugging Face06wchai /AuroraCap-recaption AuroraCap-recaption Resources Website arXiv: Paper GitHub: Code Huggingface: AuroraCap Model Huggingface: VDC Benchmark Huggingface: Trainset Features Video recaption data by AuroraCap. Continue updating... For some video source, we could upload the raw videos but for the others we could only provide the url since the well-known reason. Citation @article{chai2024auroracap, title={AuroraCap: Efficient, Performant Video Detailed… See the full description on the dataset page: https://huggingface.co/datasets/wchai/AuroraCap-recaption.textvisual-question-answering10K<n<100K5 likes61 downloads2y agoHugging Face07aurora-m /aurora-m-dataset-part-1gatedThis is part of the continued pretraining dataset used to train the Aurora-M model described in [Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code, COLING 2025] (https://aclanthology.org/2025.coling-industry.56/). text100M<n<1B0 likes57 downloads1y agoHugging Face08aurora-m /adversarial-promptsgatedAdding various adversrial permuations to questions in the aurora-redteam dataset. text10K<n<100K3 likes56 downloads1y agoHugging Face09TeichAI /Aurora-Alpha-15.5k Aurora Alpha 15.5k This is a non-reasoning dataset generated using the stealth model Aurora Alpha. The prompts from this dataset were almost all generated by GPT 5.1 and Gemini 3 (flash and pro). The categories covered include academia, multi-lingual creative writing, finance, health, law, marketing/SEO, programming, philosophy, web dev, python scripting, and science. Stats: Cost: $ 0 (USD) Tokens (input + output): 54.1 M texttext-generation10K<n<100K8 likes48 downloads7mo agoHugging Face10henriporto /auroratextn<1K0 likes36 downloads7d agoHugging Face11AuroraH456 /apps-small APPS Dataset Dataset Description APPS is a benchmark for code generation with 10000 problems. It can be used to evaluate the ability of language models to generate code from natural language specifications. You can also find APPS metric in the hub here codeparrot/apps_metric. Languages The dataset contains questions in English and code solutions in Python. Dataset Structure from datasets import load_dataset load_dataset("codeparrot/apps")… See the full description on the dataset page: https://huggingface.co/datasets/AuroraH456/apps-small.texttext-generationn<1K0 likes29 downloads2y agoHugging Face12cohort-rlwm /Liquid_V1_7B-pico-aurora-vidgen-multiturn-annotationstext100K<n<1M0 likes26 downloads9mo agoHugging Face13Roy229 /pdf-tools_filesystem_github_terminal_huggingface_2118_x94fu9_manifest_aurora Aurora Flow - Release Manifest Candidate code name: aurora Version: 2.1.0 Status: stable License: MIT Language: Python Description: Real-time streaming data processing library Maintainer: NovaTech Data Engineering Last release: 2026-07-30 This dataset contains manifest.json (the release manifest) and validate.py (the standardized validation script). Run python validate.py from this directory to validate the release manifest. textn<1K0 likes22 downloads1mo agoHugging Face14wilsondesouza /aurora-dataset-roleplay-ptbr Aurora Dataset Roleplay 🌌 (PT-BR) O que é É um dataset que contém mais de 2 mil diálogos em português do Brasil. Ainda que tenha sido gerado sinteticamente, foi utilizado apenas modelos SOTA, então os diálogos são muito próximos da naturalidade e espontaneidade de um ser humano. Foi feito pensando em roleplay, por isso os diálogos contém nuances psicológicas, cenários diversos e personagens com estilos de falas e motivações complexas. Modos de Geração… See the full description on the dataset page: https://huggingface.co/datasets/wilsondesouza/aurora-dataset-roleplay-ptbr.texttext-generation1K<n<10K0 likes21 downloads10mo agoHugging Face15naimulislam /aurora_programmer_data My Awesome Dataset A comprehensive description of my awesome dataset. Dataset Description This dataset contains images of cats and dogs. The images were collected from [mention data source(s), e.g., a specific website, scraped from the internet]. It is intended for use in image classification tasks. The dataset consists of [number] images, with approximately [percentage]% allocated to the training set and [percentage]% to the test set. [Add more details about the… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/aurora_programmer_data.texttext-classificationn<1K0 likes18 downloads2y agoHugging Face16naimulislam /aurora-gpt4texttext-classification10K<n<100K0 likes16 downloads2y agoHugging Face17Roy229 /auroraai-model-registry-5315tabularn<1K0 likes15 downloads1mo agoHugging Face18open-llm-leaderboard /DreadPoor__Aurora_faustus-8B-LINEAR-detailsgated Dataset Card for Evaluation run of DreadPoor/Aurora_faustus-8B-LINEAR Dataset automatically created during the evaluation run of model DreadPoor/Aurora_faustus-8B-LINEAR The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Aurora_faustus-8B-LINEAR-details.tabular10K<n<100K0 likes13 downloads2y agoHugging Face19open-llm-leaderboard /DreadPoor__Aurora_faustus-8B-LORABLATED-detailsgated Dataset Card for Evaluation run of DreadPoor/Aurora_faustus-8B-LORABLATED Dataset automatically created during the evaluation run of model DreadPoor/Aurora_faustus-8B-LORABLATED The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Aurora_faustus-8B-LORABLATED-details.tabular10K<n<100K0 likes13 downloads2y agoHugging Face20aurora-m /aurora-m-dataset-part-2gatedThis is part of the continued pretraining dataset used to train the Aurora-M model described in [Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code, COLING 2025] (https://aclanthology.org/2025.coling-industry.56/). text100M<n<1B0 likes12 downloads1y agoHugging Face21open-llm-leaderboard /DreadPoor__Aurora_faustus-8B-LORABLATED_ALT-detailsgated Dataset Card for Evaluation run of DreadPoor/Aurora_faustus-8B-LORABLATED_ALT Dataset automatically created during the evaluation run of model DreadPoor/Aurora_faustus-8B-LORABLATED_ALT The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Aurora_faustus-8B-LORABLATED_ALT-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face22Roy229 /huggingface_4062_aurora_audio_w6rn6vr1textn<1K0 likes12 downloads1mo agoHugging Face23AlecKloss /auroratextn<1K0 likes10 downloads10mo agoHugging Face24aurora-m /aurora-m2text100K<n<1M0 likes8 downloads11mo agoHugging Face25chenuneris /aurora-mix-data-baize-formattext100K<n<1M1 likes7 downloads3y agoHugging Face26open-llm-leaderboard /jaspionjader__Kosmos-Aurora_faustus-8B-detailsgated Dataset Card for Evaluation run of jaspionjader/Kosmos-Aurora_faustus-8B Dataset automatically created during the evaluation run of model jaspionjader/Kosmos-Aurora_faustus-8B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaspionjader__Kosmos-Aurora_faustus-8B-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face27AuroraH456 /apps-small-editedtextn<1K0 likes7 downloads1y agoHugging Face28b4knoix /aurora-worldtextn<1K0 likes7 downloads4mo agoHugging Face29open-llm-leaderboard /jaspionjader__Kosmos-Elusive-VENN-Aurora_faustus-8B-detailsgated Dataset Card for Evaluation run of jaspionjader/Kosmos-Elusive-VENN-Aurora_faustus-8B Dataset automatically created during the evaluation run of model jaspionjader/Kosmos-Elusive-VENN-Aurora_faustus-8B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaspionjader__Kosmos-Elusive-VENN-Aurora_faustus-8B-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face30chenuneris /lora-aurora-v2text100K<n<1M0 likes3 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.