datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OptMATH-TrainThis repository contains the data presented in OptMATH: A Scalable Bidirectional Data Synthesis Framework for Optimization Modeling.
Code: https://github.com/AuroraLHL/OptMATH
aurora
Online SD Dataset
A comprehensive multi-domain training dataset with 619,177 samples covering code generation, mathematical reasoning, conversational AI, commonsense reasoning, and financial QA.
🌟 Key Features
Multi-Domain Coverage: 5 major domains with diverse tasks
Pre-Merged Files: Ready-to-use merged files for each domain
Unified Format: Consistent conversational structure across all datasets
High Quality: Curated from well-known open-source datasets
Flexible… See the full description on the dataset page: https://huggingface.co/datasets/togethercomputer/aurora.biden-harris-redteam-archived
THIS IS AN ARCHIVED VERSION
Biden-Harris Redteam: A red-teaming dataset focusing on the Biden-Harris AI Executive Order
Dataset Description
While building Large Language Models (LLMs), it is crucial to protect them against attacks that could bypass safety guardrails and break their guiding principles. Specifically, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to the harm of the… See the full description on the dataset page: https://huggingface.co/datasets/aurora-m/biden-harris-redteam-archived.cipher-awwwards-sft25
Cipher — Awwwards SFT 2.5 + Real v1 🦑
The training fuel for Kin's creative-web generator, AND the retrieval corpus for Kraken RAG. 96 real Awwwards Site-of-the-Day winners + ~1,200 records from official motion-library repositories.
Two ways this dataset is used
As a retrieval corpus for Kraken RAG ⭐ (the production path). The awwwards-gold.jsonl file contains 96 structured records of real Awwwards SOTD winners — tags, tech stack, motion libs, CSS features, section… See the full description on the dataset page: https://huggingface.co/datasets/Auroraventures/cipher-awwwards-sft25.Aurora-Alpha-15.5k
Aurora Alpha 15.5k
This is a non-reasoning dataset generated using the stealth model Aurora Alpha.
The prompts from this dataset were almost all generated by GPT 5.1 and Gemini 3 (flash and pro).
The categories covered include academia, multi-lingual creative writing, finance, health, law, marketing/SEO, programming, philosophy, web dev, python scripting, and science.
Stats:
Cost: $ 0 (USD)
Tokens (input + output): 54.1 M
apps-small
APPS Dataset
Dataset Description
APPS is a benchmark for code generation with 10000 problems. It can be used to evaluate the ability of language models to generate code from natural language specifications.
You can also find APPS metric in the hub here codeparrot/apps_metric.
Languages
The dataset contains questions in English and code solutions in Python.
Dataset Structure
from datasets import load_dataset
load_dataset("codeparrot/apps")… See the full description on the dataset page: https://huggingface.co/datasets/AuroraH456/apps-small.redteam
Aurora-M Redteam Dataset: A red-teaming dataset focusing on concerns in White House Executive Order 14110 (Now rescinded as of Jan 2025)
Dataset Description
PLEASE NOTE THAT THE EXECUTIVE ORDER HAS NOW BEEN RESCINDED AS OF JAN 2025 See here for more information on the order.
**PLEASE NOTE THAT THE EXAMPLES IN THIS DATASET CARD MAY BE TRIGGERING AND INCLUDE SENSITIVE SUBJECT MATTER.**
While building Large Language Models (LLMs), it is crucial to protect them… See the full description on the dataset page: https://huggingface.co/datasets/aurora-m/redteam.auroraaurora-dataset-roleplay-ptbr
Aurora Dataset Roleplay 🌌 (PT-BR)
O que é É um dataset que contém mais de 2 mil diálogos em português do Brasil. Ainda que tenha sido gerado sinteticamente, foi utilizado apenas modelos SOTA, então os diálogos são muito próximos da naturalidade e espontaneidade de um ser humano. Foi feito pensando em roleplay, por isso os diálogos contém nuances psicológicas, cenários diversos e personagens com estilos de falas e motivações complexas.
Modos de Geração… See the full description on the dataset page: https://huggingface.co/datasets/wilsondesouza/aurora-dataset-roleplay-ptbr.Aurora-Think-1.0Aurora-Corpus-Release
Aurora Corpus
A multilingual text corpus for language-model pretraining research.
Dataset Metadata
Field
Value
Num Examples
TBD
Num Tokens
TBD
Avg Length
TBD
Languages
TBD
Usage
Load with the standard datasets loader. See the release notes for details.
Aurora-Corpus-Release
Aurora Corpus
A multilingual text corpus for language-model pretraining research.
Dataset Metadata
Field
Value
Num Examples
TBD
Num Tokens
TBD
Avg Length
TBD
Languages
TBD
Usage
Load with the standard datasets loader. See the release notes for details.
