datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OptMATH-TrainThis repository contains the data presented in OptMATH: A Scalable Bidirectional Data Synthesis Framework for Optimization Modeling.
Code: https://github.com/AuroraLHL/OptMATH
aurora
Online SD Dataset
A comprehensive multi-domain training dataset with 619,177 samples covering code generation, mathematical reasoning, conversational AI, commonsense reasoning, and financial QA.
🌟 Key Features
Multi-Domain Coverage: 5 major domains with diverse tasks
Pre-Merged Files: Ready-to-use merged files for each domain
Unified Format: Consistent conversational structure across all datasets
High Quality: Curated from well-known open-source datasets
Flexible… See the full description on the dataset page: https://huggingface.co/datasets/togethercomputer/aurora.biden-harris-redteam-archived
THIS IS AN ARCHIVED VERSION
Biden-Harris Redteam: A red-teaming dataset focusing on the Biden-Harris AI Executive Order
Dataset Description
While building Large Language Models (LLMs), it is crucial to protect them against attacks that could bypass safety guardrails and break their guiding principles. Specifically, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to the harm of the… See the full description on the dataset page: https://huggingface.co/datasets/aurora-m/biden-harris-redteam-archived.multilingual_chatThis is a quick dataset card which we will update. This dataset is a multilingual chat dataset made from translations of a subset of the biden-harris_redteam dataset, chatbot arena conversations dataset and ultrachat_200k.
To the extent we have any copyrights under this data, we license it under cc-by-nc-4.0. See the underlying datasets themselves for their licenses.
cipher-awwwards-sft25
Cipher — Awwwards SFT 2.5 + Real v1 🦑
The training fuel for Kin's creative-web generator, AND the retrieval corpus for Kraken RAG. 96 real Awwwards Site-of-the-Day winners + ~1,200 records from official motion-library repositories.
Two ways this dataset is used
As a retrieval corpus for Kraken RAG ⭐ (the production path). The awwwards-gold.jsonl file contains 96 structured records of real Awwwards SOTD winners — tags, tech stack, motion libs, CSS features, section… See the full description on the dataset page: https://huggingface.co/datasets/Auroraventures/cipher-awwwards-sft25.AuroraCap-recaption
AuroraCap-recaption
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: AuroraCap Model
Huggingface: VDC Benchmark
Huggingface: Trainset
Features
Video recaption data by AuroraCap. Continue updating...
For some video source, we could upload the raw videos but for the others we could only provide the url since the well-known reason.
Citation
@article{chai2024auroracap,
title={AuroraCap: Efficient, Performant Video Detailed… See the full description on the dataset page: https://huggingface.co/datasets/wchai/AuroraCap-recaption.aurora-m-dataset-part-1This is part of the continued pretraining dataset used to train the Aurora-M model described in [Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code, COLING 2025] (https://aclanthology.org/2025.coling-industry.56/).
adversarial-promptsAdding various adversrial permuations to questions in the aurora-redteam dataset.
Aurora-Alpha-15.5k
Aurora Alpha 15.5k
This is a non-reasoning dataset generated using the stealth model Aurora Alpha.
The prompts from this dataset were almost all generated by GPT 5.1 and Gemini 3 (flash and pro).
The categories covered include academia, multi-lingual creative writing, finance, health, law, marketing/SEO, programming, philosophy, web dev, python scripting, and science.
Stats:
Cost: $ 0 (USD)
Tokens (input + output): 54.1 M
auroraapps-small
APPS Dataset
Dataset Description
APPS is a benchmark for code generation with 10000 problems. It can be used to evaluate the ability of language models to generate code from natural language specifications.
You can also find APPS metric in the hub here codeparrot/apps_metric.
Languages
The dataset contains questions in English and code solutions in Python.
Dataset Structure
from datasets import load_dataset
load_dataset("codeparrot/apps")… See the full description on the dataset page: https://huggingface.co/datasets/AuroraH456/apps-small.Liquid_V1_7B-pico-aurora-vidgen-multiturn-annotationspdf-tools_filesystem_github_terminal_huggingface_2118_x94fu9_manifest_aurora
Aurora Flow - Release Manifest
Candidate code name: aurora
Version: 2.1.0
Status: stable
License: MIT
Language: Python
Description: Real-time streaming data processing library
Maintainer: NovaTech Data Engineering
Last release: 2026-07-30
This dataset contains manifest.json (the release manifest) and validate.py
(the standardized validation script). Run python validate.py from this
directory to validate the release manifest.
aurora-dataset-roleplay-ptbr
Aurora Dataset Roleplay 🌌 (PT-BR)
O que é É um dataset que contém mais de 2 mil diálogos em português do Brasil. Ainda que tenha sido gerado sinteticamente, foi utilizado apenas modelos SOTA, então os diálogos são muito próximos da naturalidade e espontaneidade de um ser humano. Foi feito pensando em roleplay, por isso os diálogos contém nuances psicológicas, cenários diversos e personagens com estilos de falas e motivações complexas.
Modos de Geração… See the full description on the dataset page: https://huggingface.co/datasets/wilsondesouza/aurora-dataset-roleplay-ptbr.aurora_programmer_data
My Awesome Dataset
A comprehensive description of my awesome dataset.
Dataset Description
This dataset contains images of cats and dogs. The images were collected from [mention data source(s), e.g., a specific website, scraped from the internet]. It is intended for use in image classification tasks. The dataset consists of [number] images, with approximately [percentage]% allocated to the training set and [percentage]% to the test set. [Add more details about the… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/aurora_programmer_data.aurora-gpt4auroraai-model-registry-5315DreadPoor__Aurora_faustus-8B-LINEAR-details
Dataset Card for Evaluation run of DreadPoor/Aurora_faustus-8B-LINEAR
Dataset automatically created during the evaluation run of model DreadPoor/Aurora_faustus-8B-LINEAR
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Aurora_faustus-8B-LINEAR-details.DreadPoor__Aurora_faustus-8B-LORABLATED-details
Dataset Card for Evaluation run of DreadPoor/Aurora_faustus-8B-LORABLATED
Dataset automatically created during the evaluation run of model DreadPoor/Aurora_faustus-8B-LORABLATED
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Aurora_faustus-8B-LORABLATED-details.aurora-m-dataset-part-2This is part of the continued pretraining dataset used to train the Aurora-M model described in [Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code, COLING 2025] (https://aclanthology.org/2025.coling-industry.56/).
DreadPoor__Aurora_faustus-8B-LORABLATED_ALT-details
Dataset Card for Evaluation run of DreadPoor/Aurora_faustus-8B-LORABLATED_ALT
Dataset automatically created during the evaluation run of model DreadPoor/Aurora_faustus-8B-LORABLATED_ALT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Aurora_faustus-8B-LORABLATED_ALT-details.huggingface_4062_aurora_audio_w6rn6vr1auroraaurora-m2aurora-mix-data-baize-formatjaspionjader__Kosmos-Aurora_faustus-8B-details
Dataset Card for Evaluation run of jaspionjader/Kosmos-Aurora_faustus-8B
Dataset automatically created during the evaluation run of model jaspionjader/Kosmos-Aurora_faustus-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaspionjader__Kosmos-Aurora_faustus-8B-details.apps-small-editedaurora-worldjaspionjader__Kosmos-Elusive-VENN-Aurora_faustus-8B-details
Dataset Card for Evaluation run of jaspionjader/Kosmos-Elusive-VENN-Aurora_faustus-8B
Dataset automatically created during the evaluation run of model jaspionjader/Kosmos-Elusive-VENN-Aurora_faustus-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaspionjader__Kosmos-Elusive-VENN-Aurora_faustus-8B-details.lora-aurora-v2
