datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMMLU
Multilingual Massive Multitask Language Understanding (MMMLU)
The MMLU is a widely recognized benchmark of general knowledge attained by AI models. It covers a broad range of topics from 57 different categories, covering elementary-level knowledge up to advanced professional subjects like law, physics, history, and computer science.
We translated the MMLU’s test set into 14 languages using professional human translators. Relying on human translators for this evaluation increases… See the full description on the dataset page: https://huggingface.co/datasets/openai/MMMLU.BrowseCompLongContext
BrowseComp Long Context
BrowseComp Long Context is a dataset based on BrowseComp to benchmark LLM’s capability to retrieve relevant information from noisy data in its context. It converts the agentic question answering tasks from Browsecomp into long context tasks.
For each of the questions in a subset of BrowseComp, a list of urls are attached. Each url will be paired with an indicator indicating whether the content of the web page is required to answer the question or is… See the full description on the dataset page: https://huggingface.co/datasets/openai/BrowseCompLongContext.openai-terra-batch-wiki-brazil-1000-partial-20260724-01
OpenAI Terra Batch — Wikipédia PT-BR (run parcial)
Checkpoint publicável de uma execução real e interrompida do fluxo
document_task_matrix. A execução planejou gerar uma matriz de 1.000
documentos da Wikipédia em português por 25 tasks canônicas usando a Responses
API Batch e o modelo gpt-5.6-terra.
Este repositório não representa a conclusão dos 25.000 pares planejados. Ele
contém somente os 1.282 candidatos aceitos após a reconciliação offline de
todos os resultados Batch já… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/openai-terra-batch-wiki-brazil-1000-partial-20260724-01.medmcqa-openai-native
MedMCQA — OpenAI-native, with a usable test split
MedMCQA is one of the most downloaded medical QA datasets on the Hub. Its test split has been unusable since release: all 6,150 rows carry cop=-1 (no label) and an empty explanation. You cannot score a model on it.
This release rebuilds a labelled, leak-free test split and converts everything to the native messages format, so it loads straight into TRL with no custom parsing.
What was actually wrong
Measured on the… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/medmcqa-openai-native.openai-gsm8k-enhanced-using-together-ai-deepseek-train8k-test1k-v1
OpenAI GSM8K Enhanced with DeepSeek API
🔥 Exciting News! We're thrilled to release a new dataset, meticulously curated using the open-source #OpenAI #GSM8K dataset and enhanced with chain-of-thought reasoning (CoT) via the DeepSeek API from #TogetherAI.
🔗 Access the Dataset: OpenAI GSM8K Enhanced
What’s Cooking? 🍳
Dataset Specifications
Total Samples: ~10K, with about 8K training and 1K testing entries.
Enhancements: Each sample is enhanced with CoT… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/openai-gsm8k-enhanced-using-together-ai-deepseek-train8k-test1k-v1.MetaMathQA-decontaminated-openai-native
MetaMathQA — decontaminated, OpenAI-native
MetaMathQA is a widely used math fine-tuning corpus. Its README states:
"None of the augmented data is from the testing set."
That is false, and this release proves it with measurements. 24,334 rows (6.16%) overlap with standard evaluation splits. If you fine-tune on the original and report MATH or GSM8K scores, those scores are inflated.
This release removes the leakage, converts to native messages, and documents every rejection.… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/MetaMathQA-decontaminated-openai-native.smoke-openai-terra-batch-brasil-25-20260724-01
Smoke OpenAI Terra Batch — Brasil × 25 tasks
Run real de validação do fluxo matricial document_task_matrix, executada
sobre um único documento da Wikipédia em português com o título Brasil.
Cada uma das 25 tasks canônicas recebeu exatamente um slot inicial.
Resultado
status: completed
documentos: 1
pares planejados: 25
exemplos aceitos: 25
pares pulados: 0
pares esgotados: 0
resultados reais do backend: 27
retries com nova chamada: 2
backend: openai_api… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/smoke-openai-terra-batch-brasil-25-20260724-01.gpt-oss-120B-distilled-math-OpenAI-Harmony
📚 Dataset Overview
Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines (.jsonl)Fields: Generator, Category, Input, Output
Note: If you are using this template for training, please make sure the format is correct before starting.Since this template is still under continuous improvement and learning, it may not be fully complete yet. I appreciate your understanding.
📈 Core Statistics
Generated complete reasoning processes… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120B-distilled-math-OpenAI-Harmony.Arabic_Openai_MMMLU
Arabic Multilingual Massive Multitask Language Understanding (MMMLU)
The MMLU is a widely recognized benchmark for assessing general knowledge attained by AI models. It covers a broad range of topics across 57 different categories, from elementary-level knowledge to advanced professional subjects like law, physics, history, and computer science.
We have extracted the Arabic subset from the MMMLU test set, which was translated by professional human translators. This dataset, now… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic_Openai_MMMLU.open-australian-legal-embeddings-openaiopenai_humaneval-th
HumanEval-th
A Thai translation of all 164 problems of OpenAI's HumanEval. Every row corresponds
1:1, in order, to a row of the English original, so the Thai and English scores of a
model are directly comparable.
Only the prompt column is Thai. canonical_solution, test and entry_point are
Python rather than prose and were never translated; they are byte-identical to
openai/openai_humaneval in all 164 rows, and validate.py checks that on every run.
Every revision of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/iapp/openai_humaneval-th.Formated-openai-function-invocations-20k-with-greetings
About
This dataset is the formated version of the Isaak-Carter/Openai-function-invocations-20k-with-greetings dataset.
This dataset, uniquely structured with custom special tokens, is meticulously crafted to train language models in complex function invocation and time-contextualized interactions. Each "sample" in the dataset contains a sequence of elements: function definitions, user prompts, function calls, function responses, and the assistant's responses. These elements are… See the full description on the dataset page: https://huggingface.co/datasets/Goekdeniz-Guelmez/Formated-openai-function-invocations-20k-with-greetings.ScienticDatasetArxiv-openAI-FormatV3
📚 Scientific Dataset Arxiv OpenAI Format
This dataset contains scientific data transformed for use with OpenAI models. It includes detailed descriptions and structures designed for machine learning applications. The original data was taken from:
from datasets import load_dataset
dataset = load_dataset("taesiri/arxiv_qa")
📂 Dataset Structure
The dataset is organized into a training split with comprehensive features tailored for scientific document processing:
Train… See the full description on the dataset page: https://huggingface.co/datasets/ejbejaranos/ScienticDatasetArxiv-openAI-FormatV3.
