datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multi-Opthalingua
Cite
Accepted to AAAI 2025 (https://openreview.net/group?id=AAAI.org/2025/Conference#tab-recent-activity)
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs:
@misc{restrepo2024multiophthalinguamultilingualbenchmarkassessing,
title={Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs},
author={David Restrepo and Chenwei Wu and Zhengxu Tang and Zitao Shuai and Thao… See the full description on the dataset page: https://huggingface.co/datasets/AAAIBenchmark/Multi-Opthalingua.deep-space-optical-chip-thermal-dataset
🚀 Deep Space Optical Chip Thermal Dataset 🪐
🌡️ 40,000 scenario-based prompt and response pairs on thermal mitigation for photonic chips in scientific instruments aboard deep-space probes, covering refractive index drift, waveguide misalignment, and thermal stress across materials, instruments, and environments.
⚠️ Disclaimer: All entries are synthetically generated. Material coefficients are drawn from published typical values, but no row is based on mission logs or flight… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/deep-space-optical-chip-thermal-dataset.medqa-5-opt-MedGENIE
Dataset Card for "medqa-5-opt-MedGENIE"
Dataset Description
The data is a part of the MedGENIE collection of medical datasets augmented with artificial contexts generated by PMC-LLaMA-13B. Specifically, up to 5 artificial contexts were generated for each question in MedQA-USMLE (5 options), employing a multi-view approach to encompass various perspectives associated with the given question.
The dataset has been used to train MedGENIE-fid-flan-t5-base-medqa allowing it to… See the full description on the dataset page: https://huggingface.co/datasets/disi-unibo-nlp/medqa-5-opt-MedGENIE.optimot-linguistic-data
Optimot Linguistic Data
This dataset contains 4,011 entries extracted from the public Optimot linguistic consultation service of the Departament de Política Lingüística, Generalitat de Catalunya.
Each record addresses a Catalan language question or linguistic topic and includes an explanation, source metadata, and a direct source URL when available.
Data
The dataset is provided as JSON Lines:
optimot.jsonl
Each row contains:
Fitxa: Optimot card identifier.… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/optimot-linguistic-data.health-optimization-bench-sample
Health Optimization Bench (Sample)
A 30-task public sample of Health Optimization Bench,
a rubric-graded benchmark measuring how well frontier language models handle current clinical
evidence in preventive and optimization medicine. Three tasks from each of the benchmark's ten
micro benches.
The full benchmark is 977 authored tasks with 346 released across ten micro benches. On the
current leaderboard no model scores above 71 of 100 and the field spans 66 points. Rankings:… See the full description on the dataset page: https://huggingface.co/datasets/Arcophos/health-optimization-bench-sample.quantum-optimization
Neura Parse — Quantum Optimization, Annealing & Finance: QAOA, Adiabatic Methods & the Advantage Question
A research-plus-practitioner vertical on quantum approaches to combinatorial and continuous optimization and their most-piloted enterprise use cases. Covers QAOA theory and variants, adiabatic/annealing methods and D-Wave, QUBO/Ising encodings, amplitude-estimation Monte Carlo for finance, and the rigorous question of whether and where quantum beats classical (including… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-optimization.MedQA-USMLE-4-options-clean
MedQA-USMLE-4-options-clean Dataset
Overview
MedQA-USMLE-4-options-clean is an enhanced medical question-answering benchmark that builds upon the MedQA-USMLE dataset. Physicians analyzed the 1373 questions in the original dataset and moved 52 questions that were either malformed or incomplete to another split incomplete.
Key Features
Relabeled malformed/incorrect questions
Dataset Details
Size: 1373
Language: English
Data Source… See the full description on the dataset page: https://huggingface.co/datasets/maximegmd/MedQA-USMLE-4-options-clean.mmlu-5-options-rl-ready
MMLU – 5-Options RL-Ready
A standardized, RL-friendly remix of MMLU with explicit negatives and a unified five-option presentation string for each question. Ideal for DPO and other RL setups while remaining drop-in for classic multiple-choice evaluation.
What’s inside
Splits & size: ~97.8k train + 2k test ≈ 99.8k total.
Schema (core fields):
question: str
choices: list[str] (canonical options, typically 4 as in original MMLU)
answer: int (0-based index)
task: str… See the full description on the dataset page: https://huggingface.co/datasets/openmed-community/mmlu-5-options-rl-ready.MedQA-USMLE-4-options-hfoptimuskg-neo4j-graph
OptimusKG → Neo4j Graph
A property-graph conversion of OptimusKG,
ready for direct import into Neo4j or general tabular/graph use, published in
two formats (see "Files" below).
OptimusKG is a modern biomedical knowledge graph developed by the
Zitnik Lab, Harvard Medical School
(Department of Biomedical Informatics). It integrates 65 heterogeneous
source resources — spanning molecular, anatomical, clinical, and
environmental domains — grounded in 18 ontologies, built with the… See the full description on the dataset page: https://huggingface.co/datasets/ksk-1729/optimuskg-neo4j-graph.Arabic-Optimized-Reasoning-Dataset
Arabic Optimized Reasoning Dataset
Dataset Name: Arabic Optimized ReasoningLicense: Apache-2.0Formats: CSVSize: 1600 rowsBase Dataset: cognitivecomputations/dolphin-r1Libraries Used: Datasets, Dask, Croissant
Overview
The Arabic Optimized Reasoning Dataset helps AI models get better at reasoning in Arabic. While AI models are good at many tasks, they often struggle with reasoning in languages other than English. This dataset helps fix this problem by:
Using fewer tokens… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/Arabic-Optimized-Reasoning-Dataset.MedQA-USMLE-4-options_preprocess
MedQA-USMLE Preprocessed Dataset
This dataset is a preprocessed version of GBaker/MedQA-USMLE-4-options.
The data has been formatted into a question and answer structure suitable for training or evaluating instruction-following language models.
Data Structure
question: The original medical question combined with the four multiple-choice options.
answer: The correct answer index, prefixed with ####.
Example
Question:
A 60-year-old woman comes to the emergency… See the full description on the dataset page: https://huggingface.co/datasets/LLMcompe-Team-Watanabe/MedQA-USMLE-4-options_preprocess.synthetic-code-optimization-1synthetic-code-optimization-1 is a synthetic dataset with a total of ~1136 Question and Answer pairs.
This dataset was generated using the following models:
ChatGPT:
Whatever is hosted on their website
Claude:
Fable 5
Deepseek:
Deepseek "Instant"
Deepseek "Expert"
Gemini:
3.1 Flash Lite
3.5 Flash
3.1 Pro
Grok:
Fast
Mistral:
Thinking enabled
Qwen 3.7 Plus:
Thinking enabled
GLM 5.2:
Thinking "high"
Perplexity:
Whatever is on their website
This dataset follows the following… See the full description on the dataset page: https://huggingface.co/datasets/takenusername32/synthetic-code-optimization-1.fr-medqa-5_optionsoptic_QA_pdf_russianswiss-web-premium-ch
*.ch Swiss Web Premium (A+)
Grade A+ -- 112.4M tokens -- 110,491 records -- 29 languages -- Full provenance -- PII-redacted -- RAG-ready -- SFT-formatted
A production-grade Swiss web corpus from the .ch TLD namespace. 110,491 documents independently quality-scored (avg 93.3/100, minimum 90), PII-redacted, and SHA256-verified. Built for LLM training, RAG pipelines, SFT fine-tuning, and multilingual NLP.
OptiTransferData Portfolio
Premium sovereign web corpora for… See the full description on the dataset page: https://huggingface.co/datasets/OptiTransferData/swiss-web-premium-ch.qa-dataset-20250411qa-dataset-20250430swiss-web-premium-ch-full
*.ch Swiss Web Premium (A+) -- Full Dataset
Grade A+ -- 112.4M tokens -- 110,491 records -- 29 languages -- 22 production files -- 554.1 MB
The complete production release of OptiTransfer's Swiss web corpus. 110,491 documents from the .ch TLD namespace, independently quality-scored (avg 93.3/100, minimum 90), PII-redacted, and SHA256-verified. Delivered in Parquet, JSONL, language splits, and pre-built RAG chunks.
This is the full commercial dataset. For evaluation, see the free… See the full description on the dataset page: https://huggingface.co/datasets/OptiTransferData/swiss-web-premium-ch-full.adaption-marketing-optimized-neural-titans
Adaption Marketing Optimized Dataset - Neural Titans
Competition: Adaption AutoScientist Challenge ($50,000 Prize Pool)Track: MarketingTeam: Neural Titans (HackIndia)
Dataset Details
Metric
Value
Rows
5,000
Size
22.5 MB
Format
JSONL (instruction-tuning)
Pipeline Configuration
Recipes Applied
Deduplication - Removes duplicate and near-duplicate entries
Prompt Rephrasing - Diversifies prompt formulations for robust… See the full description on the dataset page: https://huggingface.co/datasets/rishini/adaption-marketing-optimized-neural-titans.
