datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hle-multimodaloptions-IV-SP500
Downloading the Options IV SP500 Dataset
This document will guide you through the steps to download the Options IV SP500 dataset from Hugging Face Datasets. This dataset includes data on the options of the S&P 500, including implied volatility.
To start, you'll need to install Hugging Face's datasets library if you haven't done so already. You can do this using the following pip command:
!pip install datasets
Here's the Python code to load the Options IV SP500 dataset from Hugging… See the full description on the dataset page: https://huggingface.co/datasets/gauss314/options-IV-SP500.btc-inverse-option-snapshotsstocks-optionsoptical-neuromorphic-eikonal-benchmarks
Optical Neuromorphic Eikonal Solver - Benchmark Datasets
Overview
Benchmark datasets for evaluating the Optical Neuromorphic Eikonal Solver, a GPU-accelerated pathfinding algorithm achieving 30-300× speedup over CPU Dijkstra.
🎯 Key Results
134.9× average speedup vs CPU Dijkstra
0.64% mean error (sub-1% accuracy)
1.025× path length (near-optimal paths)
2-4ms per query on 512×512 grids
📊 Dataset Content
5 synthetic pathfinding test cases covering… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/optical-neuromorphic-eikonal-benchmarks.Multi-Opthalingua
Cite
Accepted to AAAI 2025 (https://openreview.net/group?id=AAAI.org/2025/Conference#tab-recent-activity)
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs:
@misc{restrepo2024multiophthalinguamultilingualbenchmarkassessing,
title={Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs},
author={David Restrepo and Chenwei Wu and Zhengxu Tang and Zitao Shuai and Thao… See the full description on the dataset page: https://huggingface.co/datasets/AAAIBenchmark/Multi-Opthalingua.AI-Code-Optimization-for-Sustainability-Dataset
AI Code Optimization for Sustainability: Dataset
Refactoring Python Code for Energy-Efficiency using Qwen3: Dataset based on HumanEval, MBPP, and Mercury
📄 Read the Paper | Zenodo Mirror | DOI: 10.5281/zenodo.18377893 | About the author
This dataset is a part of a Master thesis research internship investigating the use of LLMs to optimize Python code for energy efficiency.
The research was conducted as part of the Greenify My Code (GMC) project at the Netherlands Organisation for… See the full description on the dataset page: https://huggingface.co/datasets/BambusControl/AI-Code-Optimization-for-Sustainability-Dataset.historical_option_dataoptogpt_dataWelcome!
This is the dataset for OptoGPT: A Foundation Model for Multilayer Thin Film Inverse Design (https://www.oejournal.org/article/doi/10.29026/oea.2024.240062)
Code for OptoGPT: https://github.com/taigaoma1997/optogpt
Model for OptoGPT: https://huggingface.co/mataigao/optogpt
File description:
/train
/data_train_0.csv /data_train_1.csv /data_train_2.csv /data_train_3.csv /data_train_4.csv /data_train_5.csv /data_train_6.csv /data_train_7.csv /data_train_8.csv… See the full description on the dataset page: https://huggingface.co/datasets/mataigao/optogpt_data.geometric_optics_physical_consistency_eval
Geometric Optics – Physical Consistency Evaluation Dataset
This dataset contains visual failure cases in geometric optics for multimodal image generation models.
Scope
The dataset focuses on physical and geometric inconsistencies related to:
mirror reflections (law of reflection)
refraction and dispersion in prisms
light ray direction consistency
shadow direction vs light source
camera–object–light spatial coherence
Motivation
Current image generation models… See the full description on the dataset page: https://huggingface.co/datasets/8Planetterraforming/geometric_optics_physical_consistency_eval.optioncharts.iodeep-space-optical-chip-thermal-dataset
🚀 Deep Space Optical Chip Thermal Dataset 🪐
🌡️ 40,000 scenario-based prompt and response pairs on thermal mitigation for photonic chips in scientific instruments aboard deep-space probes, covering refractive index drift, waveguide misalignment, and thermal stress across materials, instruments, and environments.
⚠️ Disclaimer: All entries are synthetically generated. Material coefficients are drawn from published typical values, but no row is based on mission logs or flight… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/deep-space-optical-chip-thermal-dataset.optimized_adult_census
Documentation of the Dataset - Adult Census Income Dataset Optimized
1. General Description of the Dataset
This dataset, called Adult Census Income Dataset Optimized, is an optimized version of the Adult Census Income Dataset. The latter comes from the UCI Machine Learning Repository and is commonly used in classification tasks to predict whether a person earns more or less than $50,000 per year based on various demographic characteristics.
We optimized the dataset by… See the full description on the dataset page: https://huggingface.co/datasets/Databoost/optimized_adult_census.Portfolio-Optimizationstock_options_tradingAI_BATTERY_OPTIMIZER
Dataset Card for AI Battery Optimizer
The AI Battery Optimizer Dataset contains synthetic smartphone battery usage logs created during the development of the AI Battery Optimizer App.It is intended for research and experimentation on battery prediction, app usage forecasting, and adaptive resource management.
Dataset Details
This dataset logs:
Battery percentage over time
Power usage (mW)
Estimated time remaining
Predicted app usage with confidence score
Screen… See the full description on the dataset page: https://huggingface.co/datasets/yu743/AI_BATTERY_OPTIMIZER.Optimizer-cluadequestiongendata-mlops-infra-predictive-optimization-finalclinical-counterfactual-coherence-optimal-path-selection-v0.1What this dataset tests
Whether a model can choose the counterfactual path that maximizessystemic coherence versus the real outcome.
Required outputs
optimal_intervention_choice
coherence_score_delta
justification_narrative
Choice set
REAL
CF1
CF2
CF3
Coherence means
explains and stabilizes multi-stream signals
avoids iatrogenic oscillation
reduces complication risk
supports a stable recovery basin
Typical failures
picking the highest single metric
skipping… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-counterfactual-coherence-optimal-path-selection-v0.1.Token_Optimization_Org
AI Safety & Bias Evaluation Conversations
Dataset Summary
This dataset contains simulated multi-turn conversations designed to evaluate AI language model behavior across two safety-critical domains: self-harm response handling and political bias. Each row represents a single evaluation scenario where an AI model's responses are assessed for safety compliance or neutrality. The dataset is intended to support research and development of safer, less biased AI systems.
All… See the full description on the dataset page: https://huggingface.co/datasets/token-opt-org/Token_Optimization_Org.AI_BATTERY_OPTIMIZER
Dataset Card for AI Battery Optimizer
The AI Battery Optimizer Dataset contains synthetic smartphone battery usage logs created during the development of the AI Battery Optimizer App.It is intended for research and experimentation on battery prediction, app usage forecasting, and adaptive resource management.
Dataset Details
This dataset logs:
Battery percentage over time
Power usage (mW)
Estimated time remaining
Predicted app usage with confidence score
Screen… See the full description on the dataset page: https://huggingface.co/datasets/laomalaiwan/AI_BATTERY_OPTIMIZER.Banking-77-OptimizedArabic-Optimized-Reasoning-Dataset
Arabic Optimized Reasoning Dataset
Dataset Name: Arabic Optimized ReasoningLicense: Apache-2.0Formats: CSVSize: 1600 rowsBase Dataset: cognitivecomputations/dolphin-r1Libraries Used: Datasets, Dask, Croissant
Overview
The Arabic Optimized Reasoning Dataset helps AI models get better at reasoning in Arabic. While AI models are good at many tasks, they often struggle with reasoning in languages other than English. This dataset helps fix this problem by:
Using fewer tokens… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/Arabic-Optimized-Reasoning-Dataset.optimization-trap-detection-v0.1
What this dataset does
This dataset tests whether a model can detect optimization traps.
The task is simple:
Given a scenario and an optimization-trap claim, predict whether the claim is supported.
Core stability idea
Optimization traps occur when improvement of a local metric damages the larger system.
Common patterns include:
metric fixation
local optimization
hidden tradeoffs
invariant violation
delayed costs
system-wide degradation
The optimized metric improves.
The… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/optimization-trap-detection-v0.1.engine-optimization-reference
ThatDevPro Engine Optimization Reference Corpus
A curated snapshot of canonical reference documentation for engine optimization (the merge of SEO + AEO + AIO + GEO), maintained by ThatDevPro.
Overview
This dataset indexes 26 canonical reference pages from the ThatDevPro Reference Library, each covering one specific crawler signal a website emits. The dataset is intended for training and evaluating AEO/GEO scoring systems, llms.txt validators, and structured-data… See the full description on the dataset page: https://huggingface.co/datasets/ThatDeveloperGuy13/engine-optimization-reference.AI_BATTERY_OPTIMIZER
Dataset Card for AI Battery Optimizer
The AI Battery Optimizer Dataset contains synthetic smartphone battery usage logs created during the development of the AI Battery Optimizer App.It is intended for research and experimentation on battery prediction, app usage forecasting, and adaptive resource management.
Dataset Details
This dataset logs:
Battery percentage over time
Power usage (mW)
Estimated time remaining
Predicted app usage with confidence score
Screen… See the full description on the dataset page: https://huggingface.co/datasets/wycwsjdw/AI_BATTERY_OPTIMIZER.legal-advice-email-risk-option-instruction-coherence-v0.1What this dataset does
You receive
case position
facts used
risk analysis
options
recommendation
client instruction
consistency flags
You decide
coherent
or
incoherent
Daily use
advice QC
risk gap detection
instruction capture check
contradiction flag
optics-syntheticCode-Optimization
Dataset Description
This dataset contains 3,000 pairs of Python code snippets designed to demonstrate common performance bottlenecks and their optimized counterparts. Each entry includes complexity analysis (Time and Space) and a technical explanation for the optimization.
Total Samples: 3,000
Language: Python 3.x
Focus: Performance engineering, Refactoring, and Algorithmic Efficiency.
Dataset Summary
The dataset is structured to help train or fine-tune models on… See the full description on the dataset page: https://huggingface.co/datasets/SeifElden2342532/Code-Optimization.Optimizer-llama370bgeneratedquestion
