datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OR-Space
OR-Space
A full-lifecycle workspace benchmark for industrial optimization agents.
OR-Space evaluates whether language-model agents can work reliably with
operations research problems represented as executable, multi-file workspaces.
Rather than presenting a self-contained mathematical prompt, each task
distributes evidence across business requirements, structured data, source
code, execution logs, and solver records.
The benchmark contains 100 optimization topologies. Each… See the full description on the dataset page: https://huggingface.co/datasets/Chenyu-Zhou/OR-Space.SpatialReasoning
The Spatial Reasoning Dataset
The Spatial Reasoning Dataset comprises semantically meaningful question-answer pairs focused on the relative locations of geographic divisions within the United States — including states, counties, and ZIP codes.
The dataset is designed for spatial question answering and includes three types of questions:
Binary (Yes/No)
Single-choice (Radio)
Multi-choice (Checkbox)
All questions include correct answers for training and evaluation purposes.… See the full description on the dataset page: https://huggingface.co/datasets/Rammen/SpatialReasoning.NewsLensSync
Dataset Card for NewsLensSync
Dataset Description
This dataset, named NewsLensSync, contains a curated collection of news articles, sourced from trusted domains such as BBC, Reuters, AP News, NPR, PBS, The Guardian, WSJ, NY Times, and ProPublica. Each article includes both the original content and a synthetic "falsified" version of the article description, generated using a transformer-based negation model. The dataset is designed for research in misinformation… See the full description on the dataset page: https://huggingface.co/datasets/sparklessszzz/NewsLensSync.SpaceOmicsBench-v3
SpaceOmicsBench v3
A Multi-Omics AI Benchmark for Spaceflight Biomedical Data
SpaceOmicsBench v3 provides standardized ML and LLM evaluation infrastructure for spaceflight biomedical data from 4 human spaceflight missions (NASA Twins Study, Inspiration4, JAXA cfRNA, Axiom-2).
Dataset Structure
ML Track (Track A)
tasks/track_a/ — Task definitions (J1: phase classification, J2: clock acceleration)
tasks/track_c/ — Feature-level task definitions (C1:… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/SpaceOmicsBench-v3.deep-space-optical-chip-thermal-dataset
🚀 Deep Space Optical Chip Thermal Dataset 🪐
🌡️ 40,000 scenario-based prompt and response pairs on thermal mitigation for photonic chips in scientific instruments aboard deep-space probes, covering refractive index drift, waveguide misalignment, and thermal stress across materials, instruments, and environments.
⚠️ Disclaimer: All entries are synthetically generated. Material coefficients are drawn from published typical values, but no row is based on mission logs or flight… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/deep-space-optical-chip-thermal-dataset.SpanishCasualChat
SpanishCasualChat
tags: conversational-AI, dialogue-system, Spanish-language
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'SpanishCasualChat' dataset is a collection of Spanish conversational text aimed at training machine learning models for dialogue systems. It contains dialogues of various topics such as hobbies, travel experiences, and daily routines, presented in a casual and lengthy format to simulate real-life… See the full description on the dataset page: https://huggingface.co/datasets/mariagrandury/SpanishCasualChat.OR-Space
OR-Space
A full-lifecycle workspace benchmark for industrial optimization agents.
OR-Space evaluates whether LLM agents can do reliable operations research work
inside executable, multi-file workspaces. Each instance keeps business
requirements, parameter files, source code, solver artifacts, and evaluation
metadata as separate files, forcing the agent to recover and maintain the
optimization model through workspace interaction rather than one-shot text
generation.… See the full description on the dataset page: https://huggingface.co/datasets/YiYao7017/OR-Space.El-TARA_Spanish_LLM_Benchmark
El-Tara: Evaluación de Razonamiento Avanzado en Español
Dataset Summary
El-Tara (Evaluación de Razonamiento Avanzado en Español) is a benchmark dataset designed to assess the advanced reasoning capabilities of Large Language Models (LLMs) in Spanish. It is adapted from the original TARA (Turkish Advanced Reasoning Assessment) dataset.
Similar to TARA, El-Tara aims to test higher-order cognitive skills across multiple domains, using synthetically generated questions… See the full description on the dataset page: https://huggingface.co/datasets/emre/El-TARA_Spanish_LLM_Benchmark.spacy-standardSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: Spacy
Data Source Link: https://spacy.io/usage
Data Source License: https://github.com/explosion/spaCy/blob/master/LICENSE
Data Source Authors: ExplosionAI GmbH, 2016 spaCy GmbH, 2015 Matthew Honnibal
AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
wiki_sparql
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/farhadali/wiki_sparql.spacy-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: Spacy
Data Source Link: https://spacy.io/usage
Data Source License: https://github.com/explosion/spaCy/blob/master/LICENSE
Data Source Authors: ExplosionAI GmbH, 2016 spaCy GmbH, 2015 Matthew Honnibal
AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
OpenSpatialLogic
OpenSpatialLogic
Dataset Card for OpenSpatialLogic
Dataset Summary
OpenSpatialLogic is a handcrafted dataset of 50 riddles which test understanding of spatial relationships in reality.
These include questions about compass directions, ordering of bricks within towers after transformations, and permeability of objects in certain configurations.
The only way for a model to get better at something is to train on data about it. Large Language Models are bad at… See the full description on the dataset page: https://huggingface.co/datasets/spacekat99/OpenSpatialLogic.
