datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aopsOR-Space
OR-Space
A full-lifecycle workspace benchmark for industrial optimization agents.
OR-Space evaluates whether language-model agents can work reliably with
operations research problems represented as executable, multi-file workspaces.
Rather than presenting a self-contained mathematical prompt, each task
distributes evidence across business requirements, structured data, source
code, execution logs, and solver records.
The benchmark contains 100 optimization topologies. Each… See the full description on the dataset page: https://huggingface.co/datasets/Chenyu-Zhou/OR-Space.SPADE-customer-service-dialogue
SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection
Paper | Code
SPADE contains a repository of customer service line synthetic user dialogues with goals, augmented from MultiWOZ 2.1 using GPT-3.5 and Llama 70B.
The datasets are intended for training and evaluating machine generated text detectors in dialogue settings.
There are 15 English datasets generated using 5 different augmentation methods and 2 large language models… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/SPADE-customer-service-dialogue.deep-space-optical-chip-thermal-dataset
🚀 Deep Space Optical Chip Thermal Dataset 🪐
🌡️ 40,000 scenario-based prompt and response pairs on thermal mitigation for photonic chips in scientific instruments aboard deep-space probes, covering refractive index drift, waveguide misalignment, and thermal stress across materials, instruments, and environments.
⚠️ Disclaimer: All entries are synthetically generated. Material coefficients are drawn from published typical values, but no row is based on mission logs or flight… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/deep-space-optical-chip-thermal-dataset.song-lyrics-artist-classifierOR-Space
OR-Space
A full-lifecycle workspace benchmark for industrial optimization agents.
OR-Space evaluates whether LLM agents can do reliable operations research work
inside executable, multi-file workspaces. Each instance keeps business
requirements, parameter files, source code, solver artifacts, and evaluation
metadata as separate files, forcing the agent to recover and maintain the
optimization model through workspace interaction rather than one-shot text
generation.… See the full description on the dataset page: https://huggingface.co/datasets/YiYao7017/OR-Space.spanish_imdb_synopsis
Dataset Card for Spanish IMDb Synopsis
Dataset Description
4969 movie synopsis from IMDb in spanish.
Dataset Summary
[N/A]
Languages
All descriptions are in spanish, the other fields have some mix of spanish and english.
Dataset Structure
[N/A]
Data Fields
description: IMDb description for the movie (string), should be spanish
keywords: IMDb keywords for the movie (string), mix of spanish and english
genre: The genres of the… See the full description on the dataset page: https://huggingface.co/datasets/mathigatti/spanish_imdb_synopsis.El-TARA_Spanish_LLM_Benchmark
El-Tara: Evaluación de Razonamiento Avanzado en Español
Dataset Summary
El-Tara (Evaluación de Razonamiento Avanzado en Español) is a benchmark dataset designed to assess the advanced reasoning capabilities of Large Language Models (LLMs) in Spanish. It is adapted from the original TARA (Turkish Advanced Reasoning Assessment) dataset.
Similar to TARA, El-Tara aims to test higher-order cognitive skills across multiple domains, using synthetically generated questions… See the full description on the dataset page: https://huggingface.co/datasets/emre/El-TARA_Spanish_LLM_Benchmark.awesome_chatgpt_prompts_ar
📦 Awesome Arabic Chatgpt Prompts
📝 Overview
This repository contains a collection of Arabic prompts designed for use with AI language models (such as ChatGPT).
The goal is to provide a lightweight dataset that helps Arabic-speaking users quickly get started with generative AI.
🔗 Website / Demo
Check out the live demo site:omarnj-lab.github.io/awesome_chatgpt_prompts_ar
✨ Features
Entirely in Arabic 🕌
Suitable for educational and… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/awesome_chatgpt_prompts_ar.
