datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu_greek
Dataset Card for MMLU Greek
The MMLU Greek dataset is a set of 15858 examples from the MMLU dataset [available from here and here], machine-translated into Greek. The original dataset consists of multiple-choice questions from 57 tasks including elementary mathematics, US history, computer science, law, etc.
Dataset Details
Bias, Risks, and Limitations
This dataset is the result of machine translation.
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/mmlu_greek.SimBarca
🚥 SimBarca: Simulated Barcelona Traffic
This is the dataset accompanying https://github.com/Weijiang-Xiong/OpenSkyTraffic
📚 Publications and Citation
Weijiang Xiong, Robert Fonod, Alexandre Alahi and Nikolas Geroliminis, "Multi-Source Urban Traffic Flow Forecasting With Drone and Loop Detector Data," in IEEE Transactions on Intelligent Transportation Systems.
[Paper]
[Preprint]
[Exp. Config]
Weijiang Xiong, Robert Fonod and Nikolas Geroliminis, "Unveiling… See the full description on the dataset page: https://huggingface.co/datasets/Greatriver/SimBarca.leetcodeGreenChallengeData
GreenChallengeData
Demonstrations of a bimanual humanoid robot performing manipulation tasks in simulation, for training vision-language-action (VLA) policies. The repository gathers four collections: three task-specific deliveries recorded by human teleoperation and augmented with synthetic trajectories, plus a large set of scripted-policy episodes covering 15 tasks.
Format: LeRobot v2.1 · 30 fps · 3 cameras, 448×448 (H.264) · 51-dim state, 52-dim action
Total: 24,055 episodes… See the full description on the dataset page: https://huggingface.co/datasets/SberRoboticsCenter/GreenChallengeData.GreekMMLU
GreekMMLU
GreekMMLU is a native-sourced benchmark for evaluating massive multitask language understanding in Greek, built from authentic Greek exam-style multiple-choice questions (MCQ) rather than machine-translated English benchmarks.
21,805 questions across 45 subjects
4 high-level groups: STEM, Humanities, Social Sciences, Other
Difficulty/education levels spanning Primary → Secondary → University → Professional (+ an N/A bucket)
Public vs. private split for… See the full description on the dataset page: https://huggingface.co/datasets/dascim/GreekMMLU.synthetic_text_to_sql
Image generated by DALL-E. See prompt for more details
synthetic_text_to_sql
gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples,
designed and generated using Gretel Navigator, and released under Apache 2.0.
Please see our release blogpost for more details.
The dataset includes:
105,851 records partitioned into 100,000 train and 5,851 test records
~23M total tokens, including ~12M SQL tokens
Coverage across 100 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_text_to_sql.greek-cc
Greek Common Crawl
A FineWeb-style Greek-language text dataset extracted from Common Crawl, following the FineWeb-2 recipe adapted for Greek (ell_Grek).
Pipeline source: github.com/alexliap/greek-cc.
Crawl coverage starts at CC-MAIN-2024-22 rather than Common Crawl's earliest snapshots: this project picks up right where the FineWeb-2 dataset's own Greek (ell_Grek) subset leaves off (2013 through April 2024), so it extends FineWeb-2's Greek coverage forward instead of… See the full description on the dataset page: https://huggingface.co/datasets/alexliap/greek-cc.openmarket-btc-polymarket
OpenMarket BTC Polymarket
OpenMarket BTC Polymarket is a high-frequency research dataset pairing
Binance BTC/USDT market data with Polymarket BTC binary-market order book
events.
The dataset is released to support reproducible prediction-market
research, feature engineering, market microstructure analysis, and
backtesting.
Paper
This dataset accompanies OpenMarket: A Synchronized Polymarket-Binance
Dataset for High-Frequency Prediction-Market
Research… See the full description on the dataset page: https://huggingface.co/datasets/gregyoung14/openmarket-btc-polymarket.photoTgdatagreen-books-thumbnails
African American Travel Guides: Listing Thumbnails
113,053 pre-cropped thumbnail images — one per listing — from 50 volumes of mid-20th-century African American travel guides (1930–1966). Each image is a small snippet of a scanned directory page cropped down to a single business or lodging listing: the exact region a traveler would have read.
This is the image companion to the structured-listings dataset hadro/green-books-travel-guides. That dataset holds the transcribed text of… See the full description on the dataset page: https://huggingface.co/datasets/hadro/green-books-thumbnails.synthetic_pii_finance_multilingual
Image generated by DALL-E. See prompt for more details
💼 📊 Synthetic Financial Domain Documents with PII Labels
gretelai/synthetic_pii_finance_multilingual is a dataset of full length synthetic financial documents containing Personally Identifiable Information (PII), generated using Gretel Navigator and released under Apache 2.0.
This dataset is designed to assist with the following use cases:
🏷️ Training NER (Named Entity Recognition) models to detect and label PII in… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual.gretel-pii-masking-en-v1
Gretel Synthetic Domain-Specific Documents Dataset (English)
This dataset is a synthetically generated collection of documents enriched with Personally Identifiable Information (PII) and Protected Health Information (PHI) entities spanning multiple domains.
Created using Gretel Navigator with mistral-nemo-2407 as the backend model, it is specifically designed for fine-tuning Gliner models.
The dataset contains document passages featuring PII/PHI entities from a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-pii-masking-en-v1.kimi-k3-coding-and-debugging-traces
Kimi K3 Coding, Tool Use & Instruction Following Traces
582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables
below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/kimi-k3-coding-and-debugging-traces.cc-re-2020-filtered
Auto-Generated FastDetector Dataset
Model Name: google/gemma-4-E4B-it
Sampling Params (as sent to the engine): {"temperature": 0.0, "top_p": 1.0, "presence_penalty": 0.0}
Ignored Params (unsupported by this engine): None
Prompt File: prompts/filter_contiguous_subset.json
Total Train Prompts: 1
Source Dataset: G-reen/cc-re-2020-raw-sharded
Source Column: text
Target Num Samples: all
Dropped Samples (over length limit 15000 tokens): 430
Failed API Requests: 495
Total… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-re-2020-filtered.grenGalaxea_R1_Lite_classify_object_green_tablecloth
Galaxea_R1_Lite_classify_object_green_tablecloth
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 224
Total Frames: 251221
FPS: 30
Dataset Size: 25.62 GB
Robot Name: Galaxea_R1_Lite
End-Effector Type: two_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Galaxea_R1_Lite_classify_object_green_tablecloth.gpt-5.6-sol-coding-and-debugging-traces
GPT-5.6 Sol Coding & Debugging Traces
Verified software-engineering, independent model-judging, seed-authoring,
defensive-security, and training-harness trajectories from
GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an
autonomous coding agent. Sessions show the observable development loop:
inspecting repositories, reproducing failures, explaining evidence, editing
files, running compilers and test suites, correcting mistakes, and verifying
the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/gpt-5.6-sol-coding-and-debugging-traces.seizure_eeg_greyscale_224x224_6secWindow
Dataset Card for "seizure_eeg_greyscale_224x224_6secWindow"
More Information needed
open-greek-corpus-annotations
Open Greek Corpus Annotations
Token-level linguistic annotations for the
Open Greek Corpus:
lemma, part of speech (UD UPOS), and morphology (UD features) for every
served token. Three provenance classes, never confused thanks to per-token
provenance and confidence tiers: gold treebank annotations where an openly
licensed MANUAL treebank covers a work (GLAUx's treebank layers, MACULA
Greek for the NT), GLAUx's own automatic annotation as the middle auto:
class, and model… See the full description on the dataset page: https://huggingface.co/datasets/ciscoriordan/open-greek-corpus-annotations.green-vla-sft-trackio-datasymptom_to_diagnosis
Dataset Summary
This dataset contains natural language descriptions of symptoms labeled with 22 corresponding diagnoses. Gretel/symptom_to_diagnosis provides 1065 symptom descriptions in the English language labeled with 22 diagnoses, focusing on fine-grained single-domain diagnosis.
Data Fields
Each row contains the following fields:
input_text : A string field containing symptoms
output_text : A string field containing a diagnosis
Example:
{
"output_text": "drug… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/symptom_to_diagnosis.Agilex_Cobot_Magic_move_object_green_tablecloth
Agilex_Cobot_Magic_move_object_green_tablecloth
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 197
Total Frames: 85405
FPS: 30
Dataset Size: 5.33 GB
Robot Name: Agilex_Cobot_Magic
End-Effector Type: two_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_move_object_green_tablecloth.GreenChallengeAssetsimagenet-latents-imagesimport os
os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
from datasets import load_dataset
dataset = load_dataset("G-REPA/imagenet-latents-images", split="train")
cc-re-2021-stat-val
Auto-Generated FastDetector Dataset
Dataset: G-reen/cc-re-2021-stat-val
Globals Config: config/globals_re2021.toml
Analysis Config: config/analysis_nofilter.toml
Rows: 9,822
Evaluation Results
Prompt Subsets: 4 (direct_reference, indirect_reference, revise, rewrite)
Generator Configs: 5 (Hy3-NVFP4-FP8 (Temp: 0.9), Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6), Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7), claude-sonnet-5 (Temp: Unknown), gpt-5.6-luna (Temp:… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-re-2021-stat-val.grefgRefCOCO
AT_Great
EMG dataset for gesture recognition with arm translation
https://doi.org/10.5061/dryad.8sf7m0czv
Corresponding author: Iris Kyranou, email: iriskyr@gmail.com
Description of the data and file structure
Subjects
8 intact subjects.
Acquisition Setup
The sEMG data are acquired using four Trigno Quattro Sensors (https://delsys.com/trigno-quattro/), while kinematic data are acquired using a Cyberglove II data glove… See the full description on the dataset page: https://huggingface.co/datasets/KasiaAI/AT_Great.rcs_utn_green_box
Dataset Card for RCS UTN Green Box (FiftyOne)
rcs_utn_green_box is a grouped FiftyOne video dataset of a multi-view robot
manipulation task — "pick the green box" — collected with the
Robot Control Stack (RCS) ecosystem from the University of Technology
Nuremberg. Each episode is a group with one synchronized video per camera, plus
dense robot proprioception and action data on every frame.
Installation
If you haven't already, install FiftyOne:
pip install -U… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/rcs_utn_green_box.greek-contents
