datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Llama3-SSL4EO-S12-v1.1-captions
Llama3-SSL4EO-S12-Captions
The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model.
Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper.
Code: https://github.com/IBM/MS-CLIP
Data Structure
We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.LithoSimThe benchmark of "LithoSim: A Large, Holistic Lithography Simulation Benchmark for AI-Driven Semiconductor Manufacturing"
The corresponding GitHub repo can be found at https://dw-hongquan.github.io/LithoSim/
Data Construction
4 in-distributed dataset (OPC_Metal/Metal/OPC_Via/Via).
1 out-of-distribution (OOD) dataset.
Each main dataset has a train_val and a test folder with compressed data file.
Each set of data contains a source_simple.src description of the source, a layout.png, a… See the full description on the dataset page: https://huggingface.co/datasets/s11ss/LithoSim.acc_rd_s1-gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/stewy33/acc_rd_s1-gpqa.multilingual-s1
Multilingual s1
A multilingual extension of the s1K-1.1 reasoning dataset. The original English reasoning questions and DeepSeek-R1 distilled solutions were translated into 9 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config.
We filter the upstream simplescaling/s1K-1.1 corpus to keep only samples whose DeepSeek-R1 trajectories were marked as correctly distilled, then translate the resulting subset.… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-s1.idt5-v4-results-final-lora-s123-20260912T013040606815Z
final-lora-s123-20260912T013040606815Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 94.7565543071161,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-lora-s123-20260912T013040606815Z.sauatai-ertegiler-kz-misspellings-kk-s170-len60-n6-m1-3-v1
SauatAI — Kazakh Misspelled Sentences from Ertegiler.kz
SauatAI is a grammar-focused dataset built from 170 children’s stories scraped from ertegiler.kz on July 5, 2025. The dataset was designed to support Kazakh language grammar correction, error detection, and text augmentation research.
📌 Dataset Details
s170 — 170 unique stories were scraped and sentence-tokenized.
len60 — Only sentences with ≤60 characters were retained.
n6 — Each correct sentence has 5… See the full description on the dataset page: https://huggingface.co/datasets/alphazhan/sauatai-ertegiler-kz-misspellings-kk-s170-len60-n6-m1-3-v1.demandpulseS1-BenchThe benchmark constructed in paper S1-Bench: A Simple Benchmark for Evaluating System 1 Thinking Capability of Large Reasoning Models.
Introduction
S1-Bench is a novel benchmark designed to evaluate Large Reasoning Models' performance on simple tasks that favor intuitive system 1 thinking rather than deliberative system 2 reasoning.
S1-Bench comprises 422 question-answer pairs across four major categories and 28 subcategories, balanced with 220 English and 202 Chinese questions.… See the full description on the dataset page: https://huggingface.co/datasets/WYRipple/S1-Bench.idt5-v4-results-final-fft-s123-20260910T215641650583Z
final-fft-s123-20260910T215641650583Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 93.25842696629213,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-fft-s123-20260910T215641650583Z.sauatai-ertegiler-kz-misspellings-kk-s170-len60-n6-mprob-v1
SauatAI — Kazakh Misspelled Sentences from Ertegiler.kz
SauatAI is a grammar-focused dataset built from 170 children’s stories scraped from ertegiler.kz on July 5, 2025. The dataset was designed to support Kazakh language grammar correction, error detection, and text augmentation research.
📌 Dataset Details
s170 — 170 unique stories were scraped and sentence-tokenized.
len60 — Only sentences with ≤60 characters were retained.
n6 — Each correct sentence has 5… See the full description on the dataset page: https://huggingface.co/datasets/alphazhan/sauatai-ertegiler-kz-misspellings-kk-s170-len60-n6-mprob-v1.mpl_s14_dataset
Dataset Card for MPL ID & PH Season 14 game Mobile Legends: Bang Bang
This is a match dataset from MPL ID & PH Season 14 from the Regular round to the Playoffs.
Dataset Sources
source for creating the dataset:
Liquipedia
MLBB Fandom
also with analytical discussion with my friends in mlbb_heroes_attribute.csv dataset in powerspike column (early, mid, late)
sauatai-ertegiler-kz-misspellings-kk-s170-len60-n6-m1-v1
SauatAI — Kazakh Misspelled Sentences from Ertegiler.kz
SauatAI is a grammar-focused dataset built from 170 children’s stories scraped from ertegiler.kz on July 5, 2025. The dataset was designed to support Kazakh language grammar correction, error detection, and text augmentation research.
📌 Dataset Details
s170 — 170 unique stories were scraped and sentence-tokenized.
len60 — Only sentences with ≤60 characters were retained.
n6 — Each correct sentence has 5… See the full description on the dataset page: https://huggingface.co/datasets/alphazhan/sauatai-ertegiler-kz-misspellings-kk-s170-len60-n6-m1-v1.s1k-freedomThis dataset is a merged and shuffled version of two separate datasets named FreedomIntelligence/Medical-R1-Distill-Data and simplescaling/s1K-1.1
Three columns from simplescaling/s1K-1.1 named "question", "deepseek_thinking_trajectory", "deepseek_attempt" were taken and renamed to "question", "resoning", "response"
ZS-train_S1-AURORA01_S2-SDGtitle_Negative_Sample_Filter-AURORA01huggingface_4073_tickets_s101465
Customer Support Tickets Batch
Source dataset for huggingface_4073 ticket triage task.
ai-vs-real-text-generated-datasetsauatai-ertegiler-kz-misspellings-kk-s170-len60-n6-m1-2-v1
SauatAI — Kazakh Misspelled Sentences from Ertegiler.kz
SauatAI is a grammar-focused dataset built from 170 children’s stories scraped from ertegiler.kz on July 5, 2025. The dataset was designed to support Kazakh language grammar correction, error detection, and text augmentation research.
📌 Dataset Details
s170 — 170 unique stories were scraped and sentence-tokenized.
len60 — Only sentences with ≤60 characters were retained.
n6 — Each correct sentence has 5… See the full description on the dataset page: https://huggingface.co/datasets/alphazhan/sauatai-ertegiler-kz-misspellings-kk-s170-len60-n6-m1-2-v1.ZS-train_SDG_Descriptions_S1-sentence_S2-SDGtitle_Negative_Sample_Filter-Only_Titlelabel_ids:
(0) contradiction
(2) entailment
ZS-train_SDG_Descriptions_S1-sentence_S2-SDGtitle_Negative_Sample_Filter-Only_Targetslabel_ids:
(0) contradiction
(2) entailment
ZS-train_SDG_Descriptions_S1-sentence_S2-subject_Negative_Sample_Filter-Only_HeadlinesZS-train_SDG_headlines_S1-headline_S2-titles_Negative_Sample_Filter-Only_Title_and_HeadlineZS-train_SDG_Descriptions_S1-sentence_S2-SDGtitle_Negative_Sample_Filter-SDG_DescriptionsZS-train_SDG_Descriptions_S1-sentence_S2-subject_Negative_Sample_Filter-Only_Title_and_HeadlineZS-train_SDG_Descriptions_S1-sentence_S2-SDGtitle_Negative_Sample_Filter-Alllabel_ids:
(0) contradiction
(2) entailment
s1_20kContains 20k uniform at random samples of Sentinel-1 SAR imagery with VV/VH polarizations as the two bands. Every image is a 256x256 patch.
All meta-data (such as location and image acquisition time) can be found in our index.csv file. The images themselves are found in s1_20k_images.tar.gz.
ZS-train_SDG_Descriptions_S1-sentence_S2-SDGtitle_Negative_Sample_Filter-Only_Indicatorslabel_ids:
(0) contradiction
(2) entailment
ZS-train_S1-SDG_Descriptions-AURORA01_S2-SDGtitle_Negative_Sample_Filter-SDG_DescriptionsZS-train_S1-SDGdescriptions-AURORA01_S2-SDGdescriptions-SDGtitle_Negative_Sample_Filter-AURORA01diffing-stats-spp_normal10_3b-spp10-L13-Crosscoder-s1-t100-k100-lr1e-04-x32LeslieKnope_S1E1_S2E8
