datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
firstpass-peer-review
FirstPass
FirstPass is a multi-domain, multi-round scientific peer-review dataset built from
Nature Communications transparent peer-review files. It contains 3,668 complete
peer-review dialogues across five scientific domains, with editorial outcome labels
derived from real revision-cycle outcomes.
The dataset was introduced alongside a fine-tuned model (Qwen2.5-7B-Instruct + LoRA)
that achieves 80.5% accuracy and F1-macro 78.2% on revision-cycle prediction
(STANDARD vs.… See the full description on the dataset page: https://huggingface.co/datasets/Prabhjotschugh/firstpass-peer-review.indo-bloom-corpus
🇮🇩 Indo-Bloom-AQG: A Unified Framework for Controllable Indonesian AQG
⚠️ RESEARCH ARTIFACT STATUS: SILVER VERSION (Work in Progress)
This dataset serves as the preliminary corpus (Silver Standard) for the ongoing Doctoral Dissertation at Universitas Negeri Malang (UM).
Current State: Unannotated / Pre-validation with Heuristic Bloom Labels
Target Final State: Gold Standard (Expert Validated with Bloom's Taxonomy Labels)
🔒 FROZEN — v0.1 Silver
This version is permanently… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/indo-bloom-corpus.LLMs-First-Task
Super easy task for humans that All SOTA LLM fail to retrieve the correct answer from context. Including SOTA models: GPT5, Grok4, DeepSeek, Gemini 2.5PRO, Mistral, Llama4...etc
Update: Accepted to COLM 2026 (San Francisco).
AAAI 2026 Worshop Oral: Jan/2026 LaMAS (LLM-based Multi-Agent Systems: Towards Responsible, Reliable, and Scalable Agentic Systems) Jan/2026 Singapole
ICML 2025 Long-Context Foundation Models Workshop Accepted.(https://arxiv.org/abs/2506.08184)
Update: This dataset… See the full description on the dataset page: https://huggingface.co/datasets/giantfish-fly/LLMs-First-Task.idea-first-code-later-cp
Idea First, Code Later: CP Benchmark
Benchmark dataset for the paper: "Idea First, Code Later: Disentangling Problem Solving from Code Generation in Evaluating LLMs for Competitive Programming"
A curated benchmark of 83 competitive programming problems designed for evaluating LLMs on algorithmic problem-solving separately from code generation.
Motivation
We curate problems from seven contests that are not hosted on major public CP platforms (e.g., Codeforces, AtCoder).… See the full description on the dataset page: https://huggingface.co/datasets/samahadhoud/idea-first-code-later-cp.IndoBloom-AQG-Benchmark-Corpus
📚 Indo-Bloom AQG Benchmark Corpus (All Models)
⚠️ RESEARCH ARTIFACT STATUS: BENCHMARK / SILVER CORPUS (Stage 1)
This dataset serves as the comparative benchmark corpus for the Indo-Bloom research project at Universitas Negeri Malang (UM).
Current State: LLM-Generated QA Pairs Evaluated via Rule-Based Evaluator
Next Stage: Expert Annotation (Stage 2) → Gold Standard
🔒 FROZEN — Benchmark v1.0
This version is permanently frozen to ensure reproducibility of the experimental… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/IndoBloom-AQG-Benchmark-Corpus.indo-bloom-raw-bse
📚 Indo-Bloom BSE RAW Corpus
⚠️ RESEARCH ARTIFACT STATUS: RAW CORPUS (Stage 0)
This dataset serves as the raw material corpus for the Indo-Bloom research project at Universitas Negeri Malang (UM).
Current State: Extracted & Cleaned Context from BSE Textbooks
Next Stage: QA Pair Generation (Stage 1) → Silver Corpus
🔒 FROZEN — Raw v1.0
This version is permanently frozen to ensure reproducibility. This corpus will be used as input for QA generation pipeline.
📄 Associated… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/indo-bloom-raw-bse.
