datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aya_evaluation_suite
Dataset Summary
Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes:
human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated.
machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.ropedia-xperience-10m-task-suite-artifacts
Ropedia Xperience-10M Task Suite Artifacts
This dataset repository stores small derived artifacts for the Ropedia
Xperience-10M task-suite project: metrics, predictions, manifests, reports,
figures, website JSON, public-safe Qwen3-Omni diagnostic outputs, and the
Cosmos3-Nano plus Cosmos3-Super diagnostic packages.
Project Identity
The Project identity mark is shared across the GitHub README, GitHub Pages
dashboard, Hugging Face Space, artifact dataset, model… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/ropedia-xperience-10m-task-suite-artifacts.bybit-linear-perps-suiusdtBellaTurca
Dataset Card for BellaTurca
BellaTurca is the first large-scale Turkish corpus collection for training Turkish language models. The total size is around 245GB and 30 billion words. BellaTurca's focus is high quality, diversity as well as the size.
This collection is made up of five datasets: AkademikDerlem, OzenliDerlem, ForumSohbetleri, Temiz OSCAR and Temiz mC4. Originally there was a book corpus included, but it is excluded due to containing copyrighted material.
AkademikDerlem… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/BellaTurca.storiesctc-suite-eval
CTC suite eval ladders
The 22-task corpus-tracking-capacity suite: per-task context ladders from 2k to 1M tokens,
consumed by the ctc_suite task family on the prasann/ctc-suite branch of allenai/olmo-eval
(ctc_nq:r64k, suites ctc:figure / ctc:xlong / ctc:r128k / ...). One config per task, one
split per rung; each row is one unified-format example (documents + queries + answers + gold).
Public release note (2026-08-14). Gold answers are included — training on this data… See the full description on the dataset page: https://huggingface.co/datasets/PrasannSinghal/ctc-suite-eval.gemma-2b-suite-explanationsMultiViewBench
MultiView-Bench
Evaluation data for MultiView-Bench: A Diagnostic Benchmark for
World-Centric Multi-View Integration in VLMs
(arXiv:2607.08970).
This repository currently contains the synthetic image-and-text portion of the
benchmark. The real-world subset is not included; see the scope note below.
Contents
2,400 evaluation samples across 24 benchmark variants
7,900 PNG views (2.94 GiB logical image data)
One, three, or six images per sample
English question text… See the full description on the dataset page: https://huggingface.co/datasets/suijinru/MultiViewBench.OzenliDerlem
Dataset Card for OzenliDerlem
OzenliDerlem (a.k.a CraftedCrawl) is a carefully assembled collection of top-notch web crawl data from handpicked websites, featuring articles, journals, and magazines. It focuses on gathering rich and detailed text content, especially longer articles. The collection covers a wide range of topics, including travel, news, culture, fairy tales and folklore, movie reviews, popular science, product and service complaints, fashion and self-care, trendy… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/OzenliDerlem.SUITE
SUITE
SUITE (Selective Unlearning of Isolated Topics and Events) is a fine-grained benchmark for
machine unlearning in LLMs: removing a specific set of facts from a model while
preserving everything else. It is the benchmark introduced in the paper
Forget Narrowly, Retain Broadly: Unlearning as an Asymmetric Generalization Problem
(see Citation).
We frame unlearning as an asymmetric generalization problem. Forgetting must generalize
intensively: it has to hold across every… See the full description on the dataset page: https://huggingface.co/datasets/apeleg/SUITE.tabula-8b-eval-suiteEvaluation suite used in our paper "Large Scale Transfer Learning for Tabular Data via Language Modeling."
This suite includes our preprocessed versions of benchmark datasets except the AutoML Multimodal Benchmark, which can be accessed by following the installation instructions in their repo here.
We recommend using rtfm when evaluating models with these datasets.
See the rtfm repo for more information on using this data for evaluation.
bea-eval-suiteTrGLUE
TrGLUE - The First Non-Translate Natural Language Understanding Benchmark for Turkish
Dataset Card for TrGLUE
TrGLUE is a natural language understanding benchmarking dataset including several single sentence and sentence pair classification tasks.
The inspiration is clearly the original GLUE benchmark.
Tasks
Single Sentence Tasks
TrCOLA The original Corpus of Linguistic Acceptability consists of sentences compiled from English literature textbooks.… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/TrGLUE.Detection-for-Suicideqwen38-27b-fidelity-suite-v5
Qwen3.8-27B fidelity suite v5 — evaluation inputs, the BF16 reference, and every per-shard report
This dataset exists because we deleted expensive artifacts once and had to remake them. Every
tree here is replayable input for a future candidate, not a finished result — the finished
results live as receipts in the research repo. Publishing the inputs means the next candidate
costs one download instead of a fresh BF16 capture, and it means anyone can check our numbers
without our… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen38-27b-fidelity-suite-v5.suicide_depression_detectiongemma-2b-suite-explanations-residualtemiz-OSCAR
Dataset Card for Temiz OSCAR
Temiz OSCAR is a corpora collection consisting of cleaned versions of original OSCAR corpora.
This collection is made up of four datasets: OSCAR-2019, OSCAR-2109, OSCAR-2201 and OSCAR-2301
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
Dataset
num instances
size
num of words
OSCAR-2019
3.671.430
7.7G
976M
OSCAR-2109
8.472.809
18G
2.22B
OSCAR-2201… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-OSCAR.NLP_SUITEsuicide_prediction_dataset_phr
Dataset Card for "vibhorag101/suicide_prediction_dataset_phr"
The dataset contains text with binary labels for suicide or non-suicide.
The dataset was cleaned and following steps were applied
Converted to lowercase
Removed numbers and special characters.
Removed URLs, Emojis and accented characters.
Removed any word contractions.
Remove any extra white spaces and any extra spaces after a single space.
Removed any consecutive characters repeated more than 3 times.
Tokenised the… See the full description on the dataset page: https://huggingface.co/datasets/vibhorag101/suicide_prediction_dataset_phr.finqa_suitegemma-2b-suite-maxacts-attn_out
SUITE-rephrasings
SUITE: rephrasings
This is the robustness companion to apeleg/SUITE,
the benchmark introduced in the paper
Forget Narrowly, Retain Broadly: Unlearning as an Asymmetric Generalization Problem (see Citation).
It holds the SUITE forget-evaluation questions together with the paraphrased rewordings of each
one, produced by a held-out generator (Gemini) that is not used to build the training
augmentations.
It is kept as a separate repo because its schema differs from the core… See the full description on the dataset page: https://huggingface.co/datasets/apeleg/SUITE-rephrasings.Indic-Rag-Suite
🌏 Multilingual Indic RAG Suite
A comprehensive multilingual question-answering dataset covering 18 Indian languages with 21,439,886 total samples, designed for RAG (Retrieval-Augmented Generation) applications and multilingual NLP research.
🚀 Quick Start
from datasets import load_dataset
# Load specific language (recommended)
dataset = load_dataset("ai4bharat/Indic-Rag-Suite", "as")
train_data = dataset['train']
print(f"Loaded {len(train_data)} samples")
# Access… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Indic-Rag-Suite.gemma-2b-suite-maxacts-residual
CSSR-S_labelled_suicidewatch_posts_reddit
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
Full code and supplementary materials are available at https://github.com/av9ash/llm_cssrs_code.
License and Citation
This project is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.Any use or reuse of this work please cite the following:
@article{patil2025evaluating,
title={Evaluating Reasoning LLMs for Suicide Screening with the… See the full description on the dataset page: https://huggingface.co/datasets/av9ash/CSSR-S_labelled_suicidewatch_posts_reddit.suicidePredicting-Experts-for-MOE-Suite
Predicting Experts for MoE: A Coverage-First Routing Benchmark
Goal: predict the complete set of experts that a future
Mixture-of-Experts layer will activate, using only information that is
causally available before that layer executes.
This dataset turns expert prefetch prediction into a standalone machine-learning
problem. It contains 98,292 routed generated tokens and 7,371,900 ordered
expert-route labels from 12 synthetic, realistic coding tasks evaluated with… See the full description on the dataset page: https://huggingface.co/datasets/mistrjirka/Predicting-Experts-for-MOE-Suite.Havadis
Dataset Card for Havadis
Havadis is a high quality and large Turkish news corpus, indeed the largest Turkish news corpus ever.
This corpus is scraped from online news sebsites and includes text from popular newspapers such as
CNN Türk
Habertürk
Hürriyet
Millyet
NTV
Posta
Sabah
Star
Sözcü
Takvim
. The instances are first crawled from the corresponding websites, then went throught an extensive cleaning process. We eliminated instances that are too short, too repetetive… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Havadis.juliet_test_suite_c_1_3
Dataset Card for the Juliet Test Suite 1.3
Dataset Summary
This Datasets contains all test cases from the NIST's Juliet test suite for the C and C++ programming languages. The dataset contains a benign and a defective implementation of each sample, which have been extracting by means of the OMITGOOD and OMITBAD preprocessor macros of the Juliet test suite.
Supported Tasks and Leaderboards
Software defect prediction, code clone detection.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/LorenzH/juliet_test_suite_c_1_3.
