datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
toxicity_classification_jigsaw
Dataset info
Training Dataset:
You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are:
toxic
severe_toxic
obscene
threat
insult
identity_hate
The original dataset can be found here: jigsaw_toxic_classification
Our training dataset is a sampled version from the original dataset, containing equal number of samples for both clean and toxic classes.
Dataset creation:… See the full description on the dataset page: https://huggingface.co/datasets/Arsive/toxicity_classification_jigsaw.ArSAScontext-ucurve-coding-agents
Context U-curve: 36 coding-agent runs under six context-clearing policies
How often should an LLM coding agent's context be cleared? This dataset holds every run behind the report
"Clear Every Third Task: A Measured U-Curve in the Context Economy of Coding Agents"
(Evgenii Arsentev, 2026; corrected version 1.2, DOI 10.5281/zenodo.22759217; version 1.0: DOI 10.5281/zenodo.22699668).
A fixed suite of twelve programming tasks was run under six session-length policies — a fresh… See the full description on the dataset page: https://huggingface.co/datasets/arsentev-ai/context-ucurve-coding-agents.qwen-base-error-analysis
Qwen3.5-4B-Base Error Analysis Dataset
This dataset contains 10 diverse examples where the base language model Qwen3.5-4B-Base makes incorrect or unexpected predictions. It was created as part of an exploration of base model blind spots and failure modes.
Model Tested
Model: Qwen/Qwen3.5-4B-Base
Type: Pre-trained base model (not instruction-tuned)
Parameters: 4B
Architecture: Causal Language Model with hybrid Gated DeltaNet + Attention layers
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/ArslanMZahid/qwen-base-error-analysis.youtube-sentiment-dataset
YouTube Comments Sentiment Dataset
A comprehensive, large-scale dataset featuring 1,032,225 English YouTube comments, curated and labeled for 3-class sentiment analysis (Negative, Neutral, and Positive). This dataset is optimized for training, evaluating, and fine-tuning Transformer-based NLP models and sentence encoders.
🔗 Related Resources
Hugging Face Dataset: Arshia82sbn/youtube-sentiment-dataset
Hugging Face Model:… See the full description on the dataset page: https://huggingface.co/datasets/Arshia82sbn/youtube-sentiment-dataset.medquad_scraped_contextArSRED
AREEj: Arabic Relation Extraction with Evidence
This dataset was made by adding evidence annotations to the Arabic subset of SREDFM. The dataset is from the Proceedings of The Second Arabic Natural Language Processing Conference paper AREEj: Arabic Relation Extraction with Evidence. If you use the dataset or the model, please reference this work in your paper:
@inproceedings{mraikhat-etal-2024-areej,
title = "{AREE}j: {A}rabic Relation Extraction with Evidence",
author =… See the full description on the dataset page: https://huggingface.co/datasets/U4RASD/ArSRED.Iran_GoldArSASLMentalBert-1weathervit-datasetMentalBert-3NE_ArSASweather-dataset-indiallm-reasoning-evaluation-examples
LLM Reasoning Evaluation Examples
Overview
This dataset contains 20 prompts designed to evaluate potential blind spots in a base language model.It covers 10 reasoning categories: Logic, Arithmetic, Factual, Language/Translation, Pattern/Analogy, Commonsense, Comparison, Negation/Uncertainty, Math Sequence, and Analogies.
The dataset is intended for educational and research purposes to illustrate where a base model may produce outputs that differ from expected answers.… See the full description on the dataset page: https://huggingface.co/datasets/hareem-arshad/llm-reasoning-evaluation-examples.KargeniaGeorgian-Parallel-Corpora
English-Georgian Parallel Dataset 🇬🇧🇬🇪
📄 Dataset Overview
The English-Georgian Parallel Dataset is sourced from OPUS, a widely used open collection of parallel corpora. This dataset contains aligned sentence pairs in English and Georgian, extracted from Wikipedia translations.
Corpus Name: Wikimedia
Package: wikimedia.en-ka (Moses format)
Publisher: OPUS (Open Parallel Corpus)
Release: v20230407
Release Date: April 13, 2023
License: CC–BY-SA 4.0
🔗… See the full description on the dataset page: https://huggingface.co/datasets/Arseniy-Sandalov/Georgian-Parallel-Corpora.MentalBert-2MentalBert-4TestDatasetSchool_qaARSpeechGeorgian-Sentiment-Analysis
Sentiment Analysis for Georgian 🇬🇪
📄 Dataset Overview
The Sentiment Analysis for Georgian dataset was created by the European Commission, Joint Research Centre (JRC). It is designed for training and evaluating sentiment analysis models in the Georgian language.
Authors: Nicolas Stefanovitch, Jakub Piskorski, Sopho Kharazi
Publisher: European Commission, Joint Research Centre (JRC)
Year: 2023
Dataset Type: Sentiment Analysis
License: Check official source… See the full description on the dataset page: https://huggingface.co/datasets/Arseniy-Sandalov/Georgian-Sentiment-Analysis.Meezan_Bank_Ten_Year_Stockspell_correction_HW
