datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi-wiki-qa-high-quality-subset
multi-wiki-qa-high-quality-subset
A quality-filtered subset of the Danish (da) split of
alexandrainst/multi-wiki-qa,
a Wikipedia-based extractive question-answering dataset.
Configs
Config
Samples
Description
da
4,767
All LLM-verified correct samples
da-short
3,527
Correct samples where the answer is at most 3 words
Filtering methodology
Starting from the 5,000 samples in the original Danish split:
Span validation -- deterministic check that… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/multi-wiki-qa-high-quality-subset.ifc-bim-high-quality-alpaca
IFC BIM High-Quality Dataset (Alpaca Format)
Dataset Description
This is a high-quality, curated dataset for training language models on IFC (Industry Foundation Classes) and BIM (Building Information Modeling) tasks. The dataset has been filtered for quality and is provided in the Alpaca instruction-following format.
Dataset Summary
Total entries: 42,680
Format: Alpaca (instruction, input, output)
Language: English
Domain: IFC/BIM technical documentation and… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-high-quality-alpaca.korean-quality-cleaned
Korean Quality Dataset (Cleaned)
고품질 한국어 Instruction 데이터셋 (정제 버전)
English
Dataset Description
This is a cleaned and standardized Korean instruction dataset, combining multiple high-quality open-source Korean datasets with unified formatting and quality filtering.
Key Features
✅ Unified Format: Standardized messages format (OpenAI-compatible)
✅ Quality Filtering: Length, special characters, repetition filtering
✅ Clean Structure: Removed redundant… See the full description on the dataset page: https://huggingface.co/datasets/MyeongHo0621/korean-quality-cleaned.echidna-round5-response-quality
echidna-round5-response-quality
Echidna — round 5 response-quality training examples.
Contents
round5_response_quality.jsonl (13 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Echidna RAG assistant (Michael Anthony Falabella).
quality-search-dataset
QuALITY Search Dataset
This dataset is derived from the QuALITY (Question Answering with Long Input Texts, Yes!) dataset, specifically designed for search-based question answering environments.
Dataset Overview
Source: QuALITY v1.0.1 dev set
Samples: 453 question-article pairs from 50 articles
Task: Multiple-choice question answering with search capabilities
Embeddings: OpenAI text-embedding-3-small (1536 dimensions)
Quick Start
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/bhogan/quality-search-dataset.High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning
PyReason-7k: Advanced Python Chain-of-Thought Dataset
Dataset Description
This dataset contains 7,000+ high-quality Python programming examples designed for LLM fine-tuning.
Each entry includes a detailed thought_process (Chain-of-Thought) to teach models logical reasoning before coding.
Key Features:
Chain-of-Thought: Step-by-step reasoning traces.
Error Handling: Solutions include try-except blocks and logging.
Diverse Tasks: Algorithms, API handling, Data Structures.… See the full description on the dataset page: https://huggingface.co/datasets/xTayyub/High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning.SuperDataset-QualityRelease
SuperDataset
1. Introduction
SuperDataset represents a new generation of curated training data for natural language processing tasks. Our latest release incorporates advanced data curation techniques including automated quality filtering, cross-validation with multiple annotators, and comprehensive bias detection. The dataset demonstrates exceptional quality metrics across all evaluation dimensions.
Compared to the previous version, the… See the full description on the dataset page: https://huggingface.co/datasets/toolevalxm/SuperDataset-QualityRelease.SuperDataset-QualityTest
SuperDataset
1. Introduction
SuperDataset is a comprehensive, high-quality dataset designed for natural language processing tasks. This latest version includes significant improvements in data quality, coverage, and annotation accuracy. The dataset has been curated using state-of-the-art data validation pipelines and human verification processes.
Compared to previous versions, this release features enhanced data cleaning procedures… See the full description on the dataset page: https://huggingface.co/datasets/toolevalxm/SuperDataset-QualityTest.sophistry-bench-quality-dev
Sophistry-Bench QuALITY Dev Slice
A 50-item curated subset of the QuALITY
multiple-choice reading-comprehension dev set, used as the evaluation
distribution for sophistry-bench —
an asymmetric-information debate RL environment reproducing the protocol from
Khan et al. 2024 (Debating with More Persuasive LLMs Leads to More Truthful
Answers).
What this slice is for
Sophistry-Bench debates run two LLMs (one defending the gold answer, one
defending a distractor) over a… See the full description on the dataset page: https://huggingface.co/datasets/anushaacharya/sophistry-bench-quality-dev.Elite_quality_cybersecuritydevelopers-high-quality-mozgach
developers-high-quality-mozgach
Описание
Высококачественные примеры для разработчиков, сгенерированные mozgach108.
Датасет содержит отборные примеры для различных задач программирования:
Написание кода
Отладка
Рефакторинг
Архитектурные решения
Code review
Тестирование
Особенность: высокое качество ответов, сгенерированных специализированной моделью mozgach108.
Сгенерировано через Ollama (mozgach108:latest).
Статистика
Всего примеров: 1200… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/developers-high-quality-mozgach.dickens_data_quality_checks
