datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
video-quality-scored
Image-to-Video Quality-Scored Clips
A collection of prompted image-to-video samples with quality-evaluation metadata.
Each sample pairs a first frame (the I2V conditioning image) with one or both
of:
a generated video produced by a video model from the first frame + prompt
an original clip (the reference/source video the prompt was authored around)
A subset of the samples also carry per-clip quality scores: an overall
quality_score, six per-aspect breakdowns… See the full description on the dataset page: https://huggingface.co/datasets/mohantesting/video-quality-scored.prompt-quality
Prompt Quality Assessment
Prompt quality strongly affects how well large language models (LLMs) perform, especially when user inputs are vague or incomplete. A good prompt is clear, specific, and complete, giving the model enough relevant context to produce accurate and useful responses.
This report describes a dataset created by evaluating prompts with several different LLMs. These evaluations can be used to train prompt-quality classifiers and to improve methods for prompt… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-quality.polymarket-settlement-quality-register
Polymarket settlement-quality register
Frozen summaries of 123,499 settled UMA requests, window 2023-12-05 to 2026-08-12. Among settled disputes, 7.12% changed the proposal. Group summaries cover category and rule-text features.
Files and viewer
The viewer loads the canonical aggregate snapshot only. The dated files preserve export history and are not independent observations.
Method and source
See the embedded metadata and repository inventory.… See the full description on the dataset page: https://huggingface.co/datasets/ailinsun/polymarket-settlement-quality-register.backln-guest-post-quality-public-mirror
Backln Guest Post Quality Public Mirror
Public-safe mirror for validating Hugging Face Dataset Viewer indexing and release gates. This dataset is not the private training corpus.
Full text, titles, and snippets are removed by default. The mirror keeps labels, coarse metadata, feature buckets, and hash prefixes so the public Hub can verify schema and distribution without exposing customer content.
Schema
label: one of published, manual_review, rejected.
source: coarse… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/backln-guest-post-quality-public-mirror.translation-quality
Multilingual Translation Quality Dataset
This dataset provides multilingual text chunks translated into English, accompanied by automated quality evaluations generated by multiple large language models.
Dataset Details
Source Data: agentlans/HuggingFaceFW-finetranslations-100-languages-sample
Target Language: English
Content: Multilingual chunks mapped to their English translations alongside automated judge scores.
Evaluation Methodology
The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/translation-quality.high-quality-crash-courserephrased-web-data-quality-study
Rephrased Web Data Quality Study
LLM-as-judge evaluation of ~4,000 examples from HuggingFaceFW/finephrase (1,000 sampled per split, 86 dropped due to judge parse failures, 3,914 successfully evaluated).
Judge: Claude Sonnet 4.6 via OpenRouter | Cost: ~$45
Quality Scores (1-5 scale)
Metric
FAQ (n=965)
Table (n=979)
Tutorial (n=976)
Math (n=994)
Faithfulness
1.82
1.72
1.90
1.49
Info preservation
1.93
1.64
1.99
1.47
Appropriateness
3.54
2.87
2.48
1.67… See the full description on the dataset page: https://huggingface.co/datasets/ratishsp/rephrased-web-data-quality-study.bwb-quality-scores
BWB Quality Scores (QE + Arena)
中文说明
Quality scores for 600,000 Chinese→English sentence pairs sampled from the train split of the BWB bilingual web-novel corpus, produced with a locally deployed Qwen3.8-27B judge.
This repository contains scores only — no original text. Each record is keyed by a positional index (book, ch, sn) plus a sha1 fingerprint of the normalized text, so anyone who has obtained the official BWB release can re-attach the scores to the text losslessly and… See the full description on the dataset page: https://huggingface.co/datasets/umeiko/bwb-quality-scores.model-quality-release-gate
Model Quality Release Gate Evaluation Dataset
Reproducible evaluation evidence for comparing baseline and candidate AI code-generation models before release.
Phase 3 introduces explicit benchmark versioning so release evidence can identify exactly which dataset definition produced a decision.
Versioned benchmark
Current benchmark release:
Name: CodeBench-Safety
Version: 1.0.0
Manifest: versions/v1.0.0/manifest.json
Cases: versions/v1.0.0/cases.jsonl
Compatible… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/model-quality-release-gate.review_qualityQualityVision-walking-sample
QualityVision Walking Sample (compact)
Need production-scale pose data? Browse ready-made JSONL bundles (thousands of HQ frames, full manifests, schema docs) and Dataset Lab plans on qvision.space — Dataset pricing. This Hub repo is a small non-commercial sample so you can validate parsing and quality before you buy.
Small high-quality 2D pose clip for walking, exported from the Quality Vision Motion Dataset Engine pipeline: HQ frame filtering, optional temporal smoothing on… See the full description on the dataset page: https://huggingface.co/datasets/Alaaharoun/QualityVision-walking-sample.repro-ski-rental-with-distributional-predictions-of-unknown-quality-traces
Agent traces
Agent sessions published from a Trackio Logbook.
argilla_ultrafeedback-multi-binarized-quality-preferences-cleaned-PreferenceShareGPTazerbaijani-text-quality-labeled
Azerbaijani Text Quality — Labeled Dataset
249,949 Azerbaijani web documents annotated with a quality score 0-3.
Used to train a document-level quality classifier for filtering a web
corpus before language-model pretraining.
Source and labeling
Texts: sampled from LocalDoc/community_oscar_azerbaijani,
an OSCAR-derived Common Crawl corpus. The texts are NOT original to this dataset.
Labels: generated by the LLM Mistral-Small-24B-Instruct-2501, not by humans.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-text-quality-labeled.code-quality-corpus
CatQualia code-quality corpus — semantic smell classes with before/after fixes
39,383 rows · 25,541,792 bytes · JSON Lines, one object per line.
What this is
Real code smells paired with the fix: a smell_class that names the semantic problem (not just the syntax), the original lines, the corrected lines, the file and line it came from, and a rationale explaining why the original was wrong. Useful for code-review or repair training where the label has to say what… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/code-quality-corpus.chat-quality
Chat Quality
Collection of conversations evaluated using Qwen 3 series.
Prompt template:
You are an AI evaluator tasked with rating the overall quality of a complete human–AI conversation (all user messages and AI responses) on a 1–10 scale based on how effectively it serves the user’s stated and implied goals.
<conversation>
[CONVERSATION]
</conversation>
Evaluate the above conversation as a whole, considering:
* Understanding and fulfillment of the user’s goals
* Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/chat-quality.QualityVision-Locomotion-Pose-Dataset-Walking-Jogging-Running
QualityVision Locomotion Pose Dataset (Walking + Jogging + Running) — Sample
This is a compact, viewer-friendly sample extracted from a much larger HQ locomotion export generated by the QualityVision Motion Dataset Engine.
Looking for the full commercial export or custom delivery? See pricing & ready-made bundles on qvision.space/dataset-pricing.
What’s inside
data.jsonl: one JSON object per line (one frame per row) with 33 MediaPipe/BlazePose landmarks (x,y,z… See the full description on the dataset page: https://huggingface.co/datasets/Alaaharoun/QualityVision-Locomotion-Pose-Dataset-Walking-Jogging-Running.QualityVision-Jogging-Pose-Dataset-61-Videos-14550-Frames
QualityVision Jogging Pose Dataset (61 videos, 14,550 frames) — Sample
This Hugging Face dataset is a compact sample extracted from the full QualityVision Jogging Pose export.
Action label: jogging
Keypoints: 33 landmarks per person (MediaPipe / BlazePose) with x, y, z, visibility
Post-processing (as exported): temporal smoothing + body normalization flags are included in metadata
Use this sample to validate the schema and quality before purchasing larger exports.
Pricing &… See the full description on the dataset page: https://huggingface.co/datasets/Alaaharoun/QualityVision-Jogging-Pose-Dataset-61-Videos-14550-Frames.reddit-political-discourse-qualityqualitycode-quality-poor
Low Quality Code Dataset
Overview
This dataset contains 444 Python code samples with poor quality scores (62-74 out of 100).
These samples can be used for:
Training models to recognize bad code patterns
Contrastive learning (good vs bad code)
Code quality classification tasks
Statistics
Metric
Value
Total samples
444
Quality range
62-74
Average quality
~70
Source Distribution
Source
Count
BigOBench
425… See the full description on the dataset page: https://huggingface.co/datasets/happylife365/code-quality-poor.
