datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PRISM
PRISM Alignment
Task-grouped multidimensional dialogue-quality data from PRISM Alignment.
Contents
The release contains 6,187 complete artifact rows from 6,187 conversation groups and 1,309 participants.
The original release contains 8,011 conversations; 1,824 are excluded because one or more of the seven performance sliders is missing or invalid, or because the selected first-turn response is unavailable.
The targets are values, fluency, factuality, safety… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/PRISM.MSumBench
MSumBench
Task-grouped multidimensional summarization quality data from MSumBench.
Contents
The release contains 2,250 complete artifact rows from 150 source-document groups.
The three prediction targets are faithfulness, completeness, and conciseness.
Split organization
Each seed has task-grouped train, validation, and test splits. All language-specific summaries and model generations for one group_id remain in exactly one split. The seeds change… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/MSumBench.RATE
WMT24 English-Russian RATE
Task-grouped multidimensional machine-translation quality data from RATE.
Contents
The release contains 3,975 complete artifact rows from 497 source-segment groups.
One curated row was removed from the 3,976-row release because its translation was blank and its three scores were zero. The targets are score_accuracy, score_fluency, and score_style, each on a 0--100 scale.
Split organization
Each seed has task-grouped train… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/RATE.CreativeEval
CreativeEval
Paper-grouped multidimensional research-ideation evaluation data from CreativeEval.
Contents
The release contains 1,026 complete paper rows. Each row contains a human-written research-paper introduction, the raw reviewer score arrays for provenance, and four mean prediction targets: contribution_mean, soundness_mean, presentation_mean, and overall_score_mean.
All four targets are derived from the human reviewer scores released with the paper. The raw… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/CreativeEval.netryx-new-york-5km
New York 5km
Pre-computed MegaLoc index for Netryx Drishti geolocation.
Coverage
Center: 40.712800, -74.006000
Radius: 5.0 km
Panoramas: 196,824
Index entries: 787,296
Descriptor model: MegaLoc
Descriptor dim: 1024 (PCA from 8448)
Usage
from netryx_hub import NetryxHub
hub = NetryxHub()
hub.download("new-york-5km", output_dir="./netryx_data/index")
# Now open Netryx and search!
Or download manually and use Import Index in the Netryx GUI.
Details… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-new-york-5km.netryx-new-york-city-13km
Nyc-Core-Usethis 13km
Pre-computed MegaLoc index for Netryx Drishti geolocation.
Coverage
Center: 40.713200, -74.002500
Radius: 13.0 km
Panoramas: 663,084
Index entries: 2,652,336
Descriptor model: MegaLoc
Descriptor dim: 1024 (PCA from 8448)
Usage
from netryx_hub import NetryxHub
hub = NetryxHub()
hub.download("nyc-core-usethis-13km", output_dir="./netryx_data/index")
# Now open Netryx and search!
Or download manually and use Import Index in… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-new-york-city-13km.netryx-delhi-mixvpr-5km
Delhi 5km (MixVPR)
Pre-computed MegaLoc index for Netryx Drishti geolocation.
Coverage
Center: 28.644800, 77.216721
Radius: 5.0 km
Panoramas: 69,838
Index entries: 279,352
Descriptor model: MixVPR
Descriptor dim: 512 (PCA from 512)
Usage
from netryx_hub import NetryxHub
hub = NetryxHub()
hub.download("delhi-5km-(mixvpr)", output_dir="./netryx_data/index")
# Now open Netryx and search!
Or download manually and use Import Index in the Netryx GUI.… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-delhi-mixvpr-5km.adaption-defi-wallet-risk-classification
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-defi_wallet_risk_classification
This dataset contains prompt-completion pairs for classifying the 14-day risk outcomes of DeFi wallets on various EVM networks. Each entry provides behavioral features such as transaction counts, action ratios, and concentration metrics within a specific feature window to predict a binary risk label. The completions offer a concise justification for… See the full description on the dataset page: https://huggingface.co/datasets/samscript18/adaption-defi-wallet-risk-classification.solidity-audit-cot
solidity-audit-cot
Long-CoT audit traces for Solidity contracts, generated by Claude Opus 4.7 (adaptive thinking, xhigh effort) over the spec→contract corpus from the Qwopus3.6-27B-solidity training pipeline.
This dataset is the Stage 2 training corpus for the multi-stage Qwopus3.6-27B-solidity model — designed to teach long-form security reasoning (8-15 paragraph chain-of-thought) anchored to real Solidity contracts.
Why this dataset exists
Public Solidity audit… See the full description on the dataset page: https://huggingface.co/datasets/samscrack/solidity-audit-cot.postmortems-exploitssamsum-fa
Dataset Summary
SAMSum-Fa is a Persian (Farsi) dataset designed for the Summary Retrieval task. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and is a translated version of the original English SAMSum dataset, which contains dialogue-summary pairs. In FaMTEB, SAMSum-Fa is used to benchmark how well models can retrieve a human-written summary corresponding to a Persian chat conversation.
Language(s): Persian (Farsi)
Task(s): Summary Retrieval
Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/samsum-fa.samsum
Samsum Dataset (Reorganized for Easy Access)
Link to the original dataset: Samsum Dataset on Hugging Face
This is a reorganized version of the Samsum Dataset, originally created and published by Gliwa, Bogdan, et al. The dataset has been structured for easier access, but the original content remains unchanged.
Dataset Information
Original Authors: Gliwa, Bogdan, et al.
Paper: Samsum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization… See the full description on the dataset page: https://huggingface.co/datasets/nyamuda/samsum.SAMSUM-eu
SAMSUM-eu dataset for Summarization in Basque.
SAMSUM-eu was created by automatically translating SAMSum, a human-annotated dialogue dataset for abstractive summarization, using a proprietary document-level MT system based on Llama-eus-8B. We then filtered out examples with incomplete translations
or non-Basque outputs. The translated test set was further refined by a native speaker to obtain 100 high-quality, manually curated test examples. In total, we obtained 11,313
training… See the full description on the dataset page: https://huggingface.co/datasets/orai-nlp/SAMSUM-eu.APPS
APPS
This dataset is the MO-RELISH APPS code-generation quality-estimation
collection derived from akhauriyash/Code-Regression. Each row pairs an APPS
programming problem with a Python candidate solution and execution-derived
measurements.
Configs And Splits
Config
Train
Validation
Test
full
78,065
9,758
9,758
subset_10k
10,000
1,000
1,000
The full config preserves the deterministic APPS splits from the local
multi-output curation. The subset_10k… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/APPS.GraphArch
GraphArch
This dataset is the MO-RELISH graph architecture regression collection. Each
row contains serialized neural-network graph text and execution/benchmark
measurements from neural architecture search spaces.
Configs And Splits
Config
Train
Validation
Test
full
475,257
10,000
10,000
subset_10k
10,000
1,000
1,000
The full config keeps all strict-clean rows, with 10,000 validation rows and
10,000 test rows sampled deterministically and… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/GraphArch.SummEval
SummEval
SummEval contains generated CNN/DailyMail summaries with expert human judgments on coherence, consistency, fluency, and relevance.
This dataset is staged for MO-RELISH as single-file JSONL splits on Hugging
Face. All dimensions in targets are prediction targets.
Configs And Splits
Config
Train
Validation
Test
default
1,120
160
320
Splits are grouped by source document to avoid putting summaries for the same
source document in different… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/SummEval.UniSumEval
UniSumEval
UniSumEval contains generated summaries with normalized quality labels derived from key-fact coverage and factual-support annotations.
This dataset is staged for MO-RELISH as single-file JSONL splits on Hugging
Face. All dimensions in targets are prediction targets.
Configs And Splits
Config
Train
Validation
Test
en
1,291
179
358
Splits are grouped by source document to avoid putting summaries for the same
source document in different… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/UniSumEval.samsum
Samsum Dataset (Reorganized for Easy Access)
Link to the original dataset: Samsum Dataset on Hugging Face
This is a reorganized version of the Samsum Dataset, originally created and published by Gliwa, Bogdan, et al. The dataset has been structured for easier access, but the original content remains unchanged.
Dataset Information
Original Authors: Gliwa, Bogdan, et al.
Paper: Samsum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization… See the full description on the dataset page: https://huggingface.co/datasets/fastlayne/samsum.samsum-thinkingsamsum-llamafiedreact_single_componentscyfrin-audit-findingsEmilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour… See the full description on the dataset page: https://huggingface.co/datasets/samson-ailabs/Emilia-Dataset.sam-style-rolesreact_oss_instructreact_two_componentsSystemC-TrainingSet-1StatLLM
StatLLM
MO-RELISH curation of StatLLM for multi-output regression over generated SAS
programs.
Splits
train: 321
validation: 99
test: 201
Splits are grouped by statistical task id. The source dev split is mapped to
validation.
Prediction Targets
statllm_code_quality
statllm_executability
statllm_output_quality
statllm_total_score is retained under measurements for analysis but is not
a default prediction target.
samsum_enSamSum-Pref
SamSum-Pref Dataset
SamSum-Pref is a preference-aligned dialogue summarization dataset constructed by sampling from dadastory/SummOrchestra-Qwen3-8B-GRPO-BRL-SAMSUM, and filtering samples using DeepSeek-V3 as the evaluator. Preference scoring follows the AnythingReward evaluation paradigm, adapted to a strict rubric for dialogue-summary quality.
Evaluation Principles
Each sampled summary is scored according to the following weighted criteria:
Key Information Coverage… See the full description on the dataset page: https://huggingface.co/datasets/NebulaPixel/SamSum-Pref.
