datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LMLM-pretrain-dwiki6.1M_v2systemone-lite-general
systemone-lite-general
Synthetic typed-decision rows for
systemone-lite
(letter-alias choice labels for causal LM SFT).
Not affiliated with TypeSafe AI / Jev. Labels are rule-based, not human prefs.
Splits
Split
Rows
Notes
train
32 400
Stratified mix of 3 gyms
test
3 600
iid held-out by task
test_hard
5 400
layout / paraphrase / option-subset shift
full
36 000
train + iid test
Gyms
TicketDungeon: ticket.route… See the full description on the dataset page: https://huggingface.co/datasets/dwidlee/systemone-lite-general.pad-auto-solver-reviewed
PAD Reviewed Dataset
Canonical reviewed PAD board/orb artifacts for dw-indie/pad-auto-solver-reviewed. This repository
contains immutable reviewed package revisions and does not contain raw captures,
training runs, checkpoints, or model binaries.
Packages exported: 28
Active catalog datasets: 14
Catalog schema: 3
Layout
packages/<dataset_id>.tar: deterministic self-contained reviewed package
catalog.json: active revision heads and coverage summary… See the full description on the dataset page: https://huggingface.co/datasets/dw-indie/pad-auto-solver-reviewed.systemone-lite-phase2
systemone-lite-phase2
Typed System One distill rows (task / state / instructions / criteria /
label_alias) for systemone-lite.
Critical: train / test hygiene (2026-09-22)
Earlier local mixes had severe train∩eval state leakage (debate ~91%,
word_games ~87%, connect4 ~37% state_task overlap). This Hub revision is rebuilt
with 0.00% train∩test overlap on state_task fingerprints
(scripts/audit_train_eval_overlap.py).
Mechanism
Detail
word_games
Disjoint… See the full description on the dataset page: https://huggingface.co/datasets/dwidlee/systemone-lite-phase2.LMLM-pretrain-dwiki6.1MDWIE
Dataset Card for DWIE
Dataset Summary
DWIE (Deutsche Welle corpus for Information Extraction) is a new dataset for document-level multi-task Information Extraction (IE).
It combines four main IE sub-tasks:
1.Named Entity Recognition: 23,130 entities classified in 311 multi-label entity types (tags).
2.Coreference Resolution: 43,373 entity mentions clustered in 23,130 entities.
3.Relation Extraction: 21,749 annotated relations between entities classified in 65… See the full description on the dataset page: https://huggingface.co/datasets/DFKI-SLT/DWIE.LMLM-pretrain-dwiki6.1M_cleaneddwitter
Dataset Card for "dwitter"
More Information needed
AmazonReviewsCleanedDataset2023News-Article-Categorization_IAB
Article and Category Dataset
Overview
This dataset contains a collection of articles, primarily news articles, along with their respective IAB (Interactive Advertising Bureau) categories. It can be a valuable resource for various natural language processing (NLP) tasks, including text classification, text generation, and more.
Dataset Information
Number of Samples: 871,909
Number of Categories: 26
Column Information
text: The text of the article.… See the full description on the dataset page: https://huggingface.co/datasets/shishir-dwi/News-Article-Categorization_IAB.isbndb-dumpitems_raw_liteindonesia-telecomunication-sentiment-datasetDataset Contain sentimen for Indonesia Communication Industry. Source from Twitter and manually annotated in prodigy spacy
dwikitheduck__gemma-2-2b-id-institems_raw_fullthe_office_dwight_uncleaneddwikitheduck__gen-inst-1-details
Dataset Card for Evaluation run of dwikitheduck/gen-inst-1
Dataset automatically created during the evaluation run of model dwikitheduck/gen-inst-1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/dwikitheduck__gen-inst-1-details.processed_dwi_cropped
Dataset Card for "processed_dwi_cropped"
More Information needed
processed_dwi_with_adc
Dataset Card for "processed_dwi_with_adc"
More Information needed
chess-dataset-by-dwirizaldyitems_prompts_liteprocessed_dwi_sobel_thresh
Dataset Card for "processed_dwi_sobel_thresh"
More Information needed
processed_dwi_sobel_all_b_values_large_mask
Dataset Card for "processed_dwi_sobel_all_b_values_large_mask"
More Information needed
dwikitheduck__gen-try1-details
Dataset Card for Evaluation run of dwikitheduck/gen-try1
Dataset automatically created during the evaluation run of model dwikitheduck/gen-try1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/dwikitheduck__gen-try1-details.dwikitheduck__gemma-2-2b-id-instruct-details
Dataset Card for Evaluation run of dwikitheduck/gemma-2-2b-id-instruct
Dataset automatically created during the evaluation run of model dwikitheduck/gemma-2-2b-id-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/dwikitheduck__gemma-2-2b-id-instruct-details.processed_dwi_soft_edge_from_semantic
Dataset Card for "processed_dwi_soft_edge_from_semantic"
More Information needed
the_office_ds_dwight_qaNOTE This dataset is under development.
To load the dataset and prepare it for DPO training.
from datasets import load_dataset
ds = load_dataset("plaguss/the_office_ds_dwight_qa")
def return_prompt_and_responses(samples) -> dict[str, str]:
return {
"prompt": "Question: " + samples["prompt"] + "\n\nAnswer: ",
"chosen": samples[samples["label"]],
"rejected": samples["response_2" if samples["label"] == "response_1" else "response_1"],
}
rm_columns = ["response_1"… See the full description on the dataset page: https://huggingface.co/datasets/plaguss/the_office_ds_dwight_qa.processed_dwi_fixed
Dataset Card for "processed_dwi_fixed"
More Information needed
processed_dwi_cropped_soft_edge
Dataset Card for "processed_dwi_cropped_soft_edge"
More Information needed
processed_dwi_all_b_values_semantic
Dataset Card for "processed_dwi_all_b_values_semantic"
More Information needed
