CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kilian-group /LMLM-pretrain-dwiki6.1M_v2text1M<n<10M0 likes665 downloads1y agoHugging Face02dwidlee /systemone-lite-general systemone-lite-general Synthetic typed-decision rows for systemone-lite (letter-alias choice labels for causal LM SFT). Not affiliated with TypeSafe AI / Jev. Labels are rule-based, not human prefs. Splits Split Rows Notes train 32 400 Stratified mix of 3 gyms test 3 600 iid held-out by task test_hard 5 400 layout / paraphrase / option-subset shift full 36 000 train + iid test Gyms TicketDungeon: ticket.route… See the full description on the dataset page: https://huggingface.co/datasets/dwidlee/systemone-lite-general.texttext-classification10K<n<100K0 likes242 downloads7d agoHugging Face03dw-indie /pad-auto-solver-reviewed PAD Reviewed Dataset Canonical reviewed PAD board/orb artifacts for dw-indie/pad-auto-solver-reviewed. This repository contains immutable reviewed package revisions and does not contain raw captures, training runs, checkpoints, or model binaries. Packages exported: 28 Active catalog datasets: 14 Catalog schema: 3 Layout packages/<dataset_id>.tar: deterministic self-contained reviewed package catalog.json: active revision heads and coverage summary… See the full description on the dataset page: https://huggingface.co/datasets/dw-indie/pad-auto-solver-reviewed.tabularimage-classification10K<n<100K1 likes208 downloads17d agoHugging Face04dwidlee /systemone-lite-phase2 systemone-lite-phase2 Typed System One distill rows (task / state / instructions / criteria / label_alias) for systemone-lite. Critical: train / test hygiene (2026-09-22) Earlier local mixes had severe train∩eval state leakage (debate ~91%, word_games ~87%, connect4 ~37% state_task overlap). This Hub revision is rebuilt with 0.00% train∩test overlap on state_task fingerprints (scripts/audit_train_eval_overlap.py). Mechanism Detail word_games Disjoint… See the full description on the dataset page: https://huggingface.co/datasets/dwidlee/systemone-lite-phase2.texttext-classification100K<n<1M0 likes128 downloads3d agoHugging Face05kilian-group /LMLM-pretrain-dwiki6.1Mtext1M<n<10M0 likes94 downloads1y agoHugging Face06DFKI-SLT /DWIE Dataset Card for DWIE Dataset Summary DWIE (Deutsche Welle corpus for Information Extraction) is a new dataset for document-level multi-task Information Extraction (IE). It combines four main IE sub-tasks: 1.Named Entity Recognition: 23,130 entities classified in 311 multi-label entity types (tags). 2.Coreference Resolution: 43,373 entity mentions clustered in 23,130 entities. 3.Relation Extraction: 21,749 annotated relations between entities classified in 65… See the full description on the dataset page: https://huggingface.co/datasets/DFKI-SLT/DWIE.textfeature-extractionn<1K4 likes91 downloads2y agoHugging Face07kilian-group /LMLM-pretrain-dwiki6.1M_cleanedtext1M<n<10M0 likes81 downloads11mo agoHugging Face08lonestar108 /dwitter Dataset Card for "dwitter" More Information needed text10K<n<100K0 likes52 downloads3y agoHugging Face09dwishank /AmazonReviewsCleanedDataset2023tabular100K<n<1M0 likes44 downloads2mo agoHugging Face10shishir-dwi /News-Article-Categorization_IAB Article and Category Dataset Overview This dataset contains a collection of articles, primarily news articles, along with their respective IAB (Interactive Advertising Bureau) categories. It can be a valuable resource for various natural language processing (NLP) tasks, including text classification, text generation, and more. Dataset Information Number of Samples: 871,909 Number of Categories: 26 Column Information text: The text of the article.… See the full description on the dataset page: https://huggingface.co/datasets/shishir-dwi/News-Article-Categorization_IAB.texttext-classification100K<n<1M4 likes35 downloads3y agoHugging Face11dwightlangston /isbndb-dumptext100M<n<1B0 likes28 downloads5mo agoHugging Face12dwin1412 /items_raw_litetabular10K<n<100K0 likes24 downloads6mo agoHugging Face13dwisaji /indonesia-telecomunication-sentiment-datasetDataset Contain sentimen for Indonesia Communication Industry. Source from Twitter and manually annotated in prodigy spacy text1K<n<10K4 likes16 downloads4y agoHugging Face14math-extraction-comp /dwikitheduck__gemma-2-2b-id-insttabular1K<n<10K0 likes14 downloads2y agoHugging Face15dwin1412 /items_raw_fulltabular100K<n<1M0 likes14 downloads6mo agoHugging Face16plaguss /the_office_dwight_uncleanedtext1K<n<10K0 likes13 downloads3y agoHugging Face17open-llm-leaderboard /dwikitheduck__gen-inst-1-detailsgated Dataset Card for Evaluation run of dwikitheduck/gen-inst-1 Dataset automatically created during the evaluation run of model dwikitheduck/gen-inst-1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/dwikitheduck__gen-inst-1-details.tabular10K<n<100K0 likes13 downloads2y agoHugging Face18deetsadi /processed_dwi_cropped Dataset Card for "processed_dwi_cropped" More Information needed imagen<1K0 likes12 downloads3y agoHugging Face19deetsadi /processed_dwi_with_adc Dataset Card for "processed_dwi_with_adc" More Information needed imagen<1K0 likes12 downloads3y agoHugging Face200xlisda /chess-dataset-by-dwirizaldytextn<1K0 likes12 downloads1y agoHugging Face21dwin1412 /items_prompts_litetext10K<n<100K0 likes12 downloads6mo agoHugging Face22deetsadi /processed_dwi_sobel_thresh Dataset Card for "processed_dwi_sobel_thresh" More Information needed imagen<1K0 likes11 downloads3y agoHugging Face23deetsadi /processed_dwi_sobel_all_b_values_large_mask Dataset Card for "processed_dwi_sobel_all_b_values_large_mask" More Information needed imagen<1K0 likes11 downloads3y agoHugging Face24open-llm-leaderboard /dwikitheduck__gen-try1-detailsgated Dataset Card for Evaluation run of dwikitheduck/gen-try1 Dataset automatically created during the evaluation run of model dwikitheduck/gen-try1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/dwikitheduck__gen-try1-details.tabular10K<n<100K0 likes11 downloads2y agoHugging Face25open-llm-leaderboard /dwikitheduck__gemma-2-2b-id-instruct-detailsgated Dataset Card for Evaluation run of dwikitheduck/gemma-2-2b-id-instruct Dataset automatically created during the evaluation run of model dwikitheduck/gemma-2-2b-id-instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/dwikitheduck__gemma-2-2b-id-instruct-details.tabular10K<n<100K0 likes11 downloads2y agoHugging Face26deetsadi /processed_dwi_soft_edge_from_semantic Dataset Card for "processed_dwi_soft_edge_from_semantic" More Information needed imagen<1K0 likes10 downloads3y agoHugging Face27plaguss /the_office_ds_dwight_qaNOTE This dataset is under development. To load the dataset and prepare it for DPO training. from datasets import load_dataset ds = load_dataset("plaguss/the_office_ds_dwight_qa") def return_prompt_and_responses(samples) -> dict[str, str]: return { "prompt": "Question: " + samples["prompt"] + "\n\nAnswer: ", "chosen": samples[samples["label"]], "rejected": samples["response_2" if samples["label"] == "response_1" else "response_1"], } rm_columns = ["response_1"… See the full description on the dataset page: https://huggingface.co/datasets/plaguss/the_office_ds_dwight_qa.text1K<n<10K0 likes10 downloads3y agoHugging Face28deetsadi /processed_dwi_fixed Dataset Card for "processed_dwi_fixed" More Information needed imagen<1K0 likes9 downloads3y agoHugging Face29deetsadi /processed_dwi_cropped_soft_edge Dataset Card for "processed_dwi_cropped_soft_edge" More Information needed imagen<1K0 likes9 downloads3y agoHugging Face30deetsadi /processed_dwi_all_b_values_semantic Dataset Card for "processed_dwi_all_b_values_semantic" More Information needed imagen<1K0 likes9 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.