datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DREAM-1K
DREAM-1K
DREAM-1K (Description
with Rich Events, Actions, and Motions) is a challenging video description benchmark. It contains a collection of 1,000 short (around 10 seconds) video clips with diverse complexities from five different origins: live-action movies, animated movies, stock videos, long YouTube videos, and TikTok-style short videos. We provide a fine-grained manual annotation for each video.
Bellow is the dataset statistics:
testscopejudge
ScopeJudge dataset card
ScopeJudge is a calibration benchmark for pre-execution gating of autonomous
offensive-security agents. It contains 100 complete ATIF v1.7 trajectories
generated across five source-agent model families. Every one of the 4,897 tool
calls was independently labeled by five professional security experts as
in-scope or out-of-scope.
The strict-majority golden contains 377 out-of-scope calls (7.7%). Reviewers
disagreed on 582 calls (11.9%);… See the full description on the dataset page: https://huggingface.co/datasets/dreadnode/scopejudge.bill_summary_us
Dataset Card for "bill_summary_us"
Dataset Summary
Dataset for summarization of summarization of US Congressional bills (bill_summary_us).
Supported Tasks and Leaderboards
More Information Needed
Languages
English
Dataset Structure
Data Instances
default
Data Fields
id: id of the bill in format(congress number + bill type + bill number + bill version).
congress: number of the congress.
bill_type: type of… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_summary_us.lm-eval-results-DreadPoor-Harpy-7B-Model_Stock-private
Dataset Card for Evaluation run of DreadPoor/Harpy-7B-Model_Stock
Dataset automatically created during the evaluation run of model DreadPoor/Harpy-7B-Model_Stock
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-DreadPoor-Harpy-7B-Model_Stock-private.sticker-studio-dressup
Sticker Studio — Dress-Up Game
A lightweight, static-PNG dress-up game (window / canvas based). Pick a base
character, tap clothing & accessory assets to add them onto the canvas, drag to
reposition (mouse and touch), and Save a flattened PNG of your creation.
Upload your own images into the Characters or Wardrobe libraries (persisted on the backend).
Tech Stack
Frontend: React 19, Tailwind, shadcn/ui, framer-motion (HTML <canvas> stage, Pointer Events for… See the full description on the dataset page: https://huggingface.co/datasets/launch-calcium/sticker-studio-dressup.bill_text_us
Dataset Card for "bill_text_us"
Dataset Summary
Dataset for US Congressional bills (bill_text_us).
Supported Tasks and Leaderboards
More Information Needed
Languages
English
Dataset Structure
Data Instances
default
Data Fields
id: id of the bill in format(congress number + bill type + bill number + bill version).
congress: number of the congress.
bill_type: type of the bill.
bill_number: number of the… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_text_us.dream-machine-ai
██████╗ █████╗ ████████╗ █████╗ ███████╗███████╗████████╗
██╔══██╗██╔══██╗╚══██╔══╝██╔══██╗██╔════╝██╔════╝╚══██╔══╝
██║ ██║███████║ ██║ ███████║███████╗█████╗ ██║
██║ ██║██╔══██║ ██║ ██╔══██║╚════██║██╔══╝ ██║
██████╔╝██║ ██║ ██║ ██║ ██║███████║███████╗ ██║
╚═════╝ ╚═╝ ╚═╝ ╚═╝ ╚═╝ ╚═╝╚══════╝╚══════╝ ╚═╝
✨ Dream-Machine · E-Commerce Synthetic Dataset ✨
1,000 Products · 1,000 Users · 5,000 Behaviors · JD /… See the full description on the dataset page: https://huggingface.co/datasets/dream-machine-ai/dream-machine-ai.dream-of-the-red-chamber-continuations
红楼梦续写 · Dream of the Red Chamber: 100 AI Continuations
项目简介
本数据集包含 92 个独立的AI续写版本,续写中国古典文学巅峰之作《红楼梦》的第八十一回至第一百零八回(共28回)。所有续写严格遵循曹雪芹前八十回中埋下的伏笔、谶语和人物命运,完全拒绝高鹗续书。
为什么做这个数据集
《红楼梦》的结局是世界文学史上最大的悬案之一。曹雪芹约于1763年去世前未能完成全书,仅留下前八十回。1791年左右,高鹗发表了一百二十回本,补写了后四十回,但红学研究日益表明高鹗续书严重违背了曹雪芹在前八十回中精心布置的伏笔。
曹雪芹原意 vs 高鹗续书
情节
曹雪芹原意
高鹗续书
黛玉之死
泪尽而亡,呼应"绛珠还泪"神话
焚稿断痴情
宝玉宝钗婚姻
"纵然是齐眉举案,到底意难平"
掉包计骗婚
贾府败落
政治牵连,锦衣军抄家,"忽喇喇似大厦倾"
败而复兴,"兰桂齐芳"
结局… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/dream-of-the-red-chamber-continuations.trace-demomsmarco-2.1-segmentedbill_labels_us
Dataset Card for "bill_labels_us"
Dataset Summary
Dataset for US Congressional bills with policy area and legislative subjects information (bill_labels_us). Contains data for bills from the 108th to the 118th Congress, approximately 119,000 documents.
Supported Tasks and Leaderboards
More Information Needed
Languages
English
Dataset Structure
Data Instances
default
Data Fields
id: id of the bill in… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_labels_us.DreadPoor__Promissum_Mane-8B-LINEAR-lorablated-details
Dataset Card for Evaluation run of DreadPoor/Promissum_Mane-8B-LINEAR-lorablated
Dataset automatically created during the evaluation run of model DreadPoor/Promissum_Mane-8B-LINEAR-lorablated
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Promissum_Mane-8B-LINEAR-lorablated-details.test_v2sticker-studio-dressup
Sticker Studio — Dress-Up Game
A lightweight, static-PNG dress-up game (window / canvas based). Pick a base
character, tap clothing & accessory assets to add them onto the canvas, drag to
reposition (mouse and touch), and Save a flattened PNG of your creation.
Upload your own images into the Characters or Wardrobe libraries (persisted on the backend).
Tech Stack
Frontend: React 19, Tailwind, shadcn/ui, framer-motion (HTML <canvas> stage, Pointer Events for… See the full description on the dataset page: https://huggingface.co/datasets/that-username-is-not-available/sticker-studio-dressup.dream_interpretationfoo_dataset_1svelte-5-sveltekit-2
Svelte 5, SvelteKit 2 and CLI Dataset
Dataset based on the new Svelte 5, SvelteKit 2 documentation. Including all topics and example codes.
This is structured with the Qwen2 chat template, too finetune Qwen2.5-Coder models.
spanglish-sentences
Spanglish Sentences
A dataset of 10,576 Spanish–English code-switched ("Spanglish") sentences paired with English translations, intended for training and evaluating code-switch translation models.
Data format
Each line of spanglish_sentences.jsonl is a JSON object with two fields:
field
description
sentence
A Spanglish utterance (mixed Spanish / English, or monolingual in either language).
english_translation
The English translation. When the source is… See the full description on the dataset page: https://huggingface.co/datasets/drewoodward/spanglish-sentences.DreadPoor__BaeZel_V3-8B-Model_Stock-details
Dataset Card for Evaluation run of DreadPoor/BaeZel_V3-8B-Model_Stock
Dataset automatically created during the evaluation run of model DreadPoor/BaeZel_V3-8B-Model_Stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__BaeZel_V3-8B-Model_Stock-details.DreadPoor__Heart_Stolen-8B-Model_Stock-details
Dataset Card for Evaluation run of DreadPoor/Heart_Stolen-8B-Model_Stock
Dataset automatically created during the evaluation run of model DreadPoor/Heart_Stolen-8B-Model_Stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Heart_Stolen-8B-Model_Stock-details.machine.dreamdreamstruct-human-UI-classes-1kDreadPoor__Summer_Rain-8B-TIES-details
Dataset Card for Evaluation run of DreadPoor/Summer_Rain-8B-TIES
Dataset automatically created during the evaluation run of model DreadPoor/Summer_Rain-8B-TIES
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Summer_Rain-8B-TIES-details.MMDA_BenchThe dataset proposed in MoDora
Some documents involving sensitive data are hidden. If you believe any content in this open source dataset infringes upon your copyright, please contact us, and we will remove it.
DreadPoor__WIP_Damascus-8B-TIES-details
Dataset Card for Evaluation run of DreadPoor/WIP_Damascus-8B-TIES
Dataset automatically created during the evaluation run of model DreadPoor/WIP_Damascus-8B-TIES
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__WIP_Damascus-8B-TIES-details.SGOCR
SGOCR
SGOCR is a spatially-grounded OCR visual question answering dataset for training and evaluating models that must read, localize, and reason about text in images.
The dataset contains grounded question-answer pairs over ChartQA, TextOCR, and COCO/COCO-Text source images. It is designed for OCR-aware VQA, text grounding, region-conditioned QA, and data-centric experiments around scene text understanding.
Project repository:… See the full description on the dataset page: https://huggingface.co/datasets/dreeseaw/SGOCR.DreadPoor__Aspire_V2-8B-Model_Stock-details
Dataset Card for Evaluation run of DreadPoor/Aspire_V2-8B-Model_Stock
Dataset automatically created during the evaluation run of model DreadPoor/Aspire_V2-8B-Model_Stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Aspire_V2-8B-Model_Stock-details.RPGPT_PublicDomain-alpacaExperimental Synthetic Dataset of Public Domain Character Dialogue in Roleplay Format
Generated using scripts from my https://github.com/practicaldreamer/build-a-dataset repo
license: mit
feature_extraction_dataset
