datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trajectory_data_dream_32
d3LLM Trajectory Dataset
Project Page | Paper | GitHub | Blog
This repository contains the pseudo-trajectory distillation data used for training d3LLM (pseuDo-Distilled Diffusion Large Language Model), as introduced in the paper "d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation".
Introduction
d3LLM is a framework designed to strike a balance between accuracy and parallelism in diffusion-based large language models (dLLMs). This dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/d3LLM/trajectory_data_dream_32.bill_text_us
Dataset Card for "bill_text_us"
Dataset Summary
Dataset for US Congressional bills (bill_text_us).
Supported Tasks and Leaderboards
More Information Needed
Languages
English
Dataset Structure
Data Instances
default
Data Fields
id: id of the bill in format(congress number + bill type + bill number + bill version).
congress: number of the congress.
bill_type: type of the bill.
bill_number: number of the… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_text_us.dream-of-the-red-chamber-continuations
红楼梦续写 · Dream of the Red Chamber: 100 AI Continuations
项目简介
本数据集包含 92 个独立的AI续写版本,续写中国古典文学巅峰之作《红楼梦》的第八十一回至第一百零八回(共28回)。所有续写严格遵循曹雪芹前八十回中埋下的伏笔、谶语和人物命运,完全拒绝高鹗续书。
为什么做这个数据集
《红楼梦》的结局是世界文学史上最大的悬案之一。曹雪芹约于1763年去世前未能完成全书,仅留下前八十回。1791年左右,高鹗发表了一百二十回本,补写了后四十回,但红学研究日益表明高鹗续书严重违背了曹雪芹在前八十回中精心布置的伏笔。
曹雪芹原意 vs 高鹗续书
情节
曹雪芹原意
高鹗续书
黛玉之死
泪尽而亡,呼应"绛珠还泪"神话
焚稿断痴情
宝玉宝钗婚姻
"纵然是齐眉举案,到底意难平"
掉包计骗婚
贾府败落
政治牵连,锦衣军抄家,"忽喇喇似大厦倾"
败而复兴,"兰桂齐芳"
结局… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/dream-of-the-red-chamber-continuations.bill_labels_us
Dataset Card for "bill_labels_us"
Dataset Summary
Dataset for US Congressional bills with policy area and legislative subjects information (bill_labels_us). Contains data for bills from the 108th to the 118th Congress, approximately 119,000 documents.
Supported Tasks and Leaderboards
More Information Needed
Languages
English
Dataset Structure
Data Instances
default
Data Fields
id: id of the bill in… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_labels_us.bill_committees_us
Dataset Card for "bill_committees_us"
Dataset Summary
Dataset for US Congressional bills with committees information (bill_committees_us). Contains data for bills from the 108th to the 118th Congress, approximately 132,000 documents.
Supported Tasks and Leaderboards
More Information Needed
Languages
English
Dataset Structure
Data Instances
default
Data Fields
id: id of the bill in format(congress number +… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_committees_us.
