dragon
Datasets
All datasets matching “dragon”banned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/Dragonegg2026/banned-historical-archives.dragon
Dataset Card for DRAGON
🧾 ArXiv Preprint
DRAGON is a large-scale Dataset of Realistic imAges Generated by diffusiON models.
The dataset includes a total of 2.5 million training images and 100,000 test images generated using 25 diffusion models, spanning both recent advancements and older, well-established architectures.
Dataset Details
Dataset Description
The remarkable ease of use of diffusion models for image generation has led to a proliferation of… See the full description on the dataset page: https://huggingface.co/datasets/lesc-unifi/dragon.dragonballdaima
Bangumi Image Base of Dragon Ball Daima
This is the image base of bangumi Dragon Ball Daima, we detected 46 characters, 8351 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/dragonballdaima.DragonData-Finance-Corpus
DragonData Finance Corpus
The world's largest open, permissively-licensed finance corpus for LLM pretraining.
Built by Dragon Limited - 100% free, publicly available sources.
Overview
Attribute
Value
Repository
dragonlimited/DragonData-Finance-Corpus
Tokenizer
bigcode/starcoder2-3b (vocab 49,152)
Format
Binary shards (data/train-XXXXX-of-00001.bin, uint16)
Target
46T tokens across 25 domains
Current
~33B tokens, 246 shards
License
Permissive… See the full description on the dataset page: https://huggingface.co/datasets/dragonlimited/DragonData-Finance-Corpus.ksponspeechCIFAKE-image-dataset
