datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-prescription-datasetchinese_modern_poetry
简介
数据集包括了近现代的中国诗人及外国诗人(中译版)作品,所有作品著作权归原作者所有,侵删请联系aa531811820@gmail.com
chinese_poems.jsonl为原数据,training_imagery2-5_maxlen256.json 分别是根据2-5个关键意象生成诗歌的相关数据集
数据来源于网络,包括但不限于
https://github.com/sheepzh/poetry
https://bedtimepoem.com/
https://poemwiki.org/
baidu、google、zhihu等
一些作品
使用此数据集训练ChatGLM、LLaMA7b模型生成的诗歌,更多诗歌查看poems目录
chinese-lips-speech-slide-probe
Chinese-LiPS Speech + Slide Probe
A self-contained probe set for testing whether visual slide context helps
simultaneous speech translation — with the input as audio, not transcripts.
Why audio matters: feeding a transcript to a text LLM deletes the acoustic
ambiguity (homophones, polysemy) that slide context is meant to resolve; the
transcript already commits to one reading. Any honest test of "does vision help
streaming ST" must consume speech.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-speech-slide-probe.ALLaVA-4V-Chinese
ALLaVA-4V for Chinese
This is the Chinese version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Chinese through ChatGPT and instructed ChatGPT not to translate content related to OCR.
The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V.
Citation
If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Chinese.anny-render-corpus-train
anny-render-corpus
A camera-controlled render corpus from the ANNY rig, and the measurements that motivated it.
Code: weftspun/anny-render-corpus, on the 6-datasource side of the hexagon.
Everything here is produced by scripts in that repository and can be regenerated from it.
What this is for
Asked in plain language for eight camera azimuths, OmniGen2 returns a body that does not
turn. Recovered azimuth tracks the request with a slope of 0.04, where 1.00 is… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/anny-render-corpus-train.traditional-chinese-historical-ocr-lo-chia-luen
Traditional Chinese Historical OCR Dataset
(Lo Chia-Lun Manuscripts)
This dataset consists of manually annotated OCR text crops derived from the Lo Chia-Lun Manuscript Collection (羅家倫文稿), hosted by the National Chengchi University Library.
The dataset is designed to support research on Traditional Chinese OCR, particularly for historical documents characterized by vertical layouts, handwritten or semi-printed glyphs, and long-form text lines.
Due to archival and… See the full description on the dataset page: https://huggingface.co/datasets/ZihCiLin/traditional-chinese-historical-ocr-lo-chia-luen.MMC4-130k-chinese-imageMMC4-130k-chinese是对MMC4中,抽样了130k左右 simliarty较高的图文pair得到的数据集
Chinese版本是对这里所有的caption进行了翻译。
我们会陆续将更多数据集发布到hf,包括
Coco Caption的中文翻译
CoQA的中文翻译
CNewSum的Embedding数据
增广的开放QA数据
WizardLM的中文翻译
如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。
骆驼(Luotuo): 开源中文大语言模型
https://github.com/LC1332/Luotuo-Chinese-LLM
骆驼(Luotuo)项目是由冷子昂 @ 商汤科技, 陈启源 @ 华中师范大学 以及 李鲁鲁 @ 商汤科技 发起的中文大语言模型开源项目,包含了一系列语言模型。
( 注意: 陈启源 正在寻找2024推免导师,欢迎联系 )
骆驼项目不是商汤科技的官方产品。
Citation
Please cite the repo if you use the data or code… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/MMC4-130k-chinese-image.chintan10kYAML Metadata:
language:
- en
pretty_name: Chintan10k Dataset
tags:
- image-captioning
- chain-of-thought
- sdxl-prompts
license: wtfpl
task_categories:
- image-to-text
- text-generation
Dataset Card for Chintan10k
Dataset Summary
The Chintan10k dataset is a collection of 10,000 entries, each comprising an image URL, a corresponding caption, a chain-of-thought analysis, and an SDXL prompt. This dataset is designed to facilitate tasks such as image captioning, chain-of-thought… See the full description on the dataset page: https://huggingface.co/datasets/bombaygamercc/chintan10k.chintan33kYAML Metadata:
language:
- en
pretty_name: Chintan33k Dataset
tags:
- image-captioning
- chain-of-thought
- sdxl-prompts
license: wtfpl
task_categories:
- image-to-text
- text-generation
Dataset Card for Chintan33k
Dataset Summary
The Chintan33k dataset is a collection of 33,000 entries, each comprising an image URL, a corresponding caption, a chain-of-thought analysis, and an SDXL prompt. This dataset is designed to facilitate tasks such as image captioning, chain-of-thought… See the full description on the dataset page: https://huggingface.co/datasets/bombaygamercc/chintan33k.chintan1kChilmpAVE
