datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
educhat-sft-002-data-osm每条数据由一个存放对话的list和与数据对应的system_prompt组成。list中按照Q,A顺序存放对话。
数据来源为开源数据,使用CleanTool数据清理工具去重。
Azure-TTS-Osman-WikipediaOsmosisProofling-SFT-Datalaravel12-alpaca-datasetosmanlica-bench-v1
Osmanlica-Bench-v1
A benchmark dataset for Ottoman Turkish transliteration systems.
Quick Start
from datasets import load_dataset
dataset = load_dataset("bilirkesi/osmanlica-bench-v1")
print(dataset["test"][0])
Installation
pip install datasets
Documentation
GitHub
Benchmark Report
hr-consultor-vendasmagibu-identity-datasetofpptv2
