datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
telly-oasis-1Telly-Oasis-1-dataset
A high-quality, multi-turn conversational dataset designed for LLM fine-tuning and instruction-following.This dataset is used to train the Telly Oasis 1 language model Which is curently under development
Telly Oasis 1 is a conversational instruction dataset stored in JSONL format. It is designed for training, fine-tuning, and evaluating Large Language Models (LLMs), particularly chat and instruction-following models.
The dataset contains multi-turn conversations between a… See the full description on the dataset page: https://huggingface.co/datasets/Trytellypls/telly-oasis-1.Oasis-Corpus
Dataset Card for Oasis-Corpus
Dataset Description
Oasis-Corpus is a 783GB high-quality bilingual corpus.
All data in Oasis-Corpus are built by Oasis and sourced from Common Crawl.
It consists of 374GB of Chinese from 17 recent dumps and 409GB of English textual data from 5 dumps.
Languages
English(409GB, 70,121,125 lines) and Chinese(374GB, 110,580,964 lines)
Data Splits
Language
Dump
docs
size
Chinese
cc-may-jun-2023-zh
5,627,020
19.31 GB… See the full description on the dataset page: https://huggingface.co/datasets/Oasis-Team/Oasis-Corpus.
