datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trivia_qa_tiny
Dataset Card for trivia_qa_tiny
This is a preprocessed version of trivia_qa_tiny dataset for benchmarks in LM-Polygraph.
Dataset Details
Dataset Description
Curated by: https://huggingface.co/LM-Polygraph
License: https://github.com/IINemo/lm-polygraph/blob/main/LICENSE.md
Dataset Sources [optional]
Repository: https://github.com/IINemo/lm-polygraph
Uses
Direct Use
This dataset should be used for performing… See the full description on the dataset page: https://huggingface.co/datasets/LM-Polygraph/trivia_qa_tiny.TinyLM
TinyLM Data
This dataset in data.txt is a collection of user/ai conversations across domains such as science, math, programming and writing. It also contains general conversation data.
The dataset is designed to be used for the training of SLMs (Small Language Models).
Format
This is an example of a conversation in the dataset:
<|data|>
<|user|> What is a comet?
<|assistant|> A comet is a big ball of ice and rock. <|endoftext|>
<|user|> Does it look cool?
<|assistant|>… See the full description on the dataset page: https://huggingface.co/datasets/AGofficial/TinyLM.tiny-ai2d-irtlmsys-chat-tiny-20kLMS_chatbot_dataset_tinymy1
