yohannabelay/Amharic
Amharic Dataset (Cloned from saillab/taco-datasets) This dataset contains the Amharic language part of the multilingual instruction tuning dataset, originally from the TACO dataset. The Amharic data is part of the multilingual-alpaca-52k-gpt-4 split, focusing on instruction-based dialogue data. Dataset Summary This dataset contains Amharic language instructions and dialogues. It is designed for fine-tuning large language models in multilingual settings. The… See the full description on the dataset page: https://huggingface.co/datasets/yohannabelay/Amharic.
Amharic Dataset (Cloned from saillab/taco-datasets)
This dataset contains the Amharic language part of the multilingual instruction tuning dataset, originally from the TACO dataset. The Amharic data is part of the multilingual-alpaca-52k-gpt-4 split, focusing on instruction-based dialogue data.
Dataset Summary
This dataset contains Amharic language instructions and dialogues. It is designed for fine-tuning large language models in multilingual settings. The dataset includes prompt-response pairs in Amharic, where the model is trained to understand instructions and generate appropriate responses.
- Languages: Amharic
- Total Number of Records: ~66K
- Dataset Size: ~137MB
Dataset Structure
The dataset contains the following fields:
instruction: The instructional prompt in Amharic (e.g., a request or command).input: The input in Amharic (e.g., a question).output: The model's expected response to the instruction in Amharic.
Example Record:
{
"instruction": "የሚከተለውን ጽሑፍ እንደ አወንታዊ፣ አሉታዊ ወይም ገለልተኛ አድርገው ይመድቡት፡-",
"instruction": "ይህን ፊልም ወድጄዋለሁ! የሚገርም ነው።",
"response": "አዎንታዊ"
}
