CoolFace
Datasetpublic

yohannabelay/Amharic

Amharic Dataset (Cloned from saillab/taco-datasets) This dataset contains the Amharic language part of the multilingual instruction tuning dataset, originally from the TACO dataset. The Amharic data is part of the multilingual-alpaca-52k-gpt-4 split, focusing on instruction-based dialogue data. Dataset Summary This dataset contains Amharic language instructions and dialogues. It is designed for fine-tuning large language models in multilingual settings. The… See the full description on the dataset page: https://huggingface.co/datasets/yohannabelay/Amharic.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes5downloads
Dataset Card

Amharic Dataset (Cloned from saillab/taco-datasets)

This dataset contains the Amharic language part of the multilingual instruction tuning dataset, originally from the TACO dataset. The Amharic data is part of the multilingual-alpaca-52k-gpt-4 split, focusing on instruction-based dialogue data.

Dataset Summary

This dataset contains Amharic language instructions and dialogues. It is designed for fine-tuning large language models in multilingual settings. The dataset includes prompt-response pairs in Amharic, where the model is trained to understand instructions and generate appropriate responses.

  • —Languages: Amharic
  • —Total Number of Records: ~66K
  • —Dataset Size: ~137MB
  • —

Dataset Structure

The dataset contains the following fields:

  • —instruction: The instructional prompt in Amharic (e.g., a request or command).
  • —input: The input in Amharic (e.g., a question).
  • —output: The model's expected response to the instruction in Amharic.

Example Record:

json
{
  "instruction": "የሚከተለውን ጽሑፍ እንደ አወንታዊ፣ አሉታዊ ወይም ገለልተኛ አድርገው ይመድቡት፡-",
  "instruction": "ይህን ፊልም ወድጄዋለሁ! የሚገርም ነው።",
  "response": "አዎንታዊ"
}