datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FISH_spots
FISH_spots Dataset
The manually verified in situ hybridization fluorescence images and point coordinate dataset.
This dataset contains images and annotations for the task of single-molecule fluorescence in situ hybridization (FISH) spot detection, supporting 2D, 3D, and simulated noisy data. The structure is designed for deep learning model development, training, and evaluation.
Directory Structure
FISH_spots/
├── 2d/
│ ├── csv/
│ ├── image/
│ ├── image_raw/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/GangCaoLab/FISH_spots.ganjoor
Dataset Card for Dataset Name
This is the csv format of the Ganjoor Database that is published in their github
Dataset Details
Curated by: Navid Abbaspoor
Language(s) (NLP): Persian (Farsi)
License: Creative Commons Attribution 4.0 International (cc-by-4.0)
Dataset Description
This dataset contains almost all of poems by Iran's great poets through many many past years till now. The original database was tabular, that I convert it to a csv format that… See the full description on the dataset page: https://huggingface.co/datasets/mabidan/ganjoor.hindi-article-summarization
Summary
hindi-article-summarization is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-article-summarization.SMS-dataset-sample-10klicense: mit
language:
en
📌 Free 10K SMS Preview Dataset
This is a preview subset of the full OTP + OTP INTENT + Phishing dataset.
👉 For full 73K dataset:
https://huggingface.co/datasets/gandharvbakshi/SMS-dataset-OTP-OTP_INTENT_Phishing
Text-Classification-and-Relation-Event-Extraction-Mix-datasetsThe paper of GIELLM dataset.
https://arxiv.org/abs/2311.06838
Cite:
@article{gan2023giellm,
title={Giellm: Japanese general information extraction large language model utilizing mutual reinforcement effect},
author={Gan, Chengguang and Zhang, Qinghao and Mori, Tatsunori},
journal={arXiv preprint arXiv:2311.06838},
year={2023}
}
The dataset constructed base in livedoor news corpus 関口宏司 https://www.rondhuit.com/download.html
additional-bed-19f678
additional-bed-19f678
Synthetic weather test data: 40 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/gangyeonggil392/additional-bed-19f678.separate-storage-f87850
separate-storage-f87850
Synthetic weather test data: 32 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/gangjeongja/separate-storage-f87850.hindi-headline-article-generation
Summary
hindi-headline-article-generation is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi Version: 1.0
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-headline-article-generation.salesdataganda-tswana_sentence-pairs
Ganda-Tswana_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ganda-Tswana_Sentence-Pairs
Number of Rows: 183016
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ganda-tswana_sentence-pairs.agriparts_structuredamharic-ganda_sentence-pairs
Amharic-Ganda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Amharic-Ganda_Sentence-Pairs
Number of Rows: 179444
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-ganda_sentence-pairs.igbo-ganda_sentence-pairs
Igbo-Ganda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Ganda_Sentence-Pairs
Number of Rows: 111322
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-ganda_sentence-pairs.ganda-chichewa_sentence-pairs
Ganda-Chichewa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ganda-Chichewa_Sentence-Pairs
Number of Rows: 187295
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ganda-chichewa_sentence-pairs.dyula-ganda_sentence-pairs
Dyula-Ganda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Ganda_Sentence-Pairs
Number of Rows: 77324
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-ganda_sentence-pairs.ganda-wolof_sentence-pairs
Ganda-Wolof_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ganda-Wolof_Sentence-Pairs
Number of Rows: 49136
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ganda-wolof_sentence-pairs.ganda-tsonga_sentence-pairs
Ganda-Tsonga_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ganda-Tsonga_Sentence-Pairs
Number of Rows: 156337
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ganda-tsonga_sentence-pairs.ganda-swati_sentence-pairs
Ganda-Swati_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ganda-Swati_Sentence-Pairs
Number of Rows: 57049
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ganda-swati_sentence-pairs.fulah-ganda_sentence-pairs
Fulah-Ganda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fulah-Ganda_Sentence-Pairs
Number of Rows: 91895
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fulah-ganda_sentence-pairs.ewe-ganda_sentence-pairs
Ewe-Ganda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ewe-Ganda_Sentence-Pairs
Number of Rows: 139118
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ewe-ganda_sentence-pairs.ganda-umbundu_sentence-pairs
Ganda-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ganda-Umbundu_Sentence-Pairs
Number of Rows: 76417
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ganda-umbundu_sentence-pairs.afrikaans-ganda_sentence-pairs
Afrikaans-Ganda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Ganda_Sentence-Pairs
Number of Rows: 477046
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-ganda_sentence-pairs.GCE_OLevel_Science_updatedganda-twi_sentence-pairs
Ganda-Twi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ganda-Twi_Sentence-Pairs
Number of Rows: 156868
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ganda-twi_sentence-pairs.ganda-pedi_sentence-pairs
Ganda-Pedi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ganda-Pedi_Sentence-Pairs
Number of Rows: 115302
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ganda-pedi_sentence-pairs.fon-ganda_sentence-pairs
Fon-Ganda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Ganda_Sentence-Pairs
Number of Rows: 64936
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-ganda_sentence-pairs.dinka-ganda_sentence-pairs
Dinka-Ganda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dinka-Ganda_Sentence-Pairs
Number of Rows: 31116
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dinka-ganda_sentence-pairs.lingala-ganda_sentence-pairs
Lingala-Ganda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Lingala-Ganda_Sentence-Pairs
Number of Rows: 111492
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-ganda_sentence-pairs.ganda-oromo_sentence-pairs
Ganda-Oromo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ganda-Oromo_Sentence-Pairs
Number of Rows: 71187
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ganda-oromo_sentence-pairs.ganda-nuer_sentence-pairs
Ganda-Nuer_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ganda-Nuer_Sentence-Pairs
Number of Rows: 29008
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ganda-nuer_sentence-pairs.
