datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Rasa
Rasa: Towards Building an Expressive Multilingual Text-To-Speech Dataset for Indian Languages
Funded by: Bhashini, Ministry of Electronics and Information Technology, Government of IndiaSupported by: EkStep Foundation and Nilekani Philanthropies
Overview
We introduce Rasa, the first high-quality multilingual expressive Text-to-Speech (TTS) dataset for any Indian language. It comprises a minimum of 20 hours per speaker with a target of covering
a female and male… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Rasa.RASAM-2
RASAM 2 — Line-Level HTR Ground-Truth for Arabic historical manuscripts (Maghrebi)
Dataset Description
RASAM 2 extends the initial initiative RASAM 1 by integrating the outcomes of the second manuscript transcription hackathon organized by the Research Consortium Middle-East and Muslim Worlds (GIS MOMM) and Calfa at BULAC (Paris) between December 2021 and March 2022.
This HuggingFace dataset provides cropped line-level images paired with their… See the full description on the dataset page: https://huggingface.co/datasets/calfa-ai/RASAM-2.RASAM-1
RASAM 1 — Line-Level HTR Ground-Truth for Arabic historical manuscripts (Maghrebi)
Dataset Description
RASAM 1 is a specialized dataset for Handwritten Text Recognition (HTR) focusing on Arabic historical manuscripts in Maghrebi script.
This HuggingFace dataset provides cropped line-level images paired with their transcriptions and rich metadata for 3 Arabic Maghrebi manuscripts from the BULAC Library. It is designed as a ready-to-use resource for… See the full description on the dataset page: https://huggingface.co/datasets/calfa-ai/RASAM-1.rasa-tts-16krasa-tts-48kRasaif-Classical-Arabic-English-Parallel-texts
Introduction
This dataset represents a curated collection of parallel Arabic-English texts, featuring the translations of 24 historically and culturally significant books. These texts provide a portal to the intellectual and literary heritage of the Arabic-speaking world during its classical period.
Content Details
Contained within this dataset are English translations of the following texts, sourced from the Rasaif website:
A Muslim Manual of War
Al-Hanin Ila'l-Awtan… See the full description on the dataset page: https://huggingface.co/datasets/ImruQays/Rasaif-Classical-Arabic-English-Parallel-texts.Rasa-Annotated-25kHz
Rasa-Annotated
Enhanced version of Rasa with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,102
Total Duration: 46.92 hours
Average Duration: 6.47 seconds
Duration Range: 0.31s - 45.34s
Average Phonemes: 18.4 per sample
Average Kanade Tokens: 530.7 per sample
Global Embedding Dimension: 128
Gender Distribution
Gender
Count
Female
12,583
Male
13,519
Style Distribution
Style
Count… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Rasa-Annotated-25kHz.ras-assess
RAS assessment
Paid measurement booking door. Never a grade or certificate. Coming—Paddle. No price on this card.
Live: https://councilof.ai/assess
Council OS: https://councilof.ai/os
Council Space: https://councilof.ai/gspc-arena
Measurement, not certification. Empty slots are not for sale. No scores on this card.
Jail is a measured floor, not a 16th pane.
The live board is the authority
GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub… See the full description on the dataset page: https://huggingface.co/datasets/csoai/ras-assess.rasa_prepro_with_audioRasa-Annotated-V1
Rasa-Annotated
Enhanced version of Rasa with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,102
Total Duration: 46.92 hours
Average Duration: 6.47 seconds
Duration Range: 0.31s - 45.34s
Average Phonemes: 18.4 per sample
Average Kanade Tokens: 264.5 per sample
Global Embedding Dimension: 128
Gender Distribution
Gender
Count
Female
12,583
Male
13,519
Style Distribution
Style
Count… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Rasa-Annotated-V1.RASAM
RASAM dataset
An Open Dataset for the Recognition and Analysis of Scripts in Arabic Maghrebi
How to cite
The paper has been presented during the ICDAR 2021 conference (ASAR workshop). To cite this work and this dataset, please use the following informations:
@InProceedings{2021rasam-dataset,
author="Vidal-Gorène, Chahan and Lucas, Noëmie and Salah, Clément and Decours-Perez, Aliénor and Dupin, Boris",
editor="Barney Smith, Elisa H. and Pal, Umapada",
title="RASAM -- A… See the full description on the dataset page: https://huggingface.co/datasets/johnlockejrr/RASAM.rasa_san_originalrasaif-translations
Dataset Source
https://rasaif.com
gemini-dataset-rasalgethicommand-generation-calm-demo-v1
Command Generator dataset for rasa-calm-demo (v1)
This is an instruction tuning dataset consisting of prompt-command pairs. These pairs can be used to train a small LLM like
Llama 3.1 8b to act as a
command generator in the CALM paradigm.
The technical details of how a CALM assistant works can be found in this paper.
Dataset Details
Dataset Description
The dataset consists prompt-command pairs, where prompt consists of an instruction for the LLM to… See the full description on the dataset page: https://huggingface.co/datasets/rasa/command-generation-calm-demo-v1.nava-rasa-myanmar-corpus
Nava-Rasa Myanmar Corpus
The Nava-Rasa Myanmar Corpus is a structurally curated, single-label text classification dataset designed for computational linguistics and sentiment analysis in the Myanmar (Burmese) language. Built upon the foundation of classical Indian poetics (Alanka) and long-standing Burmese literary theory, this corpus frames textual sentiment analysis through the Navarasa (The Nine Literary Aesthetics/Emotions).
Unlike conventional social media sentiment sets… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/nava-rasa-myanmar-corpus.rasa_model_comparisonrasa-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: rasa
Data Source Link: https://rasa.com/docs/
Data Source License: https://github.com/RasaHQ/rasa/blob/3.6.x/LICENSE.txt
Data Source Authors: Rasa Technologies Inc
AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
gemini-dataset-rasalgethi-ptptcommand-generation-calm-v2
Rasa CALM Command Generator dataset (v2)
This is an instruction tuning dataset consisting of prompt-command pairs. These pairs can be used to train a small LLM like
Llama 3.1 8b to act as a
command generator in the CALM paradigm.
The technical details of how a CALM assistant works can be found in this paper.
Dataset Details
Dataset Description
The dataset consists prompt-command pairs, where prompt consists of an instruction for the LLM to follow in… See the full description on the dataset page: https://huggingface.co/datasets/rasa/command-generation-calm-v2.rasa-malayalam-nano-codecPersian-Texts-DatasetRasarasa_appointmentrasa-nepali-simplefrenchrasa-standardSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: rasa
Data Source Link: https://rasa.com/docs/
Data Source License: https://github.com/RasaHQ/rasa/blob/3.6.x/LICENSE.txt
Data Source Authors: Rasa Technologies Inc
AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
spellingcorrectionFrenchrasa-malayalam-processedCompiled-rasa-dataset
