onyx
Datasets
All datasets matching “onyx”Urdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.EnterpriseRAG-Bench
EnterpriseRAG-Bench
A benchmark dataset of 500,000+ documents and 500 questions for evaluating RAG systems on realistic enterprise data.
See the latest leaderboard rankings. The paper is available on arXiv. For code, methodology, and evaluation tools, see the GitHub repository.
Overview
Existing RAG and IR datasets focus on publicly accessible document sets (Bing searches, Stack Overflow, etc.). EnterpriseRAG-Bench provides the first publicly accessible dataset… See the full description on the dataset page: https://huggingface.co/datasets/onyx-dot-app/EnterpriseRAG-Bench.FineWeb-Edu-Sample-BT100-1-4The 1-4 parquet files of the fineweb-edu sample bt100 mixed up and turned up to 1 jsonl file. This is for users who dont have a NASA PC.
onyx-app-releasesUrdu-ONYX-WAV-realUrdu-ONYX-WAV
Urdu-ONYX-WAV
Urdu-ONYX-WAV is a high-quality Urdu Text-to-Speech (TTS) dataset consisting of audio recordings and corresponding transcripts. This dataset has been specifically prepared for training TTS models and conducting research in Urdu speech synthesis.
📊 Dataset Structure
This dataset is distributed across multiple parts due to size constraints:
Main repository: Base dataset with initial samples
part2: Additional 2.56 GB of audio data (6 Arrow files)
part3:… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV.
