datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
japanese-casual-conversational-speech-golden-dataset-preview
Japanese Casual Conversational Speech Golden Dataset (Preview)
💼 Commercial License & Full Access
This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning.
To purchase the full dataset, please contact us:
👉 Email: info@hth-inc.com
👉 Website: https://hth-inc.com/business
🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.ragas-golden-dataset
Dataset Card for the ragas-golden-dataset
Dataset Description
The RAGAS Golden Dataset is a synthetically generated question-answering dataset designed for evaluating Retrieval Augmented Generation (RAG) systems. It contains high-quality question-answer pairs derived from academic papers on AI agents and agentic AI architectures.
Dataset Summary
This dataset was generated using Prefect and the RAGAS TestsetGenerator framework, which creates synthetic questions… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/ragas-golden-dataset.golden-dataset-2.1ragas-golden-dataset-documents
Dataset Card for RAGAS Golden Dataset Documents
A small, mixed‐format corpus to compare PDF, API, and web‐based document loader output from the LangChain ecosystem.
The code to run the Prefect prefect_docloader_pipeline.py pipeline is available in the RAGAS Golden Dataset Pipeline repository.
While several enhancements are planned for future iterations, the hands-on insights gained from this grassroots exploration of document loader behaviors proved too valuable -- things that… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/ragas-golden-dataset-documents.azure-ai-engineer-golden-datasetquery-intent-detection-golden-datasetgolden_sample_datasetragas-golden-dataset-v2
Dataset Card for the ragas-golden-dataset-v2
Dataset Description
The RAGAS Golden Dataset is a synthetically generated question-answering dataset designed for evaluating Retrieval Augmented Generation (RAG) systems. It contains high-quality question-answer pairs derived from academic papers on AI agents and agentic AI architectures.
Dataset Summary
This dataset was generated using Prefect and the RAGAS TestsetGenerator framework, which creates synthetic… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/ragas-golden-dataset-v2.FDA_Cybersecurity_Golden_DatasetHindiNER-golden-dataset-constraint1
Dataset Card for HindiNER-golden-dataset-constraint1
These dataset is a modified version of HindiNER-golden-dataset
Check out the Colab Notebook used to modify HindiNER-golden-dataset
golden_ner_dataset
Dataset Card for golden_ner_dataset
This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Using this dataset with Argilla
To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code:
import argilla as rg
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Eduardorpm24/golden_ner_dataset.HindiNER-golden-dataset-constraint3
Dataset Card for HindiNER-golden-dataset-constraint3
These dataset is a modified version of HindiNER-golden-dataset
Check out the Colab Notebook used to modify HindiNER-golden-dataset
HindiNER-golden-dataset-constraint4
Dataset Card for HindiNER-golden-dataset-constraint4
These dataset is a modified version of HindiNER-golden-dataset
Check out the Colab Notebook used to modify HindiNER-golden-dataset
HindiNER-golden-dataset2golden-dataset
Golden Dataset — Agricultural RAG Evaluation
A high-quality golden Q&A dataset for evaluating Retrieval-Augmented Generation (RAG) systems on Indian agricultural knowledge.
Dataset Summary
Metric
Value
Total Q&A pairs
1,675
Languages
English
Domain
Indian Agriculture
Sources
8 canonical sources
Question types
Factual, Conceptual, Procedural
Difficulty levels
Easy, Medium, Hard
Sources
Source
Description
Indian… See the full description on the dataset page: https://huggingface.co/datasets/AnmolNimmala0/golden-dataset.ragas-golden-dataset-colabHindiNER-golden-dataset
Dataset Card for HindiNER-golden-dataset
HindiNER-golden-dataset - a small, diverse and high quality general Hindi NER dataset
Dataset Details
Dataset Description
The HindiNER-golden-dataset includes 952 diverse source texts sampled from nisram-hindi-text-0.0 dataset. Their labels were generated using Llama-3.3-70B-Instruct and then manually corrected twice.
Curated by: nis12ram
Language(s) (NLP): Hindi
License: Apache License 2.0
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nis12ram/HindiNER-golden-dataset.fda_samd_regulations_golden_test_datasetHindiNER-golden-dataset-constraint2
Dataset Card for HindiNER-golden-dataset-constraint2
These dataset is a modified version of HindiNER-golden-dataset
Check out the Colab Notebook used to modify HindiNER-golden-dataset
HindiNER-golden-dataset-constraint5
Dataset Card for HindiNER-golden-dataset-constraint5
These dataset is a modified version of HindiNER-golden-dataset
Check out the Colab Notebook used to modify HindiNER-golden-dataset
middle_eastern_mixed_food_dataset_goldenGolden-datasetGolden_Nutrition_DatasetHindiNER-golden-dataset-constraint-neg-corrAlita-Final-Golden-Dataset-V5HindiNER-golden-eval-dataset3spk-cetuc-golden-datasetelixir-golden-dataset
