datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
science_fiction_sentences_idwip raw, feel free to continue or pull request :
https://github.com/WdangTehSeeker3245/wip-science-fiction-sentences
Fictional_Persona_Dialogs_Anonymized-benchmark
Fictional_Persona_Dialogs_Anonymized Benchmark Dataset
Short Summary:
A 68-pair synthetic Question-Answering (QA) dataset derived from anonymized fictional dialogues, specifically designed for rigorous Retrieval-Augmented Generation (RAG) system evaluation. It isolates and demonstrates the critical impact of contextual information on LLM accuracy.
Introduction & Motivation:
This dataset addresses the need for a clean, bias-minimized benchmark to accurately… See the full description on the dataset page: https://huggingface.co/datasets/TPelc/Fictional_Persona_Dialogs_Anonymized-benchmark.FictionalCharactersHathiTrust_Post45_Fictionfiction_headlines_challenge_eval_setFictionRealLabelSet
FictionRealLabelSet
tags: classification, fiction, real, label, corpus
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description: The 'FictionRealLabelSet' dataset contains a curated collection of texts extracted from various sources. These texts have been meticulously classified into two categories: 'Fiction' and 'Non-fiction'. The dataset is intended for use in natural language processing (NLP) tasks that require distinguishing between… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FictionRealLabelSet.
