datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sherlock
Dataset Card for "sherlock"
More Information needed
jurisdb-legal-documents-v2annotations_creators: found
language_creators: found
language: pt
license: mit
multilinguality: monolingual
size_categories: 100K<n<1M
task_categories:
question-answering
text-classification
retrieval-augmented-generation
pretty_name: JurisDB Brazilian Legal Documents
config_names:
default
legislation
jurisprudence
tags:
legal
jurisprudence
brazilian-law
legislation
domain:legal
region:brazil
JurisDB - Brazilian Legal Documents Dataset (v2)
Dataset Summary
This… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents-v2.sherlock-holmes-qa
Sherlock Holmes Q&A Dataset
A question-answering dataset for retrieval-augmented generation (RAG) over Sherlock Holmes short stories.
Dataset Structure
{
"question": "What deduction did Holmes make?",
"answer": "Holmes observed...",
"story_id": "a_scandal_in_bohemia",
"story_title": "A SCANDAL IN BOHEMIA"
}
Usage
from datasets importload_dataset
dataset = load_dataset("Alleinzellgaenger/sherlock-holmes-qa")
Source
Generated using… See the full description on the dataset page: https://huggingface.co/datasets/Alleinzellgaenger/sherlock-holmes-qa.Sherlock_QA_testsherlock-holmes-corpus
Sherlock Holmes Corpus
Full-text corpus of 55 Sherlock Holmes short stories for retrieval-augmented generation (RAG).
Dataset Structure
{
"id": "a_scandal_in_bohemia",
"title": "A SCANDAL IN BOHEMIA",
"collection": "The Adventures of Sherlock Holmes",
"content": "To Sherlock Holmes she is always _the_ woman..."
}
Usage
from datasets importload_dataset
corpus = load_dataset("Alleinzellgaenger/sherlock-holmes-corpus", split="train")
Contents… See the full description on the dataset page: https://huggingface.co/datasets/Alleinzellgaenger/sherlock-holmes-corpus.sherlock-debugger-datasetIdentity-SherlockHolmesSherlock_txtsherlock-annotated
Sherlock Column Type Annotations
Column-level type annotations for the Sherlock corpus, produced by the FineType distillation pipeline.
Dataset Description
Each row represents a single column from the Sherlock test set, annotated with:
Blind label — an LLM classification made without seeing FineType's prediction
FineType label — the prediction from FineType's CharCNN inference engine
Final label — adjudicated result (blind-first: the blind label is preferred… See the full description on the dataset page: https://huggingface.co/datasets/meridian-online/sherlock-annotated.Sherlock_QAsmoldrone-nav-synthetic-v1complete_sherlock_holmes-book
Dataset Card for "complete_sherlock_holmes-book"
More Information needed
sherlock_gemmasherlock
