datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jurisdb-legal-documents
JurisDB - Brazilian Legal Documents Dataset
Dataset Description
This dataset contains a comprehensive collection of Brazilian legal documents, including legislation (federal and state laws) and jurisprudence (court rulings and summaries from TSE, STJ, STF, and TNU).
Dataset Structure
.
├── legislacao_grifada_e_anotada_atualiz_em_01_01_2026/
│ ├── leis_estaduais/
│ ├── leis_federais/
│ └── ...
└── sumulas_tse_stj_stf_e_tnu_atualiz_01_01_2026_2/
├──… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents.sherlock
Sherlock
Naturalistic fMRI dataset: 16 subjects watched ~50 minutes of Sherlock across
two scanning runs (Part1, Part2) and then verbally recalled the narrative in
the scanner. TR = 1.5 s.
This repo mirrors the fmriprep-preprocessed dataset originally distributed via
DataLad at https://gin.g-node.org/ljchang/Sherlock. fmriprep version 1.2.6-1.
Layout
derivatives/fmriprep/sub-XX/
anat/ func/ figures/ log/
onsets/
Sherlock_Crop_Onsets.csv… See the full description on the dataset page: https://huggingface.co/datasets/dartbrains/sherlock.wikipedia-en-2026-07-01-passages
English Wikipedia Passages, Chunked (2026-07-01)
Every English Wikipedia article split into retrieval-sized passages with title
and section attached. A clean, dated corpus for RAG — embed it yourself, or use
the ready-made vectors and indexes in the companion repos:
embeddings
·
faiss.
Contents
17,473,199 passages, 21 GB, JSON Lines (one passage per line).
Fields: id (<pageid>#<n>), title, section, text.
Line order matches ids.txt / vector row order in the… See the full description on the dataset page: https://huggingface.co/datasets/Sherlock-Comms/wikipedia-en-2026-07-01-passages.wikipedia-en-2026-07-01-qwen3-embed-4b
English Wikipedia Passage Embeddings — Qwen3-Embedding-4B (2026-07-01)
Dense vectors for every passage of English Wikipedia, ready for retrieval-
augmented generation. Publishing these saves ~63 GPU-hours of embedding.
Pairs with the passage text at
wikipedia-en-2026-07-01-passages
and prebuilt FAISS indexes at
wikipedia-en-2026-07-01-faiss.
What is here
17,473,199 passage vectors, 1024-dim, float16, L2-normalized.
34 GB across 57 .npy shards (vecs_gpu*.npy), one… See the full description on the dataset page: https://huggingface.co/datasets/Sherlock-Comms/wikipedia-en-2026-07-01-qwen3-embed-4b.details_SherlockAssistant__Mistral-7B-Instruct-Ukrainian
Dataset Card for Evaluation run of SherlockAssistant/Mistral-7B-Instruct-Ukrainian
Dataset automatically created during the evaluation run of model SherlockAssistant/Mistral-7B-Instruct-Ukrainian on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_SherlockAssistant__Mistral-7B-Instruct-Ukrainian.Sherlock-Case-Files
Sherlock Case Files 📁
Sherlock Case Files is a synthetic multilingual dataset for schema-guided
information extraction. Each case asks a model to read a compact JSON schema and a text, then return exactly one JSON object matching that schema.
The dataset covers short snippets and long documents across varied domains and formats. It includes distractors and missing fields, represented by null, in English, Italian, Spanish, French, Portuguese, and German. Metadata supports… See the full description on the dataset page: https://huggingface.co/datasets/derogab/Sherlock-Case-Files.tiny-sherlock-audioTest Audio Dataset, 12 Hrs of Sherlock Audio Book.
sourced from https://www.digitalbook.io/audiobook/57e3ff81e7350a5135de3a3ba60779af/Adventures%20of%20Sherlock%20Holmes
MEPC
Multi-level Product Category Recognition Image Dataset
Summary
Wordcloud
Introduce
MEPC - 1000 Dataset:
Classes: 1000
Images: 164,117
Train: 131,293
Val: 32824
MEPC - 10 Dataset:
Classes: 10
Images: 2,192
Train: 1,753
Val: 439
Statistics
Statistics of the number of multi-level categories in the two datasets MEPC-10 and MEPC-1000
Label-only embeddings visualizing label connections… See the full description on the dataset page: https://huggingface.co/datasets/sherlockvn/MEPC.wikipedia-en-2026-07-01-faiss
English Wikipedia FAISS Indexes (2026-07-01)
Prebuilt FAISS indexes over the 17,473,199 Qwen3-Embedding-4B vectors, so you
can query English Wikipedia locally without embedding or building anything.
Companion repos:
embeddings
·
passages.
Files
ivfpq.faiss — 1.33 GB, compressed. OPQ64,IVF16384,PQ64x8, inner
product. ~64 bytes/vector. Best when RAM is tight; re-rank its hits against
the raw vectors or a text reranker for full accuracy.
hnsw_sq.faiss — 38 GB… See the full description on the dataset page: https://huggingface.co/datasets/Sherlock-Comms/wikipedia-en-2026-07-01-faiss.sherlock
Dataset Card for "sherlock"
More Information needed
ControlUTR_training_dataData used for training ControlUTR. All stored in Apache format.
Usage order:
5rna_pretrain ->
5rna_stage2 -> 5rna_stage2-2 -> 5rna_stage2-3 -> 5rna_stage2-4-3-merge -> 5rna_stage2-6-4-merge
5rna_stage3-1
Code: https://github.com/sherlockma11/ControlUTR
Dataset: https://huggingface.co/datasets/SherlockMa/ControlUTR_training_data
Model: https://huggingface.co/SherlockMa/ControlUTR
hf_hashcatjurisdb-legal-documents-v2annotations_creators: found
language_creators: found
language: pt
license: mit
multilinguality: monolingual
size_categories: 100K<n<1M
task_categories:
question-answering
text-classification
retrieval-augmented-generation
pretty_name: JurisDB Brazilian Legal Documents
config_names:
default
legislation
jurisprudence
tags:
legal
jurisprudence
brazilian-law
legislation
domain:legal
region:brazil
JurisDB - Brazilian Legal Documents Dataset (v2)
Dataset Summary
This… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents-v2.sherlock-thinking-alpha-11000xThis dataset is unique in the sense that it is a non-reasoning dataset that was generated by a reasoning model (the stealth model that turned into grok 4.1 fast)
The prompts from this dataset were generated by multiple models across the following domains:
Health
Legal
Programming
Marketing
Academia
Finance
Science
mini-monster_huntersherlock-holmes-qa
Sherlock Holmes Q&A Dataset
A question-answering dataset for retrieval-augmented generation (RAG) over Sherlock Holmes short stories.
Dataset Structure
{
"question": "What deduction did Holmes make?",
"answer": "Holmes observed...",
"story_id": "a_scandal_in_bohemia",
"story_title": "A SCANDAL IN BOHEMIA"
}
Usage
from datasets importload_dataset
dataset = load_dataset("Alleinzellgaenger/sherlock-holmes-qa")
Source
Generated using… See the full description on the dataset page: https://huggingface.co/datasets/Alleinzellgaenger/sherlock-holmes-qa.sherlock-trainsherlock-dash-alpha-1000xSherlock_QA_testsherlock-holmes-corpus
Sherlock Holmes Corpus
Full-text corpus of 55 Sherlock Holmes short stories for retrieval-augmented generation (RAG).
Dataset Structure
{
"id": "a_scandal_in_bohemia",
"title": "A SCANDAL IN BOHEMIA",
"collection": "The Adventures of Sherlock Holmes",
"content": "To Sherlock Holmes she is always _the_ woman..."
}
Usage
from datasets importload_dataset
corpus = load_dataset("Alleinzellgaenger/sherlock-holmes-corpus", split="train")
Contents… See the full description on the dataset page: https://huggingface.co/datasets/Alleinzellgaenger/sherlock-holmes-corpus.LifeExpectancyDataaccident-detection-from-cctv-footagesherlock_cleaned
sherlock_cleaned
The sherlock__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
13,546
QA turns
86,223
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
33
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/sherlock_cleaned.nlp-a2-sherlockA collection of Sir Arthur Conan Doyle's Sherlock Holmes books for NLP A2.
sherlock-holmes-corpus
Sherlock Holmes Corpus 🕵️
A cleaned public domain corpus of Arthur Conan Doyle's Sherlock Holmes stories.
Ideal for experimenting with retrieval, summarization, or fine-tuning small LMs.
Dataset format:
1,234 paragraphs
JSONL format with {"id": int, "text": str}
sherlock_preference_datasetThis dataset contains preference data for tuning Vision-Language models on the Sherlock Dataset for Abductive Reasoning. It is designed to evaluate the effectiveness of fine-tuning using Supervised Fine-Tuning (SFT) or Preference Optimization. Preferences are generated by prompting four models: mistralai/Pixtral-12B-2409, Qwen/Qwen2-VL-7B-Instruct, google/paligemma2-3b-ft-docci-448, and google/paligemma2-10b-ft-docci-448.
Since this dataset is intended for optimizing PaLI-Gemma models… See the full description on the dataset page: https://huggingface.co/datasets/akshayg08/sherlock_preference_dataset.sherlock-debugger-datasetIdentity-SherlockHolmessherlock-datasetSherlock-Holmes-and-Thor
