sherlock
Datasets
All datasets matching “sherlock”jurisdb-legal-documents
JurisDB - Brazilian Legal Documents Dataset
Dataset Description
This dataset contains a comprehensive collection of Brazilian legal documents, including legislation (federal and state laws) and jurisprudence (court rulings and summaries from TSE, STJ, STF, and TNU).
Dataset Structure
.
├── legislacao_grifada_e_anotada_atualiz_em_01_01_2026/
│ ├── leis_estaduais/
│ ├── leis_federais/
│ └── ...
└── sumulas_tse_stj_stf_e_tnu_atualiz_01_01_2026_2/
├──… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents.sherlock
Sherlock
Naturalistic fMRI dataset: 16 subjects watched ~50 minutes of Sherlock across
two scanning runs (Part1, Part2) and then verbally recalled the narrative in
the scanner. TR = 1.5 s.
This repo mirrors the fmriprep-preprocessed dataset originally distributed via
DataLad at https://gin.g-node.org/ljchang/Sherlock. fmriprep version 1.2.6-1.
Layout
derivatives/fmriprep/sub-XX/
anat/ func/ figures/ log/
onsets/
Sherlock_Crop_Onsets.csv… See the full description on the dataset page: https://huggingface.co/datasets/dartbrains/sherlock.wikipedia-en-2026-07-01-passages
English Wikipedia Passages, Chunked (2026-07-01)
Every English Wikipedia article split into retrieval-sized passages with title
and section attached. A clean, dated corpus for RAG — embed it yourself, or use
the ready-made vectors and indexes in the companion repos:
embeddings
·
faiss.
Contents
17,473,199 passages, 21 GB, JSON Lines (one passage per line).
Fields: id (<pageid>#<n>), title, section, text.
Line order matches ids.txt / vector row order in the… See the full description on the dataset page: https://huggingface.co/datasets/Sherlock-Comms/wikipedia-en-2026-07-01-passages.wikipedia-en-2026-07-01-qwen3-embed-4b
English Wikipedia Passage Embeddings — Qwen3-Embedding-4B (2026-07-01)
Dense vectors for every passage of English Wikipedia, ready for retrieval-
augmented generation. Publishing these saves ~63 GPU-hours of embedding.
Pairs with the passage text at
wikipedia-en-2026-07-01-passages
and prebuilt FAISS indexes at
wikipedia-en-2026-07-01-faiss.
What is here
17,473,199 passage vectors, 1024-dim, float16, L2-normalized.
34 GB across 57 .npy shards (vecs_gpu*.npy), one… See the full description on the dataset page: https://huggingface.co/datasets/Sherlock-Comms/wikipedia-en-2026-07-01-qwen3-embed-4b.details_SherlockAssistant__Mistral-7B-Instruct-Ukrainian
Dataset Card for Evaluation run of SherlockAssistant/Mistral-7B-Instruct-Ukrainian
Dataset automatically created during the evaluation run of model SherlockAssistant/Mistral-7B-Instruct-Ukrainian on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_SherlockAssistant__Mistral-7B-Instruct-Ukrainian.Sherlock-Case-Files
Sherlock Case Files 📁
Sherlock Case Files is a synthetic multilingual dataset for schema-guided
information extraction. Each case asks a model to read a compact JSON schema and a text, then return exactly one JSON object matching that schema.
The dataset covers short snippets and long documents across varied domains and formats. It includes distractors and missing fields, represented by null, in English, Italian, Spanish, French, Portuguese, and German. Metadata supports… See the full description on the dataset page: https://huggingface.co/datasets/derogab/Sherlock-Case-Files.
