CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SherlockRamos /jurisdb-legal-documents JurisDB - Brazilian Legal Documents Dataset Dataset Description This dataset contains a comprehensive collection of Brazilian legal documents, including legislation (federal and state laws) and jurisprudence (court rulings and summaries from TSE, STJ, STF, and TNU). Dataset Structure . ├── legislacao_grifada_e_anotada_atualiz_em_01_01_2026/ │ ├── leis_estaduais/ │ ├── leis_federais/ │ └── ... └── sumulas_tse_stj_stf_e_tnu_atualiz_01_01_2026_2/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents.documenttext-classificationn<1K0 likes1.7k downloads9mo agoHugging Face02dartbrains /sherlock Sherlock Naturalistic fMRI dataset: 16 subjects watched ~50 minutes of Sherlock across two scanning runs (Part1, Part2) and then verbally recalled the narrative in the scanner. TR = 1.5 s. This repo mirrors the fmriprep-preprocessed dataset originally distributed via DataLad at https://gin.g-node.org/ljchang/Sherlock. fmriprep version 1.2.6-1. Layout derivatives/fmriprep/sub-XX/ anat/ func/ figures/ log/ onsets/ Sherlock_Crop_Onsets.csv… See the full description on the dataset page: https://huggingface.co/datasets/dartbrains/sherlock.imagefeature-extractionn<1K0 likes1.2k downloads3mo agoHugging Face03Sherlock-Comms /wikipedia-en-2026-07-01-passages English Wikipedia Passages, Chunked (2026-07-01) Every English Wikipedia article split into retrieval-sized passages with title and section attached. A clean, dated corpus for RAG — embed it yourself, or use the ready-made vectors and indexes in the companion repos: embeddings · faiss. Contents 17,473,199 passages, 21 GB, JSON Lines (one passage per line). Fields: id (<pageid>#<n>), title, section, text. Line order matches ids.txt / vector row order in the… See the full description on the dataset page: https://huggingface.co/datasets/Sherlock-Comms/wikipedia-en-2026-07-01-passages.texttext-retrieval10M<n<100M0 likes307 downloads2mo agoHugging Face04Sherlock-Comms /wikipedia-en-2026-07-01-qwen3-embed-4b English Wikipedia Passage Embeddings — Qwen3-Embedding-4B (2026-07-01) Dense vectors for every passage of English Wikipedia, ready for retrieval- augmented generation. Publishing these saves ~63 GPU-hours of embedding. Pairs with the passage text at wikipedia-en-2026-07-01-passages and prebuilt FAISS indexes at wikipedia-en-2026-07-01-faiss. What is here 17,473,199 passage vectors, 1024-dim, float16, L2-normalized. 34 GB across 57 .npy shards (vecs_gpu*.npy), one… See the full description on the dataset page: https://huggingface.co/datasets/Sherlock-Comms/wikipedia-en-2026-07-01-qwen3-embed-4b.textfeature-extraction10M<n<100M0 likes284 downloads2mo agoHugging Face05open-llm-leaderboard-old /details_SherlockAssistant__Mistral-7B-Instruct-Ukrainian Dataset Card for Evaluation run of SherlockAssistant/Mistral-7B-Instruct-Ukrainian Dataset automatically created during the evaluation run of model SherlockAssistant/Mistral-7B-Instruct-Ukrainian on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_SherlockAssistant__Mistral-7B-Instruct-Ukrainian.0 likes91 downloads3y agoHugging Face06derogab /Sherlock-Case-Files Sherlock Case Files 📁 Sherlock Case Files is a synthetic multilingual dataset for schema-guided information extraction. Each case asks a model to read a compact JSON schema and a text, then return exactly one JSON object matching that schema. The dataset covers short snippets and long documents across varied domains and formats. It includes distractors and missing fields, represented by null, in English, Italian, Spanish, French, Portuguese, and German. Metadata supports… See the full description on the dataset page: https://huggingface.co/datasets/derogab/Sherlock-Case-Files.tabulartext-generationn<1K0 likes76 downloads22d agoHugging Face07srinivasbilla /tiny-sherlock-audioTest Audio Dataset, 12 Hrs of Sherlock Audio Book. sourced from https://www.digitalbook.io/audiobook/57e3ff81e7350a5135de3a3ba60779af/Adventures%20of%20Sherlock%20Holmes audion<1K7 likes63 downloads2y agoHugging Face08sherlockvn /MEPC Multi-level Product Category Recognition Image Dataset Summary Wordcloud Introduce MEPC - 1000 Dataset: Classes: 1000 Images: 164,117 Train: 131,293 Val: 32824 MEPC - 10 Dataset: Classes: 10 Images: 2,192 Train: 1,753 Val: 439 Statistics Statistics of the number of multi-level categories in the two datasets MEPC-10 and MEPC-1000 Label-only embeddings visualizing label connections… See the full description on the dataset page: https://huggingface.co/datasets/sherlockvn/MEPC.image100K<n<1M0 likes61 downloads1y agoHugging Face09Sherlock-Comms /wikipedia-en-2026-07-01-faiss English Wikipedia FAISS Indexes (2026-07-01) Prebuilt FAISS indexes over the 17,473,199 Qwen3-Embedding-4B vectors, so you can query English Wikipedia locally without embedding or building anything. Companion repos: embeddings · passages. Files ivfpq.faiss — 1.33 GB, compressed. OPQ64,IVF16384,PQ64x8, inner product. ~64 bytes/vector. Best when RAM is tight; re-rank its hits against the raw vectors or a text reranker for full accuracy. hnsw_sq.faiss — 38 GB… See the full description on the dataset page: https://huggingface.co/datasets/Sherlock-Comms/wikipedia-en-2026-07-01-faiss.texttext-retrieval10M<n<100M0 likes47 downloads2mo agoHugging Face10youngermax /sherlock Dataset Card for "sherlock" More Information needed text10K<n<100K0 likes44 downloads3y agoHugging Face11SherlockMa /ControlUTR_training_dataData used for training ControlUTR. All stored in Apache format. Usage order: 5rna_pretrain -> 5rna_stage2 -> 5rna_stage2-2 -> 5rna_stage2-3 -> 5rna_stage2-4-3-merge -> 5rna_stage2-6-4-merge 5rna_stage3-1 Code: https://github.com/sherlockma11/ControlUTR Dataset: https://huggingface.co/datasets/SherlockMa/ControlUTR_training_data Model: https://huggingface.co/SherlockMa/ControlUTR texttext-generation1M<n<10M1 likes39 downloads8mo agoHugging Face12SherlockRamos /hf_hashcattext0 likes35 downloads9mo agoHugging Face13SherlockRamos /jurisdb-legal-documents-v2annotations_creators: found language_creators: found language: pt license: mit multilinguality: monolingual size_categories: 100K<n<1M task_categories: question-answering text-classification retrieval-augmented-generation pretty_name: JurisDB Brazilian Legal Documents config_names: default legislation jurisprudence tags: legal jurisprudence brazilian-law legislation domain:legal region:brazil JurisDB - Brazilian Legal Documents Dataset (v2) Dataset Summary This… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents-v2.text1K<n<10K0 likes35 downloads8mo agoHugging Face14TeichAI /sherlock-thinking-alpha-11000xThis dataset is unique in the sense that it is a non-reasoning dataset that was generated by a reasoning model (the stealth model that turned into grok 4.1 fast) The prompts from this dataset were generated by multiple models across the following domains: Health Legal Programming Marketing Academia Finance Science text10K<n<100K4 likes31 downloads10mo agoHugging Face15SherlockYoung /mini-monster_hunterimagen<1K1 likes30 downloads3y agoHugging Face16Alleinzellgaenger /sherlock-holmes-qa Sherlock Holmes Q&A Dataset A question-answering dataset for retrieval-augmented generation (RAG) over Sherlock Holmes short stories. Dataset Structure { "question": "What deduction did Holmes make?", "answer": "Holmes observed...", "story_id": "a_scandal_in_bohemia", "story_title": "A SCANDAL IN BOHEMIA" } Usage from datasets importload_dataset dataset = load_dataset("Alleinzellgaenger/sherlock-holmes-qa") Source Generated using… See the full description on the dataset page: https://huggingface.co/datasets/Alleinzellgaenger/sherlock-holmes-qa.textquestion-answeringn<1K0 likes29 downloads11mo agoHugging Face17swaghjal /sherlock-traintext100K<n<1M0 likes20 downloads3y agoHugging Face18TeichAI /sherlock-dash-alpha-1000xtext1K<n<10K0 likes19 downloads10mo agoHugging Face19lmassaron /Sherlock_QA_testtextn<1K0 likes17 downloads1y agoHugging Face20Alleinzellgaenger /sherlock-holmes-corpus Sherlock Holmes Corpus Full-text corpus of 55 Sherlock Holmes short stories for retrieval-augmented generation (RAG). Dataset Structure { "id": "a_scandal_in_bohemia", "title": "A SCANDAL IN BOHEMIA", "collection": "The Adventures of Sherlock Holmes", "content": "To Sherlock Holmes she is always _the_ woman..." } Usage from datasets importload_dataset corpus = load_dataset("Alleinzellgaenger/sherlock-holmes-corpus", split="train") Contents… See the full description on the dataset page: https://huggingface.co/datasets/Alleinzellgaenger/sherlock-holmes-corpus.texttext-retrievaln<1K0 likes17 downloads11mo agoHugging Face21sherlockab /LifeExpectancyDatatabular1K<n<10K0 likes17 downloads5mo agoHugging Face22sherlockab /accident-detection-from-cctv-footage0 likes16 downloads5mo agoHugging Face23elliot-mllm /sherlock_cleanedgated sherlock_cleaned The sherlock__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 13,546 QA turns 86,223 answers rewritten by the cleaning pass 0 QA created by the cleaning pass (new_qa) not measured for this family shards 33 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/sherlock_cleaned.imagevisual-question-answering0 likes16 downloads26d agoHugging Face24minnbanya /nlp-a2-sherlockA collection of Sir Arthur Conan Doyle's Sherlock Holmes books for NLP A2. text10K<n<100K1 likes15 downloads3y agoHugging Face25chayandatta /sherlock-holmes-corpus Sherlock Holmes Corpus 🕵️ A cleaned public domain corpus of Arthur Conan Doyle's Sherlock Holmes stories. Ideal for experimenting with retrieval, summarization, or fine-tuning small LMs. Dataset format: 1,234 paragraphs JSONL format with {"id": int, "text": str} text1K<n<10K0 likes15 downloads11mo agoHugging Face26akshayg08 /sherlock_preference_datasetThis dataset contains preference data for tuning Vision-Language models on the Sherlock Dataset for Abductive Reasoning. It is designed to evaluate the effectiveness of fine-tuning using Supervised Fine-Tuning (SFT) or Preference Optimization. Preferences are generated by prompting four models: mistralai/Pixtral-12B-2409, Qwen/Qwen2-VL-7B-Instruct, google/paligemma2-3b-ft-docci-448, and google/paligemma2-10b-ft-docci-448. Since this dataset is intended for optimizing PaLI-Gemma models… See the full description on the dataset page: https://huggingface.co/datasets/akshayg08/sherlock_preference_dataset.texttext-generation100K<n<1M0 likes14 downloads2y agoHugging Face27shabul /sherlock-debugger-datasettextn<1K0 likes14 downloads4mo agoHugging Face28mekongai /Identity-SherlockHolmestextn<1K2 likes13 downloads2y agoHugging Face29adityaxmittal /sherlock-datasettext100K<n<1M0 likes13 downloads3mo agoHugging Face30chengli-thu /Sherlock-Holmes-and-Thortextn<1K0 likes12 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.