datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Sherlock-Case-Files
Sherlock Case Files 📁
Sherlock Case Files is a synthetic multilingual dataset for schema-guided
information extraction. Each case asks a model to read a compact JSON schema and a text, then return exactly one JSON object matching that schema.
The dataset covers short snippets and long documents across varied domains and formats. It includes distractors and missing fields, represented by null, in English, Italian, Spanish, French, Portuguese, and German. Metadata supports… See the full description on the dataset page: https://huggingface.co/datasets/derogab/Sherlock-Case-Files.LifeExpectancyDatasherlock-annotated
Sherlock Column Type Annotations
Column-level type annotations for the Sherlock corpus, produced by the FineType distillation pipeline.
Dataset Description
Each row represents a single column from the Sherlock test set, annotated with:
Blind label — an LLM classification made without seeing FineType's prediction
FineType label — the prediction from FineType's CharCNN inference engine
Final label — adjudicated result (blind-first: the blind label is preferred unless… See the full description on the dataset page: https://huggingface.co/datasets/meridian-online/sherlock-annotated.creditcard_datasetsmoltiny_sherlock_whisper_snac_combined
