datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Sherlock-Case-Files
Sherlock Case Files 📁
Sherlock Case Files is a synthetic multilingual dataset for schema-guided
information extraction. Each case asks a model to read a compact JSON schema and a text, then return exactly one JSON object matching that schema.
The dataset covers short snippets and long documents across varied domains and formats. It includes distractors and missing fields, represented by null, in English, Italian, Spanish, French, Portuguese, and German. Metadata supports… See the full description on the dataset page: https://huggingface.co/datasets/derogab/Sherlock-Case-Files.ControlUTR_training_dataData used for training ControlUTR. All stored in Apache format.
Usage order:
5rna_pretrain ->
5rna_stage2 -> 5rna_stage2-2 -> 5rna_stage2-3 -> 5rna_stage2-4-3-merge -> 5rna_stage2-6-4-merge
5rna_stage3-1
Code: https://github.com/sherlockma11/ControlUTR
Dataset: https://huggingface.co/datasets/SherlockMa/ControlUTR_training_data
Model: https://huggingface.co/SherlockMa/ControlUTR
sherlock_preference_datasetThis dataset contains preference data for tuning Vision-Language models on the Sherlock Dataset for Abductive Reasoning. It is designed to evaluate the effectiveness of fine-tuning using Supervised Fine-Tuning (SFT) or Preference Optimization. Preferences are generated by prompting four models: mistralai/Pixtral-12B-2409, Qwen/Qwen2-VL-7B-Instruct, google/paligemma2-3b-ft-docci-448, and google/paligemma2-10b-ft-docci-448.
Since this dataset is intended for optimizing PaLI-Gemma models… See the full description on the dataset page: https://huggingface.co/datasets/akshayg08/sherlock_preference_dataset.
