datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tool-use-llama-format
Open Paws Tool Use Llama Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Tool Use Data
Format: JSONL (JSON Lines)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open Paws
License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/tool-use-llama-format.visual-qa-llama-format
Open Paws Visual Qa Llama Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Multimodal Data
Format: JSONL (JSON Lines)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open Paws
License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/visual-qa-llama-format.mllm-shap
MLLM-SHAP experiment datasets
Curated test splits for studying Shapley-value explanations in multimodal large language models (text and audio inputs). Each configuration is a filtered, size-controlled subset built for reproducible benchmarking—not a full copy of the upstream corpora.
Configs follow the naming pattern {task}__{source} (for example single_sentence__voice_bench).
Quick load
Pin a dataset revision for reproducibility (replace REVISION with the commit hash… See the full description on the dataset page: https://huggingface.co/datasets/Pawlo77/mllm-shap.task400_paws_paraphrase_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task400_paws_paraphrase_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task400_paws_paraphrase_classification.task776_pawsx_japanese_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task776_pawsx_japanese_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task776_pawsx_japanese_text_modification.task770_pawsx_english_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task770_pawsx_english_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task770_pawsx_english_text_modification.claude-fable-5-code
Claude Fable 5 Coding and Math Dataset (Non-Thinking)
This repository contains a dataset of 603 coding and math-related prompts and responses from Claude Fable 5.
The generation of this dataset cost approximately $75.
Please note that this dataset is non-thinking. Fable 5 only supported adaptive thinking, and it decided not to think for these prompts, meaning there is no chain-of-thought/reasoning content in this dataset.
Origin of Prompts
The prompts in this… See the full description on the dataset page: https://huggingface.co/datasets/PawanKrd/claude-fable-5-code.task788_pawsx_korean_japanese_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task788_pawsx_korean_japanese_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task788_pawsx_korean_japanese_translation.task814_pawsx_japanese_korean_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task814_pawsx_japanese_korean_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task814_pawsx_japanese_korean_translation.task798_pawsx_spanish_german_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task798_pawsx_spanish_german_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task798_pawsx_spanish_german_translation.sanskrit-verses-gretil
Sanskrit Literature Source Retrieval Dataset (GRETIL)
This dataset contains 283,935 Sanskrit verses from the GRETIL (Göttingen Register of Electronic Texts in Indian Languages) corpus, designed for training language models on Sanskrit literature source identification.
Dataset Description
This dataset is created for a reinforcement learning task where models learn to identify the source of Sanskrit quotes, including:
Genre (kavya, epic, purana, veda, shastra, tantra… See the full description on the dataset page: https://huggingface.co/datasets/paws/sanskrit-verses-gretil.task815_pawsx_japanese_french_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task815_pawsx_japanese_french_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task815_pawsx_japanese_french_translation.task778_pawsx_english_french_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task778_pawsx_english_french_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task778_pawsx_english_french_translation.continued-pretraining-llama-format
Open Paws Continued Pretraining Llama Format
Overview
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Specialized Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.pawa-function-calling-60k
Swahili Translated Version of xlam-function-calling-60k
this is the version 0.1
small-lean
Small Lean Alpaca
Thirty filtered Alpaca-style Lean 4 theorem-proving records derived from
internlm/Lean-Workbook.
Only train.jsonl is a Hub dataset split. The hf_dataset/ directory is a
local datasets.save_to_disk() artifact and must not be interpreted as JSON
training data.
reasoning-and-chat-harmony-format
Open Paws Reasoning And Conversational Finetuning Harmony Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Reasoning Data
Format: JSONL (JSON Lines)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/reasoning-and-chat-harmony-format.conversational-finetuning-llama-format
Open Paws Conversational Finetuning Llama Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Training Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open Paws… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/conversational-finetuning-llama-format.dialectical-reasoningA specialised dialectical reasoning dataset.
contain { Thesis:, Antithesis:, Synthesis: }.
Domain are math, science, creative writing
task808_pawsx_chinese_korean_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task808_pawsx_chinese_korean_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task808_pawsx_chinese_korean_translation.reasoning-llama-format
Open Paws Reasoning Llama Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Reasoning Data
Format: JSONL (JSON Lines)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open Paws
License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/reasoning-llama-format.paws-x_fr_prompt_paraphrase_generation
paws-x_fr_prompt_paraphrase_generation
Summary
paws-x_fr_prompt_paraphrase_generation is a subset of the Dataset of French Prompts (DFP).It contains 562,728 rows that can be used for a paraphrase generation task.The original data (without prompts) comes from the dataset paws-x by Yang et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/paws-x_fr_prompt_paraphrase_generation.task782_pawsx_english_japanese_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task782_pawsx_english_japanese_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task782_pawsx_english_japanese_translation.task789_pawsx_french_english_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task789_pawsx_french_english_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task789_pawsx_french_english_translation.task810_pawsx_chinese_spanish_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task810_pawsx_chinese_spanish_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task810_pawsx_chinese_spanish_translation.task777_pawsx_english_korean_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task777_pawsx_english_korean_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task777_pawsx_english_korean_translation.task786_pawsx_korean_german_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task786_pawsx_korean_german_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task786_pawsx_korean_german_translation.task797_pawsx_spanish_french_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task797_pawsx_spanish_french_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task797_pawsx_spanish_french_translation.task813_pawsx_japanese_english_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task813_pawsx_japanese_english_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task813_pawsx_japanese_english_translation.task804_pawsx_german_spanish_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task804_pawsx_german_spanish_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task804_pawsx_german_spanish_translation.
