datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task544_alt_translation_hi_en
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task544_alt_translation_hi_en
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task544_alt_translation_hi_en.task427_hindienglish_corpora_hi-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.task558_alt_translation_en_hi
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task558_alt_translation_en_hi
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task558_alt_translation_en_hi.task434_alt_en_hi_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task434_alt_en_hi_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task434_alt_en_hi_answer_generation.task433_alt_hi_en_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task433_alt_hi_en_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task433_alt_hi_en_translation.task1329_open_subtitles_en_hi_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1329_open_subtitles_en_hi_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1329_open_subtitles_en_hi_translation.task424_hindienglish_corpora_hi_en_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task424_hindienglish_corpora_hi_en_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task424_hindienglish_corpora_hi_en_translation.task432_alt_en_hi_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task432_alt_en_hi_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task432_alt_en_hi_translation.task1323_open_subtitles_hi_en_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1323_open_subtitles_hi_en_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1323_open_subtitles_hi_en_translation.Vietnam-History-200K-ENEnglish 200,000-sample Vietnamese history dataset in the same fine-tuning format (with ≈78% reasoning and ≈22% final-only).
Format & coverage
Language: English
Scope: 905–2025 (events, figures, dynasties, wars, reforms, culture, documents)
Structure: messages (ShareGPT/ChatML style)
With reasoning (≈78%): system → user → assistant (analysis) → assistant (final)
Final-only (≈22%): system → user → assistant (final)
Learn more on GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/minhxthanh/Vietnam-History-200K-EN.Vietnam-History-500K-EnEnglish 500,000-sample Vietnamese history dataset, with ≈78% chain-of-thought (analysis) and ≈22% final-only answers.
Format & coverage
Scope: 905–2025 (events, figures, dynasties, wars, reforms, culture, documents)
Structure: ShareGPT/ChatML-style messages
With reasoning (≈78%): system → user → assistant (analysis) → assistant (final)
Final-only (≈22%): system → user → assistant (final)
Learn more on GitHub: https://github.com/MinhxThanh/Vietnam-History-Chat-Datasets
task426_hindienglish_corpora_hi-en_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task426_hindienglish_corpora_hi-en_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task426_hindienglish_corpora_hi-en_classification.task1353_hind_encorp_translation_en_hi
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1353_hind_encorp_translation_en_hi
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1353_hind_encorp_translation_en_hi.VietNam-History-100K_ENLearn more on GitHub: https://github.com/MinhxThanh/Vietnam-History-Chat-Datasets
task425_hindienglish_corpora_en_hi_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task425_hindienglish_corpora_en_hi_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task425_hindienglish_corpora_en_hi_translation.en_hi_bhoRef: https://github.com/ltrc/mini-project-slm-PranavAga
Vietnam-History-1M-EnLearn more on GitHub: https://github.com/MinhxThanh/Vietnam-History-Chat-Datasets
Vietnamese History Q&A – English – 1,000,000 samples
Size: 1M conversations (JSONL, gzip)Language: EnglishDomain: Vietnamese history 905–2025 (events, figures, dynasties, wars, reforms, documents)Format: ShareGPT/ChatML-style messages with assistant channels analysis (reasoning) and final (answer).
Record
{
"messages": [
{"role":"system","content":"…"},
{"role":"user"… See the full description on the dataset page: https://huggingface.co/datasets/minhxthanh/Vietnam-History-1M-En.alpaca-gpt4-en-high-prob-qwen-0.5b-10k
High-Probability Sentence Predictions Dataset
Dataset Description
This dataset contains sentences from llamafactory/alpaca_gpt4_en
where the model Qwen/Qwen2.5-0.5B predicts the token before the final period
with ≥90% probability.
Source Dataset Attribution
This dataset is derived from llamafactory/alpaca_gpt4_en
and inherits its license terms (apache-2.0). Please cite the original dataset when using this data.
Extraction Parameters
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/ermiaazarkhalili/alpaca-gpt4-en-high-prob-qwen-0.5b-10k.task1352_hind_encorp_translation_hi_en
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1352_hind_encorp_translation_hi_en
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1352_hind_encorp_translation_hi_en.
