datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLM-self-identification Self Identification – Give your Language model an identity
About Self Identification
Self identification is training set SupraLabs curated for developers/trainers to experiment with to let your language model know about their identity.
Self identification let your LM know these information about them:
Model ID
Model Name
Model Description
Model Creator
Model Family
Model Architecture
Parameter Count
Knowledge Cutoff
Here is an example from the dataset:
If… See the full description on the dataset page: https://huggingface.co/datasets/SupraLabs/LLM-self-identification.LLM-self-identification
LLM Identity · Give your LLM an identity
Self Identification
The Self-Identification Dataset, curated by Qyrou, is a specialized training resource designed to help developers and trainers establish clear self-identity awareness within language models. By incorporating this dataset, models can accurately learn and convey essential metadata about themselves, including their Model ID, Model Name, Model Description, Model Creator, Model Family, Model Architecture, Parameter Count… See the full description on the dataset page: https://huggingface.co/datasets/Qyrou/LLM-self-identification.task427_hindienglish_corpora_hi-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.task112_asset_simple_sentence_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task112_asset_simple_sentence_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task112_asset_simple_sentence_identification.Vertex-0.6-35M-self-identification
Vertex 0.6 35M — Self Identification
A self-identification SFT dataset for Vertex-0.6-35M-Instruct:
459 ChatML-style conversations that teach the model who it is — its name,
creator, family, architecture, parameter count, and knowledge cutoff.
Made from SupraLabs/LLM-self-identification
(Apache-2.0), with every {{SELF_ID.*}} marker replaced with the Vertex 0.6
35M identity:
Marker
Value
MODEL_ID
VertexResearch/Vertex-0.6-35M-Instruct
MODEL_NAME
Vertex 0.6 35M… See the full description on the dataset page: https://huggingface.co/datasets/VertexResearch/Vertex-0.6-35M-self-identification.task441_eng_guj_parallel_corpus_gu-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task441_eng_guj_parallel_corpus_gu-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task441_eng_guj_parallel_corpus_gu-en_language_identification.Language_Identification_v1
Dataset Card for Language Identification Dataset
Dataset Summary
A comprehensive dataset for Indian language identification and text classification. The dataset contains text samples across 10 major Indian languages, making it suitable for developing language identification systems and multilingual NLP applications.
Languages and Distribution
Language Distribution:
Urdu 1000
Hindi 1000
Odia 1000
Tamil 1000
Kannada 1000
Bengali… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Language_Identification_v1.Identification-of-paraphrasing
🇰🇿 Identification of Paraphrasing in Kazakh Context
Dataset Summary
Identification of Paraphrasing in Kazakh Context is a targeted dataset designed to train Large Language Models (LLMs) and embeddings to detect semantic equivalence between two distinct Kazakh texts.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
2,000
Total Words (approx.)
184,465
Avg. Words per Sample
92
Word Count… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Identification-of-paraphrasing.task265_paper_reviews_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task265_paper_reviews_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task265_paper_reviews_language_identification.task533_europarl_es-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task533_europarl_es-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task533_europarl_es-en_language_identification.task562_alt_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task562_alt_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task562_alt_language_identification.task315_europarl_sv-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task315_europarl_sv-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task315_europarl_sv-en_language_identification.task1574_amazon_reviews_multi_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1574_amazon_reviews_multi_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1574_amazon_reviews_multi_language_identification.task1621_menyo20k-mt_en_yo_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1621_menyo20k-mt_en_yo_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1621_menyo20k-mt_en_yo_language_identification.
