datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLM-self-identification Self Identification – Give your Language model an identity
About Self Identification
Self identification is training set SupraLabs curated for developers/trainers to experiment with to let your language model know about their identity.
Self identification let your LM know these information about them:
Model ID
Model Name
Model Description
Model Creator
Model Family
Model Architecture
Parameter Count
Knowledge Cutoff
Here is an example from the dataset:
If… See the full description on the dataset page: https://huggingface.co/datasets/SupraLabs/LLM-self-identification.Vertex-0.6-35M-self-identification
Vertex 0.6 35M — Self Identification
A self-identification SFT dataset for Vertex-0.6-35M-Instruct:
459 ChatML-style conversations that teach the model who it is — its name,
creator, family, architecture, parameter count, and knowledge cutoff.
Made from SupraLabs/LLM-self-identification
(Apache-2.0), with every {{SELF_ID.*}} marker replaced with the Vertex 0.6
35M identity:
Marker
Value
MODEL_ID
VertexResearch/Vertex-0.6-35M-Instruct
MODEL_NAME
Vertex 0.6 35M… See the full description on the dataset page: https://huggingface.co/datasets/VertexResearch/Vertex-0.6-35M-self-identification.Identification-of-paraphrasing
🇰🇿 Identification of Paraphrasing in Kazakh Context
Dataset Summary
Identification of Paraphrasing in Kazakh Context is a targeted dataset designed to train Large Language Models (LLMs) and embeddings to detect semantic equivalence between two distinct Kazakh texts.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
2,000
Total Words (approx.)
184,465
Avg. Words per Sample
92
Word Count… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Identification-of-paraphrasing.
