datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLM-self-identification Self Identification – Give your Language model an identity
About Self Identification
Self identification is training set SupraLabs curated for developers/trainers to experiment with to let your language model know about their identity.
Self identification let your LM know these information about them:
Model ID
Model Name
Model Description
Model Creator
Model Family
Model Architecture
Parameter Count
Knowledge Cutoff
Here is an example from the dataset:
If… See the full description on the dataset page: https://huggingface.co/datasets/SupraLabs/LLM-self-identification.LLM-self-identification
LLM Identity · Give your LLM an identity
Self Identification
The Self-Identification Dataset, curated by Qyrou, is a specialized training resource designed to help developers and trainers establish clear self-identity awareness within language models. By incorporating this dataset, models can accurately learn and convey essential metadata about themselves, including their Model ID, Model Name, Model Description, Model Creator, Model Family, Model Architecture, Parameter Count… See the full description on the dataset page: https://huggingface.co/datasets/Qyrou/LLM-self-identification.Vertex-0.6-35M-self-identification
Vertex 0.6 35M — Self Identification
A self-identification SFT dataset for Vertex-0.6-35M-Instruct:
459 ChatML-style conversations that teach the model who it is — its name,
creator, family, architecture, parameter count, and knowledge cutoff.
Made from SupraLabs/LLM-self-identification
(Apache-2.0), with every {{SELF_ID.*}} marker replaced with the Vertex 0.6
35M identity:
Marker
Value
MODEL_ID
VertexResearch/Vertex-0.6-35M-Instruct
MODEL_NAME
Vertex 0.6 35M… See the full description on the dataset page: https://huggingface.co/datasets/VertexResearch/Vertex-0.6-35M-self-identification.repro-box-thirding-anytime-best-arm-identification-under-insufficient-sampling
Box Thirding (B3): Anytime Best Arm Identification under Insufficient Sampling
Reproduction of ICML 2026 paper (OpenReview: XoONWh8fbL)
Tags
trackio
trackio-logbook
open-experiment
icml2026-repro
paper-XoONWh8fbL
Horse-Individual-Identification-Dataset
Horse Individual Identification Dataset
In the agricultural field, individual identification of horses is crucial for farm management and animal health monitoring. Currently, many farms face issues such as low identification efficiency and inaccurate data in horse management. Existing solutions often rely on manual labeling, which is time-consuming and error-prone. This dataset aims to address accuracy and efficiency issues in object detection by providing high-quality images of… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Horse-Individual-Identification-Dataset.Asparagus-Identification-Dataset
Asparagus Identification Dataset
The current agricultural industry faces challenges in efficient crop monitoring and recognition. Traditional manual detection methods are inefficient and prone to errors. Existing solutions often rely on empirical judgment without scientific model support. This dataset aims to provide diverse asparagus images to help train automatic recognition models, improving the accuracy and efficiency of crop monitoring. The dataset includes images of asparagus… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Asparagus-Identification-Dataset.sapient-synth-flan-niv2-fsopt-data-task265-paper-reviews-language-identification
sapient-synth-flan-niv2-fsopt-data-task265-paper-reviews-language-identification
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 1341
Task: synthetic anonymous instruction replacement
Generation… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task265-paper-reviews-language-identification.Vulnerability_IdentificationSpam-Identificationsapient-synth-flan-niv2-zsopt-data-task265-paper-reviews-language-identification
sapient-synth-flan-niv2-zsopt-data-task265-paper-reviews-language-identification
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 710
Task: synthetic anonymous instruction replacement
Generation… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-zsopt-data-task265-paper-reviews-language-identification.lmsys-self-identification-new
