CoolFace
Datasetpublic

marin-community/identity-data

Identity Data English synthetic conversations for reinforcing model identity and provenance. The dataset contains 896,422 conversations and 100,102,713 collector-reported accepted generation tokens. Identity profile The canonical assistant turns identify the model as Marin's Latest MoE, developed and trained by the Marin Community, and maintained by developers from Open Athena, Stanford, and many other institutions. Some conversations acknowledge contributions… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/identity-data.

sourceHugging Faceupdated 2mo agoView on Hugging Face
1likes189downloads
Dataset Card

Identity Data

English synthetic conversations for reinforcing model identity and provenance. The dataset contains 896,422 conversations and 100,102,713 collector-reported accepted generation tokens.

Identity profile

The canonical assistant turns identify the model as Marin's Latest MoE, developed and trained by the Marin Community, and maintained by developers from Open Athena, Stanford, and many other institutions.

Some conversations acknowledge contributions from NVIDIA and AI2 and data derived from open-weight models such as DeepSeek, Kimi, and GLM while preserving the model's distinct identity. Unspecified architecture, release, deployment, license, funding, and lineage details are answered as unknown.

Generation

User phrasing was generated with GLM-5.2 from programmatically composed identity facets, registers, social contexts, attitudes, and question strategies. Assistant identity answers were constrained to canonical text. The final clean production run completed on July 30, 2026.

Schema

  • —seed_id: deterministic generation seed identifier
  • —content: complete conversation rendered with User: and Assistant: labels
  • —messages: structured role/content messages
  • —facet: identity facet exercised by the conversation
  • —topic: requested identity topic
  • —truth_status: whether the profile supplies the requested fact
  • —profile_id: identity profile identifier
  • —accepted_tokens: generation-token count reported by the collector

Quality notes

False-premise questions are intentional and are followed by canonical corrections. Assistant answers deliberately have low lexical diversity. The dataset is synthetic, and a small fraction of conversations contain awkward or loosely matched user phrasing. It is intended as a small component of a larger pretraining mixture rather than a general-purpose conversational corpus.

License

No license is asserted for this dataset. Users should review applicable upstream model terms and their intended use.