CoolFace
Datasetpublic

marin-community/identity-data

Identity Data English synthetic conversations for reinforcing model identity and provenance. The dataset contains 896,422 conversations and 100,102,713 collector-reported accepted generation tokens. Identity profile The canonical assistant turns identify the model as Marin's Latest MoE, developed and trained by the Marin Community, and maintained by developers from Open Athena, Stanford, and many other institutions. Some conversations acknowledge contributions… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/identity-data.

sourceHugging Faceupdated 2mo agoView on Hugging Face
1likes192downloads
../
filetrain-00000-of-00010.parquet10.6 MBdownload
filetrain-00001-of-00010.parquet10.6 MBdownload
filetrain-00002-of-00010.parquet10.6 MBdownload
filetrain-00003-of-00010.parquet10.6 MBdownload
filetrain-00004-of-00010.parquet10.7 MBdownload
filetrain-00005-of-00010.parquet10.6 MBdownload
filetrain-00006-of-00010.parquet10.6 MBdownload
filetrain-00007-of-00010.parquet10.6 MBdownload
filetrain-00008-of-00010.parquet10.7 MBdownload
filetrain-00009-of-00010.parquet10.6 MBdownload

marin-community/identity-data · main · files are served by the source, never re-hosted here