tokens
Datasets
All datasets matching “tokens”obelics_seed2_tokensPart of the OBELISC data set, including 32 Million samples, please refer to dataset.py to use this data
mtpnet_tokens
模型训练过程汇总(持续更新中)
对于已收集的每一个模型,code 目录为模型定义、训练和测试的代码和脚本文件,model 目录为已收集的 epoch 模型文件,dataset.zip 为模型数据集。
下表汇总了所有收集的模型训练过程信息:
模型名称
模型简介
模型类型
Epoch数量
数据集信息
Clone-detection-BigCloneBench
基于大规模代码克隆基准数据集的代码克隆检测模型,任务是进行二元分类(0/1),其中1代表语义等价,0代表其他情况。
代码克隆检测
2个epoch
BigCloneBench数据集
Clone-detection-POJ-104
基于POJ-104数据集的代码克隆检测模型,任务是识别不同编程题目中相似的代码实现,给定一段代码和一组候选代码,任务是返回具有相同语义的Top K个代码
代码克隆检测
2个epoch (0-1)
POJ-104编程题目数据集… See the full description on the dataset page: https://huggingface.co/datasets/code-philia/mtpnet_tokens.tulu_flan_mds_incremental-tokensllava-video-178k-siglip-tokens-ftov-new
LLaVA-Video-178K SigLIP Token Cache (LLaVA-OV fine-tuned vision tower)
Derived data (vision-encoder features of video frames), not a
redistribution of the source videos. Source:
lmms-lab/LLaVA-Video-178K -- its card
restricts use to academic research and education, and its annotations come
from GPT-4-class models (see the OpenAI usage policy).
Complete: 85000 clips.
Subset
Folders: 0_30_s_academic_v0_1, 0_30_s_youtube_v0_1, 30_60_s_academic_v0_1… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new.tokenspace
tokenspace directory
This directory contains utilities for the purpose of browsing the
"token space" of CLIP ViT-L/14
Primary tools are:
"calculate-distances.py": allows command-line browsing of words and their neighbours
"graph-embeddings.py": plots graph of full values of two embeddings
(clipmodel,cliptextmodel)-calculate-distances.py
Loads the generated embeddings, reads in a word, calculates "distance" to every
embedding, and then shows the closest "neighbours".
To… See the full description on the dataset page: https://huggingface.co/datasets/ppbrown/tokenspace.cc_news_mds_incremental-tokens
