CoolFace
Datasetpublic

ericssonbear/mr-right-zhtw-embeddings

Mr. Right zh-TW — pre-computed document embeddings The deployed document bank (zhen_img_v2) for the Mr. Right zh-TW corpus: one 4096-dimensional vector per document, covering all 769,245 documents. Encoding all of them takes multiple GPU-days and requires the full 1.2 TB image tree. This repository exists so you can skip both and go straight to retrieval. file shape / size notes emb.npy (769245, 4096) float16, 6.3 GB L2-normalised — cosine similarity is a plain dot… See the full description on the dataset page: https://huggingface.co/datasets/ericssonbear/mr-right-zhtw-embeddings.

sourceHugging Facecc-by-sa-4.0updated 2mo agoView on Hugging Face
0likes50downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
ericssonbear/mr-right-zhtw-embeddings · CoolFace