sleeping-ai/arXiv-abstract-model2vec
arXiv-model2vec Datase The arXiv-model2vec dataset contains embeddings for all arXiv paper abstracts and their corresponding titles, generated using the Model2Vec proton-8M model. This dataset is released to support research projects that require high-quality semantic representations of scientific papers. Dataset Details Source: The dataset includes embeddings derived from arXiv paper titles and abstracts. Embedding Model: The embeddings are generated using the… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/arXiv-abstract-model2vec.
Upload folder using huggingface_hub
Delete paper-parquet.zip
Upload paper-parquet.zip with huggingface_hub
Delete dataset.py
Delete dataset_infos.json
Update dataset.py
Update dataset.py
Update README.md
Create dataset.py
Create dataset_infos.json
Update README.md
Upload paper.zip with huggingface_hub
initial commit
