SciPhi/AgentSearch-V1
Getting Started The AgentSearch-V1 dataset boasts a comprehensive collection of over one billion embeddings, produced using jina-v2-base. The dataset encompasses more than 50 million high-quality documents and over 1 billion passages, covering a vast range of content from sources such as Arxiv, Wikipedia, Project Gutenberg, and includes carefully filtered Creative Commons (CC) data. Our team is dedicated to continuously expanding and enhancing this corpus to improve the search… See the full description on the dataset page: https://huggingface.co/datasets/SciPhi/AgentSearch-V1.
Update README.md
Update README.md
Update README.md
add more arxiv
add more arxiv
add more arxiv
batch 48
batch 47
batch 46
batch 45
batch 6
batch 44
batch 43
batch 42
batch 41
batch 40
batch 39
batch 38
batch 37
batch 36
batch 35
batch 34
batch 33
batch 32
batch 31
batch 30
batch 29
batch 28
batch 27
batch 26
batch 25
batch 24
batch 23
batch 22
batch 21
batch 20
batch 19
batch 40
batch 39
batch 38
batch 37
batch 36
batch 2
batch 35
batch 34
batch 33
batch 32
batch 31
batch 30
batch 29
