datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco-beir-e5msmarco-beir-constbertpinecone_test
Dataset Card for "pinecone_test"
More Information needed
core-2020-05-10-deduplication
Dataset Card for CORE Deduplication
Dataset Summary
CORE 2020 Deduplication dataset (https://core.ac.uk/documentation/dataset) contains 100K scholarly documents labeled as duplicates/non-duplicates.
Languages
The dataset language is English (BCP-47 en)
Citation Information
@inproceedings{dedup2020,
title={Deduplication of Scholarly Documents using Locality Sensitive Hashing and Word Embeddings},
author={Gyawali, Bikash and Anastasiou, Lucas and… See the full description on the dataset page: https://huggingface.co/datasets/pinecone/core-2020-05-10-deduplication.sec-10k-qa-embeddings
SEC 10-K QA Embeddings
Pre-computed embeddings for the Tim-Pinecone/sec-10k-qa dataset.
What's in here
Each config is a parquet file containing pre-computed text-embedding-ada-002 embeddings
for a specific chunking strategy applied to the SEC 10-K corpus.
Config
Description
questions_ada002
All 950 evaluation questions
chunks_RecursiveChunker_512_ada002
RecursiveChunker at chunk_size=512
chunks_RecursiveChunker_1024_ada002
RecursiveChunker at… See the full description on the dataset page: https://huggingface.co/datasets/Tim-Pinecone/sec-10k-qa-embeddings.reddit-qapinecone_hackathon
Dataset Card for "pinecone_hackathon"
More Information needed
dl-doc-searchlanguage:
en
language_creators:
found
multilinguality:
monolingual
pretty_name: hello
size_categories:
'100K<n<1M
sec-10k-qa
SEC 10-K QA Dataset
A retrieval QA dataset built from SEC 10-K annual filings, designed for benchmarking
RAG chunking strategies with MTCB.
Contents
Split
Rows
Description
corpus
95
Cleaned 10-K filing text (20 companies × 5 years)
questions
950
QA pairs generated from corpus chunks
Companies
AAPL, MSFT, GOOGL, AMZN, TSLA, JPM, JNJ, UNH, V, PG,
NVDA, META, BRK, XOM, WMT, BAC, PFE, DIS, NFLX, AMD
Schema
corpus
document_id — filing… See the full description on the dataset page: https://huggingface.co/datasets/Tim-Pinecone/sec-10k-qa.movie-postersyt-transcriptionswikipedia-me5Cohere's Simple Wikipedia Embedded with Multilingual E5 Large
some_notesrefinedweb-generated-questions
Generated Questions and Answers from the Falcon RefinedWeb Dataset
This dataset contains 1k open-domain questions and answers generated using documents from Falcon's refinedweb dataset using GPT-4. You can find more details about this work in the following blogpost.
Each row consits of:
document_id - an id of a text chunk from the refined web dataset, from which the question was generated. Each id contains the original document index from the refinedweb dataset, and the chunk index… See the full description on the dataset page: https://huggingface.co/datasets/pinecone/refinedweb-generated-questions.image-setmovielens-recent-ratingsThis dataset streams recent user ratings from the MovieLens 25M dataset and adds poster URLs.test_pineconeexample-context-wealth-advisorypinecone_arknights
Dataset of pinecone/パインコーン/松果 (Arknights)
This is the dataset of pinecone/パインコーン/松果 (Arknights), containing 192 images and their tags.
The core tags of this character are long_hair, feather_hair, mole, mole_under_eye, ponytail, orange_hair, brown_eyes, bow, orange_eyes, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
List of Packages… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/pinecone_arknights.asyouwere-pagesasyouwere-ocr-p1sst2-MiniLM-embeddingspii-masking-gliner-testpinecone_productsmovie-posters-siglip-embeddingslibrispeech-whisper-testsst2-stats
Stats for stanfordnlp/sst2
Generated by dataset-stats.py over the train split.
Total rows in split: 67,349
Rows profiled: 5,000
Columns: 3
Column overview
column
type
kind
null %
highlights
idx
Value(int32)
numeric
0.0%
min=0.00 · p50=2,499.50 · max=4,999.00 · distinct=5000
sentence
Value(string)
string
0.0%
distinct=4,999 · len p50=39 · max=255
label
ClassLabel
class_label
0.0%
positive=2758 · negative=2242
Per-column detail… See the full description on the dataset page: https://huggingface.co/datasets/Tim-Pinecone/sst2-stats.testMAGISTRAL_PINECONE_973_CASES_20250728_211652movie-posters-vlm-detect-test
