datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pinecone_test
Dataset Card for "pinecone_test"
More Information needed
core-2020-05-10-deduplication
Dataset Card for CORE Deduplication
Dataset Summary
CORE 2020 Deduplication dataset (https://core.ac.uk/documentation/dataset) contains 100K scholarly documents labeled as duplicates/non-duplicates.
Languages
The dataset language is English (BCP-47 en)
Citation Information
@inproceedings{dedup2020,
title={Deduplication of Scholarly Documents using Locality Sensitive Hashing and Word Embeddings},
author={Gyawali, Bikash and Anastasiou, Lucas and… See the full description on the dataset page: https://huggingface.co/datasets/pinecone/core-2020-05-10-deduplication.reddit-qapinecone_hackathon
Dataset Card for "pinecone_hackathon"
More Information needed
sec-10k-qa
SEC 10-K QA Dataset
A retrieval QA dataset built from SEC 10-K annual filings, designed for benchmarking
RAG chunking strategies with MTCB.
Contents
Split
Rows
Description
corpus
95
Cleaned 10-K filing text (20 companies × 5 years)
questions
950
QA pairs generated from corpus chunks
Companies
AAPL, MSFT, GOOGL, AMZN, TSLA, JPM, JNJ, UNH, V, PG,
NVDA, META, BRK, XOM, WMT, BAC, PFE, DIS, NFLX, AMD
Schema
corpus
document_id — filing… See the full description on the dataset page: https://huggingface.co/datasets/Tim-Pinecone/sec-10k-qa.dl-doc-searchlanguage:
en
language_creators:
found
multilinguality:
monolingual
pretty_name: hello
size_categories:
'100K<n<1M
movie-postersyt-transcriptionswikipedia-me5Cohere's Simple Wikipedia Embedded with Multilingual E5 Large
image-setrefinedweb-generated-questions
Generated Questions and Answers from the Falcon RefinedWeb Dataset
This dataset contains 1k open-domain questions and answers generated using documents from Falcon's refinedweb dataset using GPT-4. You can find more details about this work in the following blogpost.
Each row consits of:
document_id - an id of a text chunk from the refined web dataset, from which the question was generated. Each id contains the original document index from the refinedweb dataset, and the chunk index… See the full description on the dataset page: https://huggingface.co/datasets/pinecone/refinedweb-generated-questions.test_pineconeasyouwere-pagesmovie-posters-siglip-embeddingsasyouwere-ocr-p1sst2-MiniLM-embeddingspii-masking-gliner-testpinecone_productslibrispeech-whisper-testsst2-stats
Stats for stanfordnlp/sst2
Generated by dataset-stats.py over the train split.
Total rows in split: 67,349
Rows profiled: 5,000
Columns: 3
Column overview
column
type
kind
null %
highlights
idx
Value(int32)
numeric
0.0%
min=0.00 · p50=2,499.50 · max=4,999.00 · distinct=5000
sentence
Value(string)
string
0.0%
distinct=4,999 · len p50=39 · max=255
label
ClassLabel
class_label
0.0%
positive=2758 · negative=2242
Per-column detail… See the full description on the dataset page: https://huggingface.co/datasets/Tim-Pinecone/sst2-stats.MAGISTRAL_PINECONE_973_CASES_20250728_211652movie-posters-vlm-detect-testasyouwere-extractedbio_pinecone1
