datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pinecone_test
Dataset Card for "pinecone_test"
More Information needed
sec-10k-qa
SEC 10-K QA Dataset
A retrieval QA dataset built from SEC 10-K annual filings, designed for benchmarking
RAG chunking strategies with MTCB.
Contents
Split
Rows
Description
corpus
95
Cleaned 10-K filing text (20 companies × 5 years)
questions
950
QA pairs generated from corpus chunks
Companies
AAPL, MSFT, GOOGL, AMZN, TSLA, JPM, JNJ, UNH, V, PG,
NVDA, META, BRK, XOM, WMT, BAC, PFE, DIS, NFLX, AMD
Schema
corpus
document_id — filing… See the full description on the dataset page: https://huggingface.co/datasets/Tim-Pinecone/sec-10k-qa.pinecone_hackathon
Dataset Card for "pinecone_hackathon"
More Information needed
movie-posterswikipedia-me5Cohere's Simple Wikipedia Embedded with Multilingual E5 Large
asyouwere-pagestest_pineconemovie-posters-siglip-embeddingssst2-MiniLM-embeddingspii-masking-gliner-testasyouwere-ocr-p1librispeech-whisper-testsst2-stats
Stats for stanfordnlp/sst2
Generated by dataset-stats.py over the train split.
Total rows in split: 67,349
Rows profiled: 5,000
Columns: 3
Column overview
column
type
kind
null %
highlights
idx
Value(int32)
numeric
0.0%
min=0.00 · p50=2,499.50 · max=4,999.00 · distinct=5000
sentence
Value(string)
string
0.0%
distinct=4,999 · len p50=39 · max=255
label
ClassLabel
class_label
0.0%
positive=2758 · negative=2242
Per-column detail… See the full description on the dataset page: https://huggingface.co/datasets/Tim-Pinecone/sst2-stats.movie-posters-vlm-detect-testMAGISTRAL_PINECONE_973_CASES_20250728_211652asyouwere-extracted
