CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01edanigoben /pinecone_test Dataset Card for "pinecone_test" More Information needed text1M<n<10M0 likes304 downloads3y agoHugging Face02pinecone /core-2020-05-10-deduplication Dataset Card for CORE Deduplication Dataset Summary CORE 2020 Deduplication dataset (https://core.ac.uk/documentation/dataset) contains 100K scholarly documents labeled as duplicates/non-duplicates. Languages The dataset language is English (BCP-47 en) Citation Information @inproceedings{dedup2020, title={Deduplication of Scholarly Documents using Locality Sensitive Hashing and Word Embeddings}, author={Gyawali, Bikash and Anastasiou, Lucas and… See the full description on the dataset page: https://huggingface.co/datasets/pinecone/core-2020-05-10-deduplication.textother100K<n<1M1 likes198 downloads4y agoHugging Face03pinecone /reddit-qatext1K<n<10K3 likes85 downloads4y agoHugging Face04factored /pinecone_hackathon Dataset Card for "pinecone_hackathon" More Information needed text100K<n<1M0 likes76 downloads3y agoHugging Face05Tim-Pinecone /sec-10k-qa SEC 10-K QA Dataset A retrieval QA dataset built from SEC 10-K annual filings, designed for benchmarking RAG chunking strategies with MTCB. Contents Split Rows Description corpus 95 Cleaned 10-K filing text (20 companies × 5 years) questions 950 QA pairs generated from corpus chunks Companies AAPL, MSFT, GOOGL, AMZN, TSLA, JPM, JNJ, UNH, V, PG, NVDA, META, BRK, XOM, WMT, BAC, PFE, DIS, NFLX, AMD Schema corpus document_id — filing… See the full description on the dataset page: https://huggingface.co/datasets/Tim-Pinecone/sec-10k-qa.textquestion-answering1K<n<10K0 likes71 downloads6mo agoHugging Face06pinecone /dl-doc-searchlanguage: en language_creators: found multilinguality: monolingual pretty_name: hello size_categories: '100K<n<1M text100K<n<1M0 likes69 downloads4y agoHugging Face07pinecone /movie-postersimage10K<n<100K2 likes60 downloads4y agoHugging Face08pinecone /yt-transcriptionsimage10K<n<100K1 likes43 downloads4y agoHugging Face09pinecone /wikipedia-me5Cohere's Simple Wikipedia Embedded with Multilingual E5 Large tabular100K<n<1M1 likes41 downloads2y agoHugging Face10pinecone /image-settextn<1K1 likes35 downloads4y agoHugging Face11pinecone /refinedweb-generated-questions Generated Questions and Answers from the Falcon RefinedWeb Dataset This dataset contains 1k open-domain questions and answers generated using documents from Falcon's refinedweb dataset using GPT-4. You can find more details about this work in the following blogpost. Each row consits of: document_id - an id of a text chunk from the refined web dataset, from which the question was generated. Each id contains the original document index from the refinedweb dataset, and the chunk index… See the full description on the dataset page: https://huggingface.co/datasets/pinecone/refinedweb-generated-questions.textquestion-answering1K<n<10K3 likes35 downloads3y agoHugging Face12seregadgl /test_pineconetext10K<n<100K0 likes24 downloads2y agoHugging Face13Tim-Pinecone /asyouwere-pagesimagen<1K0 likes21 downloads4mo agoHugging Face14Tim-Pinecone /movie-posters-siglip-embeddingsimage1K<n<10K0 likes18 downloads4mo agoHugging Face15Tim-Pinecone /asyouwere-ocr-p1imagen<1K0 likes17 downloads4mo agoHugging Face16Tim-Pinecone /sst2-MiniLM-embeddingstext1K<n<10K0 likes16 downloads4mo agoHugging Face17Tim-Pinecone /pii-masking-gliner-testtextn<1K0 likes16 downloads4mo agoHugging Face18Shubham09 /pinecone_productstext1K<n<10K0 likes15 downloads3y agoHugging Face19Tim-Pinecone /librispeech-whisper-testaudion<1K0 likes15 downloads4mo agoHugging Face20Tim-Pinecone /sst2-stats Stats for stanfordnlp/sst2 Generated by dataset-stats.py over the train split. Total rows in split: 67,349 Rows profiled: 5,000 Columns: 3 Column overview column type kind null % highlights idx Value(int32) numeric 0.0% min=0.00 · p50=2,499.50 · max=4,999.00 · distinct=5000 sentence Value(string) string 0.0% distinct=4,999 · len p50=39 · max=255 label ClassLabel class_label 0.0% positive=2758 · negative=2242 Per-column detail… See the full description on the dataset page: https://huggingface.co/datasets/Tim-Pinecone/sst2-stats.tabularn<1K0 likes14 downloads4mo agoHugging Face21master-laoo-dev /MAGISTRAL_PINECONE_973_CASES_20250728_211652textn<1K0 likes13 downloads1y agoHugging Face22Tim-Pinecone /movie-posters-vlm-detect-testimagen<1K0 likes13 downloads4mo agoHugging Face23Tim-Pinecone /asyouwere-extractedimagen<1K0 likes12 downloads4mo agoHugging Face24alexignite /bio_pinecone1text10K<n<100K0 likes9 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.