CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01code-search-net /code_search_net Dataset Card for CodeSearchNet corpus Dataset Summary CodeSearchNet corpus is a dataset of 2 milllion (comment, code) pairs from opensource libraries hosted on GitHub. It contains code and documentation for several programming languages. CodeSearchNet corpus was gathered to support the CodeSearchNet challenge, to explore the problem of code retrieval using natural language. Supported Tasks and Leaderboards language-modeling: The dataset can be used to… See the full description on the dataset page: https://huggingface.co/datasets/code-search-net/code_search_net.texttext-generation1M<n<10M338 likes32k downloads7mo agoHugging Face02Nan-Do /code-search-net-python Dataset Card for "code-search-net-python" Dataset Description Homepage: None Repository: https://huggingface.co/datasets/Nan-Do/code-search-net-python Paper: None Leaderboard: None Point of Contact: @Nan-Do Dataset Summary This dataset is the Python portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-python.texttext-generation100K<n<1M30 likes4.1k downloads3y agoHugging Face03Becky7777777 /polymarket-search-indextabular1M<n<10M0 likes3.6k downloads12h agoHugging Face04alvinming /hle-searchtextn<1K0 likes3.3k downloads8mo agoHugging Face05infinite-dataset-hub /FINDER_API_KEY_AI_SEARCH_2023 FINDER_API_KEY_AI_SEARCH_2023 tags: data collection, machine learning, API performance Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.tabularn<1K0 likes1.9k downloads2y agoHugging Face06Mouuns /Patent-search0 likes1.8k downloads10mo agoHugging Face07jtktolivefor /jtk0-search-ledger0 likes1.7k downloads6d agoHugging Face08axon-rl /search-evaltext1K<n<10K0 likes1.7k downloads1y agoHugging Face09image-search-2 /unsplash_lite_image_dataset The Unsplash Dataset The Unsplash Dataset is made up of over 250,000+ contributing global photographers and data sourced from hundreds of millions of searches across a nearly unlimited number of uses and contexts. Due to the breadth of intent and semantics contained within the Unsplash dataset, it enables new opportunities for research and learning. The Unsplash Dataset is offered in two datasets: the Lite dataset: available for commercial and noncommercial usage, containing 25k… See the full description on the dataset page: https://huggingface.co/datasets/image-search-2/unsplash_lite_image_dataset.3 likes1.5k downloads5y agoHugging Face10lucadiliello /searchqa Dataset Card for "searchqa" Split taken from the MRQA 2019 Shared Task, formatted and filtered for Question Answering. For the original dataset, have a look here. text100K<n<1M5 likes1.3k downloads3y agoHugging Face11adityasoni17 /SWE-bench_Verified-code-searchtextn<1K0 likes1.3k downloads9mo agoHugging Face12pantomiman /reason-over-search-eval-m5 Reason-over-Search M5: unified held-out evaluation (9 runs) Held-out 7-benchmark QA evaluation of the M5 reward-shape x seed ablation: Qwen3.5-0.8B GRPO-trained on MuSiQue with three reward shapes (F1-only / F1+format / EM-only) at three seeds (42 / 43 / 44). This repo is the uniform, light index of every checkpoint's eval scores; it carries the per-cell metric_score.txt + config.yaml and the per-run aggregated CSVs, in ONE consistent layout. (The heavy intermediate_data.json… See the full description on the dataset page: https://huggingface.co/datasets/pantomiman/reason-over-search-eval-m5.text1K<n<10K0 likes1.2k downloads4mo agoHugging Face13mixed-modality-search /MixBench25MixBench is a benchmark for evaluating mixed-modality retrieval. It contains queries and corpora from four datasets: MSCOCO, Google_WIT, VisualNews, and OVEN. Each subset provides: query, corpus, mixed_corpus, and qrel splits.imagetext-ranking0 likes1k downloads1y agoHugging Face14sohaillansarii /fashion-search0 likes982 downloads2mo agoHugging Face15Nan-Do /instructional_code-search-net-python Dataset Card for "instructional_code-search-net-python" Dataset Summary This is an instructional dataset for Python. The dataset contains two different kind of tasks: Given a piece of code generate a description of what it does. Given a description generate a piece of code that fulfils the description. Languages The dataset is in English. Data Splits There are no splits. Dataset Creation May of 2023 Curation Rationale This… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-python.texttext-generation100K<n<1M36 likes972 downloads3y agoHugging Face16davanstrien /search-v3-embeddings Hub Card Search Embeddings (v3) One-sentence summaries and 1024-d embeddings for 1,173,030 dataset and model cards on the Hugging Face Hub — 536,870 datasets and 636,160 models. It is the search index behind the revived librarian-bots/huggingface-semantic-search backend: you search over a short model-written summary of each card rather than the raw card, and retrieve against the embedding of that summary. The cards come from librarian-bots/dataset_cards_with_metadata and… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/search-v3-embeddings.tabularfeature-extraction1M<n<10M4 likes968 downloads1mo agoHugging Face17cat-searcher /minif2f-lean4Fixing the errors in some formal statements and informal proofs of minif2f-lean4. textn<1K7 likes948 downloads3y agoHugging Face18reasoning-cues /rollouts-olmo7b-cue-search rollouts-olmo7b-cue-search Model: allenai/Olmo-3-1025-7B (snapshot a81bae42). Tokenizer: allenai/Olmo-3-1025-7B (snapshot a81bae42). Protocol: RL-Zero prompt, MATH-500 x 4 rollouts, budget 31,744, T 0.6, top-p 0.95, seed 20260819 (depth-2 exhaustive and n-gram chain: seed 20260821); the top-20 beam nominee screen, ten random-opener arms, every depth-2 opener (84 shards, arm names unique across shards) and the n-gram chain arms. Rollouts generated on the CSAIL cluster for the… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo7b-cue-search.tabular100K<n<1M0 likes881 downloads8d agoHugging Face19LZYFirecn /weibo-hot-searchtabular10K<n<100K3 likes879 downloads1y agoHugging Face20OpenHands /SWE-rebench-code-searchtext10K<n<100K0 likes879 downloads6mo agoHugging Face21lnm1p /Search-Gen-V-evallog Search-Gen-V-evallog Dataset This repository contains the running logs of the experiments conducted in the paper AN EFFICIENT RUBRIC-BASED GENERATIVE VERIFIER FOR SEARCH-AUGMENTED LLMS. Related links paper: AN EFFICIENT RUBRIC-BASED GENERATIVE VERIFIER FOR SEARCH-AUGMENTED LLMS code: Search-Gen-V model: Search-Gen-V-1.7B-SFT Search-Gen-V-4B datasets: Search-Gen-V Search-Gen-V-raw Search-Gen-V-evalSearch-Gen-V-evallog Result Table 1. Results on the… See the full description on the dataset page: https://huggingface.co/datasets/lnm1p/Search-Gen-V-evallog.0 likes809 downloads11mo agoHugging Face22multimodal-reasoning-lab /Visual-Searchimage10K<n<100K4 likes756 downloads1y agoHugging Face23SALT-NLP /search_privacy_risk Searching for Privacy Risks in LLM Agents via Simulation Paper, Code Authors: Yanzhe Zhang, Diyi Yang Abstract The widespread deployment of LLM-based agents is likely to introduce a critical privacy threat: malicious agents that proactively engage others in multi-turn interactions to extract sensitive information. These dynamic dialogues enable adaptive attack strategies that can cause severe privacy violations, yet their evolving nature makes it difficult to anticipate… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/search_privacy_risk.1 likes748 downloads1y agoHugging Face24kyunghyuncho /search_qaWe publicly release a new large-scale dataset, called SearchQA, for machine comprehension, or question-answering. Unlike recently released datasets, such as DeepMind CNN/DailyMail and SQuAD, the proposed SearchQA was constructed to reflect a full pipeline of general question-answering. That is, we start not from an existing article and generate a question-answer pair, but start from an existing question-answer pair, crawled from J! Archive, and augment it with text snippets retrieved by Google. Following this approach, we built SearchQA, which consists of more than 140k question-answer pairs with each pair having 49.6 snippets on average. Each question-answer-context tuple of the SearchQA comes with additional meta-data such as the snippet's URL, which we believe will be valuable resources for future research. We conduct human evaluation as well as test two baseline methods, one simple word selection and the other deep learning based, on the SearchQA. We show that there is a meaningful gap between the human and machine performances. This suggests that the proposed dataset could well serve as a benchmark for question-answering.question-answering100K<n<1M23 likes687 downloads3y agoHugging Face25rl-rag /rl_rag_sqa_searcharena_rubrics_web_augmented_outcome_with_new_mcp_system_prompttext1K<n<10K0 likes615 downloads1y agoHugging Face26Melmaphother /crag-mm-image-search-tmpimage1K<n<10K0 likes611 downloads1y agoHugging Face27OpenSearch-VL /Search-VL-SFT-36K An Open Recipe for Frontier Multimodal Search Agents Cold-Start Agentic SFT  ·  Multi-Turn Fatal-Aware GRPO  ·  Visual Tool Use 📑 Table of Contents 📖 Introduction 🗺️ Overview 🍭 Method Overview 📊 Main Results 🔎 Case Study 📁 Repository Layout 🛠️ Prerequisites 🏋️ Agentic SFT · code/SFT 🚀 Agentic RL · code/RL 📊 Inference & Evaluation · code/infer 🚧 TODO 🙌 Acknowledgements 📖 Introduction OpenSearch-VL is a fully… See the full description on the dataset page: https://huggingface.co/datasets/OpenSearch-VL/Search-VL-SFT-36K.image10K<n<100K10 likes582 downloads5mo agoHugging Face28Kausp11 /searchlm-eval-caches searchlm-eval-caches Evaluation caches used by every doc-PPL / dynamic / QA number in RESULTS.md: v2rL_{val,train}_ob_cache.json (reservoir-sampled holdout/train docs + gold blocks), v2rL_val_shuf_cache.json (shuffled-content control), ood_cache.json + ood_fineweb*.jsonl (FineWeb OOD), rank/query caches, items.jsonl. Load with evals_kf/*.py from Search-LLM analysis/. Sources on RCP: /mloscratch/homes/ponkshe/search-lm/evals_kf/data Pushed 2026-09-15 by… See the full description on the dataset page: https://huggingface.co/datasets/Kausp11/searchlm-eval-caches.0 likes565 downloads6d agoHugging Face29jinaai /code_search_net_clean Dataset Card for "code_search_net_clean" More Information needed text1M<n<10M1 likes547 downloads3y agoHugging Face30claudios /code_search_net CodeSearchNet This is an unofficial reupload of the code_search_net dataset in the parquet format. I have also removed the columns func_code_tokens, func_documentation_tokens, and split_name as they are not relevant. The original repository relies on a Python module that is downloaded and executed to unpack the dataset, which is a potential security risk but importantly raises an annoying warning. As a plus, parquets load faster. Original model card: Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/claudios/code_search_net.texttext-generation1M<n<10M11 likes535 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.