datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_search_net
Dataset Card for CodeSearchNet corpus
Dataset Summary
CodeSearchNet corpus is a dataset of 2 milllion (comment, code) pairs from opensource libraries hosted on GitHub. It contains code and documentation for several programming languages.
CodeSearchNet corpus was gathered to support the CodeSearchNet challenge, to explore the problem of code retrieval using natural language.
Supported Tasks and Leaderboards
language-modeling: The dataset can be used to… See the full description on the dataset page: https://huggingface.co/datasets/code-search-net/code_search_net.code-search-net-python
Dataset Card for "code-search-net-python"
Dataset Description
Homepage: None
Repository: https://huggingface.co/datasets/Nan-Do/code-search-net-python
Paper: None
Leaderboard: None
Point of Contact: @Nan-Do
Dataset Summary
This dataset is the Python portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-python.polymarket-search-indexhle-searchFINDER_API_KEY_AI_SEARCH_2023
FINDER_API_KEY_AI_SEARCH_2023
tags: data collection, machine learning, API performance
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.Patent-searchjtk0-search-ledgersearch-evalunsplash_lite_image_dataset
The Unsplash Dataset
The Unsplash Dataset is made up of over 250,000+ contributing global photographers and data sourced from hundreds of millions of searches across a nearly unlimited number of uses and contexts. Due to the breadth of intent and semantics contained within the Unsplash dataset, it enables new opportunities for research and learning.
The Unsplash Dataset is offered in two datasets:
the Lite dataset: available for commercial and noncommercial usage, containing 25k… See the full description on the dataset page: https://huggingface.co/datasets/image-search-2/unsplash_lite_image_dataset.searchqa
Dataset Card for "searchqa"
Split taken from the MRQA 2019 Shared Task, formatted and filtered for Question Answering. For the original dataset, have a look here.
SWE-bench_Verified-code-searchreason-over-search-eval-m5
Reason-over-Search M5: unified held-out evaluation (9 runs)
Held-out 7-benchmark QA evaluation of the M5 reward-shape x seed ablation: Qwen3.5-0.8B GRPO-trained on MuSiQue with three reward shapes (F1-only / F1+format / EM-only) at three seeds (42 / 43 / 44). This repo is the uniform, light index of every checkpoint's eval scores; it carries the per-cell metric_score.txt + config.yaml and the per-run aggregated CSVs, in ONE consistent layout. (The heavy intermediate_data.json… See the full description on the dataset page: https://huggingface.co/datasets/pantomiman/reason-over-search-eval-m5.MixBench25MixBench is a benchmark for evaluating mixed-modality retrieval. It contains queries and corpora from four datasets: MSCOCO, Google_WIT, VisualNews, and OVEN. Each subset provides: query, corpus, mixed_corpus, and qrel splits.fashion-searchinstructional_code-search-net-python
Dataset Card for "instructional_code-search-net-python"
Dataset Summary
This is an instructional dataset for Python.
The dataset contains two different kind of tasks:
Given a piece of code generate a description of what it does.
Given a description generate a piece of code that fulfils the description.
Languages
The dataset is in English.
Data Splits
There are no splits.
Dataset Creation
May of 2023
Curation Rationale
This… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-python.search-v3-embeddings
Hub Card Search Embeddings (v3)
One-sentence summaries and 1024-d embeddings for 1,173,030 dataset and model cards on the
Hugging Face Hub — 536,870 datasets and 636,160 models. It is the search index behind the revived
librarian-bots/huggingface-semantic-search
backend: you search over a short model-written summary of each card rather than the raw card, and
retrieve against the embedding of that summary.
The cards come from librarian-bots/dataset_cards_with_metadata
and… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/search-v3-embeddings.minif2f-lean4Fixing the errors in some formal statements and informal proofs of minif2f-lean4.
rollouts-olmo7b-cue-search
rollouts-olmo7b-cue-search
Model: allenai/Olmo-3-1025-7B (snapshot a81bae42).
Tokenizer: allenai/Olmo-3-1025-7B (snapshot a81bae42).
Protocol: RL-Zero prompt, MATH-500 x 4 rollouts, budget 31,744, T 0.6, top-p 0.95, seed 20260819 (depth-2 exhaustive and n-gram chain: seed 20260821); the top-20 beam nominee screen, ten random-opener arms, every depth-2 opener (84 shards, arm names unique across shards) and the n-gram chain arms.
Rollouts generated on the CSAIL cluster for the… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo7b-cue-search.weibo-hot-searchSWE-rebench-code-searchSearch-Gen-V-evallog
Search-Gen-V-evallog Dataset
This repository contains the running logs of the experiments conducted in the paper AN EFFICIENT RUBRIC-BASED GENERATIVE VERIFIER FOR SEARCH-AUGMENTED LLMS.
Related links
paper:
AN EFFICIENT RUBRIC-BASED GENERATIVE VERIFIER FOR SEARCH-AUGMENTED LLMS
code:
Search-Gen-V
model:
Search-Gen-V-1.7B-SFT
Search-Gen-V-4B
datasets:
Search-Gen-V
Search-Gen-V-raw
Search-Gen-V-evalSearch-Gen-V-evallog
Result
Table 1. Results on the… See the full description on the dataset page: https://huggingface.co/datasets/lnm1p/Search-Gen-V-evallog.Visual-Searchsearch_privacy_risk
Searching for Privacy Risks in LLM Agents via Simulation
Paper, Code
Authors: Yanzhe Zhang, Diyi Yang
Abstract
The widespread deployment of LLM-based agents is likely to introduce a critical privacy threat: malicious agents that proactively engage others in multi-turn interactions to extract sensitive information. These dynamic dialogues enable adaptive attack strategies that can cause severe privacy violations, yet their evolving nature makes it difficult to anticipate… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/search_privacy_risk.search_qaWe publicly release a new large-scale dataset, called SearchQA, for machine comprehension, or question-answering. Unlike recently released datasets, such as DeepMind
CNN/DailyMail and SQuAD, the proposed SearchQA was constructed to reflect a full pipeline of general question-answering. That is, we start not from an existing article
and generate a question-answer pair, but start from an existing question-answer pair, crawled from J! Archive, and augment it with text snippets retrieved by Google.
Following this approach, we built SearchQA, which consists of more than 140k question-answer pairs with each pair having 49.6 snippets on average. Each question-answer-context
tuple of the SearchQA comes with additional meta-data such as the snippet's URL, which we believe will be valuable resources for future research. We conduct human evaluation
as well as test two baseline methods, one simple word selection and the other deep learning based, on the SearchQA. We show that there is a meaningful gap between the human
and machine performances. This suggests that the proposed dataset could well serve as a benchmark for question-answering.rl_rag_sqa_searcharena_rubrics_web_augmented_outcome_with_new_mcp_system_promptcrag-mm-image-search-tmpSearch-VL-SFT-36K
An Open Recipe for Frontier Multimodal Search Agents
Cold-Start Agentic SFT · Multi-Turn Fatal-Aware GRPO · Visual Tool Use
📑 Table of Contents
📖 Introduction
🗺️ Overview
🍭 Method Overview
📊 Main Results
🔎 Case Study
📁 Repository Layout
🛠️ Prerequisites
🏋️ Agentic SFT · code/SFT
🚀 Agentic RL · code/RL
📊 Inference & Evaluation · code/infer
🚧 TODO
🙌 Acknowledgements
📖 Introduction
OpenSearch-VL is a fully… See the full description on the dataset page: https://huggingface.co/datasets/OpenSearch-VL/Search-VL-SFT-36K.searchlm-eval-caches
searchlm-eval-caches
Evaluation caches used by every doc-PPL / dynamic / QA number in RESULTS.md: v2rL_{val,train}_ob_cache.json (reservoir-sampled holdout/train docs + gold blocks), v2rL_val_shuf_cache.json (shuffled-content control), ood_cache.json + ood_fineweb*.jsonl (FineWeb OOD), rank/query caches, items.jsonl. Load with evals_kf/*.py from Search-LLM analysis/.
Sources on RCP: /mloscratch/homes/ponkshe/search-lm/evals_kf/data
Pushed 2026-09-15 by… See the full description on the dataset page: https://huggingface.co/datasets/Kausp11/searchlm-eval-caches.code_search_net_clean
Dataset Card for "code_search_net_clean"
More Information needed
code_search_net
CodeSearchNet
This is an unofficial reupload of the code_search_net dataset in the parquet format. I have also removed the columns func_code_tokens, func_documentation_tokens, and split_name as they are not relevant. The original repository relies on a Python module that is downloaded and executed to unpack the dataset, which is a potential security risk but importantly raises an annoying warning. As a plus, parquets load faster.
Original model card:
Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/claudios/code_search_net.
