datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FINDER_API_KEY_AI_SEARCH_2023
FINDER_API_KEY_AI_SEARCH_2023
tags: data collection, machine learning, API performance
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.weibo-hot-searchcode_search_net_python_10000_examplesr15-ai-search-metamerism
R15: AI Search Metamerism — Cross-Cultural Brand Perception Dataset
Citation: Zharnikov, D. (2026v) | DOI: 10.5281/zenodo.19422427 | Version: v3.2.0
Dataset Summary
This dataset contains the full session logs, aggregated results, and analysis outputs from the R15 large-scale experiment testing whether Large Language Models systematically collapse multi-dimensional brand perception into Economic and Experiential dimensions ("spectral metamerism"). It comprises 21… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/r15-ai-search-metamerism.deepresearchgym-agentic-search-logs
DeepResearchGym Agentic Search Logs
This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617).
The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253.
All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/deepresearchgym-agentic-search-logs.hnm-search-data
HnM Search Dataset Created from Recommendations Dataset
This synthetic data-set is created using the recommendations dataset:
https://huggingface.co/datasets/einrafh/hnm-fashion-recommendations-data (Use of this dataset is subject to the terms and conditions set forth on the original distribution page. This dataset is intended for non-commercial and research use.)
https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/data (DATA ACCESS AND USE:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/hnm-search-data.search-source-audit
Sources of Truth — AI Search Citations for Mental Health Queries
Which external sources do consumer AI search products actually cite when people ask about mental
health? This dataset is the annotated citation corpus behind "Sources of Truth: A Multi-Platform,
Multilingual Audit of Citations in AI Mental Health Information Queries."
Twenty English mental health questions were put to three free consumer AI search products (ChatGPT, Perplexity, and Google AI Overview) under two… See the full description on the dataset page: https://huggingface.co/datasets/MindBench/search-source-audit.Object_Search
JojoQaQ/Object_Search
Instruction-conditioned retrieval triplets over Amazon Reviews 2023 Appliances.
Every column is a literal model input. These rows are produced by
src.finetune_triplets.model_input_rows, the same function that writes the
fine-tuning run's W&B tables, over triplets loaded by
load_training_triplets, the same loader the training entrypoint calls. What
you see here is what the embedder receives, field for field, not a rendering
of it.
Splits… See the full description on the dataset page: https://huggingface.co/datasets/JojoQaQ/Object_Search.Product-Search-Images-v0.1MTSBerquadMTSBerquad is a cleaned and enriched dataset SberQuAD transferred to the Generative QA task. All entities were truecased, refactored by hand to improve readability and consistency. Answers have been expanded and rearranged from MLM QA task to Generative/Long Form QA task. MTSBerquad presented in PyCon 2024 by MTS AI Search Group.
Developed by MTS AI Search Group (Krayko Nikita, Laputin Fedor, Sidorov Ivan)
autocomplete-search-datasetlinux-file-search
Linux File Search Dataset
Dataset Summary
The Linux File Search NLI Dataset is a synthetic dataset designed to train and evaluate Natural Language Inference (NLI) models that map natural language file search queries into structured representations of file attributes.
The dataset is intended to enable semantic file search on Linux systems by allowing models to extract structured constraints such as file type, extension, size, ownership, permissions, and other properties… See the full description on the dataset page: https://huggingface.co/datasets/software-si/linux-file-search.lab05-semantic-searchproduct-search-2023-queriesdeepresearchgym-agentic-search-logs
DeepResearchGym Agentic Search Logs
This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617).
The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253.
All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/ethanning/deepresearchgym-agentic-search-logs.suicide_searchvolumenegamax-search-depth-scaling
Negamax Search-Depth Scaling — What Is One Ply of Search Worth?
Measured strength gain, in Elo per extra ply of search depth, for four
depth-limited negamax + alpha-beta board-game engines: Connect 4, Checkers,
Othello, and Chess. Strength is defined as a fixed search(state, {maxDepth: N})
depth — a deterministic, hardware-independent knob — so every number reproduces from a seed.
Full write-up: https://lkforge.com/blog/search-depth-scaling/
Playable engines:… See the full description on the dataset page: https://huggingface.co/datasets/LKForge/negamax-search-depth-scaling.semantic_searchameba_faq_search
AMEBA Blog FAQ Search Dataset
This data was obtained by crawling this website.
The FAQ Data was processed to remove HTML tags and other formatting after crawling, and entries containing excessively long content were excluded.
The Query Data was generated using a Large Language Model (LLM). Please refer to the following blog for information about the generation process.
https://www.ai-shift.co.jp/techblog/3710
https://www.ai-shift.co.jp/techblog/3761
Column description… See the full description on the dataset page: https://huggingface.co/datasets/ai-shift/ameba_faq_search.semantic-job-search-dataset
Semantic Job Search Dataset (Synthetic)
Overview
This dataset contains 10,000 synthetic job postings designed for semantic search.Each record represents a job listing with structured attributes (e.g., field, location, job type) plus a short natural-language description.
The dataset was generated as part of a course final project and is used to support a Gradio app that performs semantic job search using Sentence-Transformer embeddings and cosine similarity.… See the full description on the dataset page: https://huggingface.co/datasets/aurele1/semantic-job-search-dataset.nearest_neighbor_searchgoogle_search_terms_training_data
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Dataset Name: Google Search Trends Top Rising Search Terms
Description:
The Google Search Trends Top Rising Search Terms dataset provides valuable insights into the most rapidly growing search queries on the Google search engine. It offers a comprehensive collection of trending search… See the full description on the dataset page: https://huggingface.co/datasets/hoshangc/google_search_terms_training_data.product-search-2024-queriesproduct-search-2025-test-querieslab05-semantic-searchMP16-Searchai-search-visibility-romania-electronics-market
AI Search Visibility — Romania's Electronics & IT Market (August 2026)
18 brand-free purchase questions × 5 AI engines = 87 answers. 86 of them name a major retailer. Position, not presence, decides the market. Raw data CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
Eighteen real purchase questions were put to ChatGPT, Google Gemini, Perplexity, Google AI Mode and Google AI Overviews, in Romanian, from Romania, in… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-search-visibility-romania-electronics-market.khmer-search-frequency
Khmer Search Frequency
How often each Khmer dictionary headword was the word a user was actually
looking for, measured by real lookup demand rather than corpus occurrence.
Each row is a headword and the number of search sessions that ended on it. A
session counts only when the user settled on that word: sessions abandoned by
backspacing to a fragment are excluded, since single Khmer letters are
themselves headwords and would otherwise dominate the ranking.
Columns… See the full description on the dataset page: https://huggingface.co/datasets/seanghay/khmer-search-frequency.consumer-vs-critic-wine-searcher-2026
Consumer vs Critic: Wine-Searcher 2026
Open dataset by MrBridge — 598 rows of wine data (ratings, prices and attributes), free to download, inspect and reuse.
Data source
Collected with the Wine-Searcher Scraper on Apify — a cloud actor returning clean, structured CSV/JSON with no local setup. Re-run it yourself for fresh data.
Get fresh data → https://apify.com/mrbridge/wine-searcher-scraper-from-list?fpr=mrbridge
Schema
vivino_sample_300.csv — 300… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Bridge/consumer-vs-critic-wine-searcher-2026.us-excel-template-search-demand-2026
US Search Demand for Excel & Google Sheets Templates — 1,008 Keywords (2026)
1,008 template-related keywords with estimated monthly US search volume, grouped into
18 practical categories (bookkeeping, HR & payroll, project management, invoicing,
construction, weddings...).
Columns: keyword, monthly_search_volume_us, category, has_free_template,
free_template_url. Where a free, no-signup spreadsheet implementation exists, the row
links to it on tabletemplates.com — see the
free… See the full description on the dataset page: https://huggingface.co/datasets/tresor2k/us-excel-template-search-demand-2026.
