datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FINDER_API_KEY_AI_SEARCH_2023
FINDER_API_KEY_AI_SEARCH_2023
tags: data collection, machine learning, API performance
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.search-arena-24k
Overview
This dataset contains ALL in-the-wild conversation crowdsourced from Search Arena between March 18, 2025 and May 8, 2025. It includes 24,069 multi-turn conversations with search-LLMs across diverse intents, languages, and topics—alongside 12,652 human preference votes. The dataset spans approximately 11,000 users across 136 countries, 13 publicly released models, around 90 languages (including 11% multilingual prompts), and over 5,000 multi-turn sessions.
While user… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/search-arena-24k.r15-ai-search-metamerism
R15: AI Search Metamerism — Cross-Cultural Brand Perception Dataset
Citation: Zharnikov, D. (2026v) | DOI: 10.5281/zenodo.19422427 | Version: v3.2.0
Dataset Summary
This dataset contains the full session logs, aggregated results, and analysis outputs from the R15 large-scale experiment testing whether Large Language Models systematically collapse multi-dimensional brand perception into Economic and Experiential dimensions ("spectral metamerism"). It comprises 21… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/r15-ai-search-metamerism.theorem-search-dataset
Theorem Search Dataset
The largest open corpus of informal mathematical theorems: 1,341,083 theorem statements with natural-language slogans from 209,777 papers, designed for semantic theorem retrieval.
Paper: Semantic Search over 9 Million Mathematical Theorems
Demo: huggingface.co/spaces/uw-math-ai/theorem-search
Benchmark results
On 110 test queries written by research mathematicians, our best pipeline (Qwen3-Embedding-8B on DeepSeek-V3.1 slogans) outperforms all… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/theorem-search-dataset.search-arena-v1-7k
Overview
This dataset contains 7k leaderboard conversation votes collected from Search Arena between March 18, 2025 and April 13, 2025. All entries have been redacted for PII and sensitive user information to ensure privacy.
Each data point includes:
Two model responses (messages_a and messages_b)
The human vote result
A timestamp
Full system metadata, LLM + web search trace, and post-processed metadata for controlled experiments (conv_meta)
To reproduce the leaderboard results… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/search-arena-v1-7k.theorem-search-dataset-permissive
Theorem Search Dataset
The largest open corpus of informal mathematical theorems: 1,239,720 theorem statements with natural-language slogans from 197,889 papers, designed for semantic theorem retrieval.
Paper: Semantic Search over 9 Million Mathematical Theorems
Demo: huggingface.co/spaces/uw-math-ai/theorem-search
Benchmark results
On 110 test queries written by research mathematicians, our best pipeline (Qwen3-Embedding-8B on DeepSeek-V3.1 slogans) outperforms all… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/theorem-search-dataset-permissive.ecommerce-search-extraction
Ionio E-commerce Search Query Extraction
Built with: simula — schema-driven synthetic data generation with auditable taxonomy lineage.
An English synthetic dataset for training and evaluating systems that translate natural-language
shopping requests into narrow, atomic, database-queryable JSON. It contains 10,985 accepted
examples from a 13,000-attempt generation run. No accepted rows were trimmed from this release.
Each example pairs a realistic typed or spoken shopper query… See the full description on the dataset page: https://huggingface.co/datasets/Ionio-ai/ecommerce-search-extraction.MTSBerquadMTSBerquad is a cleaned and enriched dataset SberQuAD transferred to the Generative QA task. All entities were truecased, refactored by hand to improve readability and consistency. Answers have been expanded and rearranged from MLM QA task to Generative/Long Form QA task. MTSBerquad presented in PyCon 2024 by MTS AI Search Group.
Developed by MTS AI Search Group (Krayko Nikita, Laputin Fedor, Sidorov Ivan)
job-searcher-data
AI Job Searcher Training Data (V3)
Fine-tuning dataset for a career advisor AI specializing in Nordic and European job markets.
V3 (March 2026): Added 150 freeform/narrative-style Analyze examples modeled after real
job postings from finn.no and arbeidsplassen.nav.no. Now includes startup-style, agency,
generalist, and narrative formats alongside V2 rigid-template examples.
Dataset Description
This dataset contains 1,190 training examples across 9 languages and 5… See the full description on the dataset page: https://huggingface.co/datasets/ai-colombia/job-searcher-data.natural-qa-random-67-with-AI-search-answers
Dataset Details
Dataset Description
This dataset is a refined subset of the "Natural Questions" dataset, filtered to include only high-quality answers as labeled manually. The dataset includes ground truth examples of "good" answers, defined as responses that are correct, clear, and sufficient for the given questions. Additionally, answers generated by three AI search engines (Perplexity, Gemini, Exa AI) have been incorporated to provide both raw and parsed outputs for… See the full description on the dataset page: https://huggingface.co/datasets/quotientai/natural-qa-random-67-with-AI-search-answers.ai-job-searcher-training-data
AI Job Searcher Training Data
Fine-tuning dataset for a career advisor AI specializing in Nordic and European job markets.
Dataset Description
This dataset contains 1,040 training examples across 9 languages and 5 task categories, formatted as chat conversations (system/user/assistant) suitable for fine-tuning LLMs.
Task Categories
Category
Examples
Description
Cover Letter Generation
208
Professional cover letters from job description + user profile… See the full description on the dataset page: https://huggingface.co/datasets/ai-colombia/ai-job-searcher-training-data.ai-residency-vector-search-retrieval-dataameba_faq_search
AMEBA Blog FAQ Search Dataset
This data was obtained by crawling this website.
The FAQ Data was processed to remove HTML tags and other formatting after crawling, and entries containing excessively long content were excluded.
The Query Data was generated using a Large Language Model (LLM). Please refer to the following blog for information about the generation process.
https://www.ai-shift.co.jp/techblog/3710
https://www.ai-shift.co.jp/techblog/3761
Column description… See the full description on the dataset page: https://huggingface.co/datasets/ai-shift/ameba_faq_search.bfcl_v4_web_searchai-search-visibility-romania-electronics-market
AI Search Visibility — Romania's Electronics & IT Market (August 2026)
18 brand-free purchase questions × 5 AI engines = 87 answers. 86 of them name a major retailer. Position, not presence, decides the market. Raw data CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
Eighteen real purchase questions were put to ChatGPT, Google Gemini, Perplexity, Google AI Mode and Google AI Overviews, in Romanian, from Romania, in… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-search-visibility-romania-electronics-market.ai-search-agent
AI Search Agent Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ai-search-agent.search-recommendation-ai-agent
Search Recommendation Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/search-recommendation-ai-agent.fuzzy_search_datasetai-search-visibility-romania-book-market
AI Search Visibility — Romania's Book Market (July 2026)
When someone asks ChatGPT "which online bookstore should I use for children's
books?", they get one answer, not ten blue links. This dataset measures
who is inside that answer — and who merely feeds it.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
Eighteen real purchase questions were put to five AI engines — ChatGPT,
Google Gemini, Perplexity, Google AI Mode, Google AI… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-search-visibility-romania-book-market.verified-ai-search-recommendations-telemetry
Verified AI Search Recommendations & Brand Mention Telemetry (2026)
Sample live telemetry dataset tracking B2B product search queries, cited domains, and the corresponding Share of Voice / recommendation percentage inside AI Search Engines (ChatGPT Search, Perplexity, Claude, Gemini).
Published by Pixel Office EU.
Purpose
This dataset demonstrates the correlation between website grounding (structured metadata / Fact Anchors) and the likelihood of being cited as… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/verified-ai-search-recommendations-telemetry.SearchBench-EvalAI-Paper-Similarity-Searchcomputer_science_ai_search_queriescomputer_science_non_ai_search_queriesai-marketplace-searchhadith-search-engine-dataai-search-datasetZNX-Search-5K
🚀 ZNX-Search-5K Dataset
ZNX-Search-5K adalah dataset percakapan/penalaran (Reasoning) premium yang dirancang khusus untuk mengoptimalkan kemampuan Large Language Models (LLM) seperti Qwen, Llama, dan Mistral dalam memahami instruksi kompleks, pencarian terstruktur, dan analisis siber/logika.
Dataset ini diformat menggunakan standar ShareGPT tingkat lanjut, lengkap dengan blok pemikiran internal (<think>...</think>) untuk melatih model agar memiliki kemampuan penalaran… See the full description on the dataset page: https://huggingface.co/datasets/ZNX-ai/ZNX-Search-5K.
