datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FINDER_API_KEY_AI_SEARCH_2023
FINDER_API_KEY_AI_SEARCH_2023
tags: data collection, machine learning, API performance
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.r15-ai-search-metamerism
R15: AI Search Metamerism — Cross-Cultural Brand Perception Dataset
Citation: Zharnikov, D. (2026v) | DOI: 10.5281/zenodo.19422427 | Version: v3.2.0
Dataset Summary
This dataset contains the full session logs, aggregated results, and analysis outputs from the R15 large-scale experiment testing whether Large Language Models systematically collapse multi-dimensional brand perception into Economic and Experiential dimensions ("spectral metamerism"). It comprises 21… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/r15-ai-search-metamerism.MTSBerquadMTSBerquad is a cleaned and enriched dataset SberQuAD transferred to the Generative QA task. All entities were truecased, refactored by hand to improve readability and consistency. Answers have been expanded and rearranged from MLM QA task to Generative/Long Form QA task. MTSBerquad presented in PyCon 2024 by MTS AI Search Group.
Developed by MTS AI Search Group (Krayko Nikita, Laputin Fedor, Sidorov Ivan)
ameba_faq_search
AMEBA Blog FAQ Search Dataset
This data was obtained by crawling this website.
The FAQ Data was processed to remove HTML tags and other formatting after crawling, and entries containing excessively long content were excluded.
The Query Data was generated using a Large Language Model (LLM). Please refer to the following blog for information about the generation process.
https://www.ai-shift.co.jp/techblog/3710
https://www.ai-shift.co.jp/techblog/3761
Column description… See the full description on the dataset page: https://huggingface.co/datasets/ai-shift/ameba_faq_search.ai-search-visibility-romania-electronics-market
AI Search Visibility — Romania's Electronics & IT Market (August 2026)
18 brand-free purchase questions × 5 AI engines = 87 answers. 86 of them name a major retailer. Position, not presence, decides the market. Raw data CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
Eighteen real purchase questions were put to ChatGPT, Google Gemini, Perplexity, Google AI Mode and Google AI Overviews, in Romanian, from Romania, in… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-search-visibility-romania-electronics-market.ai-search-visibility-romania-book-market
AI Search Visibility — Romania's Book Market (July 2026)
When someone asks ChatGPT "which online bookstore should I use for children's
books?", they get one answer, not ten blue links. This dataset measures
who is inside that answer — and who merely feeds it.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
Eighteen real purchase questions were put to five AI engines — ChatGPT,
Google Gemini, Perplexity, Google AI Mode, Google AI… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-search-visibility-romania-book-market.cascade-ai-adversarial-search-simulator-v0.1Clarus Adversarial Cascade Simulator (Demo)
Configuration → Risk → Adversarial Search → Redesign
This repository demonstrates automated structural red teaming using cascade geometry.
The demo shows how a system configuration can be:
• Scored for cascade probability
• Stress-searched for near-threshold instability
• Converted into a safe sandbox scenario pack
• Redesigned to reduce structural risk
What This Repo Does
Most stress tools evaluate a single configuration.
This demo goes further.
It:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cascade-ai-adversarial-search-simulator-v0.1.cascade-ai-adversarial-search-simulator-v0.2
Clarus Adversarial Cascade Simulator v0.2
Adversarial boundary discovery for cascade-prone system configurations.
You provide a configuration.The simulator maps how close it is to systemic collapse.
Interactive Demo
Live Gradio interface available in Hugging Face Spaces.
Workflow:
Input baseline configuration (6 sliders)
Score configuration → View risk assessment
Run adversarial search → Discover worst-case boundary states
View scenario pack → Executable sandbox… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cascade-ai-adversarial-search-simulator-v0.2.
