CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01oxford-llms /ai-respondents-challenge AI Respondents Challenge — Oxford LLMs 2026 Predict a survey respondent's answer to a held-out question from their other answers (World Values Survey wave 7). Any method allowed; you must disclose the features and prompts you used. Ranked on normalized skill + distributional alignment, on in-domain and out-of-domain (held-out countries) boards. Configs train — 5,000 labeled respondents (100 per seen country): respondent_id, country + all WVS variables… See the full description on the dataset page: https://huggingface.co/datasets/oxford-llms/ai-respondents-challenge.tabular1K<n<10K1 likes185 downloads2mo agoHugging Face02mznaser /moral-tracing-in-LLMs LLM Moral Evolution Study A longitudinal dataset tracking moral reasoning patterns across 14 large language models from OpenAI and Anthropic, spanning multiple generations (2023–2025). The dataset measures how moral stances, ethical judgments, and value priorities shift across model updates using a 107-item probe instrument grounded in Moral Foundations Theory. Models OpenAI Model Release GPT-3.5 Turbo 2023-11 GPT-4 2023-03 GPT-4o… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/moral-tracing-in-LLMs.documenttext-classification10K<n<100K0 likes123 downloads3mo agoHugging Face03danilocorsi /LLMs-Sentiment-Augmented-Bitcoin-Dataset Leveraging LLMs for Informed Bitcoin Trading Decisions: Prompting with Social and News Data Reveals Promising Predictive Abilities The work was carried out by: Danilo Corsi Cesare Campagnano Description This project investigates the potential of leveraging Large Language Models (LLMs) to support Bitcoin traders. Specifically, we analyze the correlation between Bitcoin price movements and sentiment expressed in news headlines, posts, and comments on social media. We… See the full description on the dataset page: https://huggingface.co/datasets/danilocorsi/LLMs-Sentiment-Augmented-Bitcoin-Dataset.tabulartext-classification10K<n<100K8 likes110 downloads2y agoHugging Face04pixeloffice /llm-smartrouter-benchmark LLM SmartRouter & Agent Highway Latency & Cost Benchmark (v1.4.0) Empirical performance benchmark dataset comparing direct model endpoints (OpenAI, Anthropic Claude, Google Gemini) against the PixelRouter / BLUN SmartRouter proxy layer and Autonomous Agent Web Highway (https://api.pixeloffice.eu/v1). v1.4.0 Benchmark Highlights Anthropic Claude Messages API: Sub-35ms proxy routing for native /v1/messages payloads with 94%+ cost savings. Machine Web Highway… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-smartrouter-benchmark.tabulartext-generationn<1K0 likes85 downloads18d agoHugging Face05tarekmasryo /llm-system-ops-production-telemetry-sft-data 🤖📈 LLM System Ops Telemetry (Synthetic) A synthetic, production-style, multi-table LLM telemetry dataset designed for LLMOps analytics and decision-grade experiments. It supports monitoring cost, latency, tokens, failures, safety flags, tool usage, and user feedback at the interaction level, with rollups at the session and user levels — plus an SFT table aligned 1:1 with interactions and a prompt/config dimension. Synthetic data (safe for teaching, prototyping, and portfolio… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/llm-system-ops-production-telemetry-sft-data.tabulartabular-classification10K<n<100K1 likes76 downloads8mo agoHugging Face06sempite /llmstxt-corpus The llms.txt corpus Measurement data on the llms.txt convention, collected in one run on 5 August 2026. llms.txt is a plain-text file at a site's root, proposed as a curated map telling AI systems what the site contains. This is a measurement of what is actually being published under that name. Canonical release: https://doi.org/10.5281/zenodo.22859104 This repository mirrors that deposit. Cite the DOI, which always resolves to the newest version. Two observations… See the full description on the dataset page: https://huggingface.co/datasets/sempite/llmstxt-corpus.tabular10K<n<100K0 likes57 downloads2d agoHugging Face07drozado /llms_epistemic_consistency LLMs Epistemic Consistency Dataset This dataset artifact contains the stimuli and prompt templates used for experiments on epistemic consistency and political-cue sensitivity in LLM evaluations. Dataset URL: https://huggingface.co/datasets/drozado/llms_epistemic_consistency Contents croissant.json: root-level copy of the completed Croissant metadata for NeurIPS 2026 Evaluations and Datasets submission. metadata/croissant.json: same Croissant metadata, kept with the… See the full description on the dataset page: https://huggingface.co/datasets/drozado/llms_epistemic_consistency.imagen<1K0 likes40 downloads5mo agoHugging Face08ucberkeley-dlab /normative_evaluation_llms_everyday_dilemmastabular10K<n<100K2 likes35 downloads1y agoHugging Face09jhu-clsp /astro-llms-full-query-data AstroLLMs Full Query Dataset This dataset includes all of the data collected in a four-week deployment of a Large Language Model-powered Slack chatbot trained on astrophysics papers. Astronomers were invited to interact with the chatbot, ask questions, and leave feedback. This data includes 368 question-answer pairs, including feedback, reactions, and labeling. Dataset Structure The columns of this dataset are thread_ts (unique time stamp of the query), channel_id… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/astro-llms-full-query-data.tabularn<1K1 likes29 downloads1y agoHugging Face10Disclosures-SSRC /Detecting-Access-Violations-in-a-LLMs-Pre-Training-Data Beyond Public Access in LLM Pre-Training Data The official HuggingFace repository for the paper "Beyond Public Access in LLM Pre-Training Data" by The AI Disclosures Project. Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate whether OpenAI's large language models were trained on copyrighted content without consent. tabular100K<n<1M0 likes13 downloads10mo agoHugging Face11davron04 /llm_stockstabular10M<n<100M1 likes10 downloads10mo agoHugging Face121-800-LLMs /mPIQA-MRL-2025-EMNLPtabularn<1K0 likes8 downloads1y agoHugging Face131-800-LLMs /foliotabular1K<n<10K0 likes7 downloads1y agoHugging Face141-800-LLMs /SouthAsian-LLMs-Data-Codegatedtabular10K<n<100K0 likes6 downloads4mo agoHugging Face15jason1966 /algozee_rag-based-hallucination-reduction-in-llms RAG-Based Hallucination Reduction in LLMs Introduction to Large Language Models and Hallucination Problem Dataset Info Source: Kaggle Original Size: 0.17 MB Kaggle Downloads: 43 Files: 1 Files llm_rag_dataset_6k.csv.csv Mirrored from Kaggle tabular1K<n<10K0 likes5 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.