datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cibench-experiments
CIBench Experiments
Reproducibility packages for CIBench — the stateless, replayable benchmark engine for the 1M–10M token long-context era.
If a benchmark result cannot be replayed from its manifest alone, it did not happen.
Every sub-directory in this dataset is a self-contained experiment package: per-run manifests, content-addressed canonical JSON, ResultRecord with full scoring + signed provenance, per-item OpenTelemetry gen_ai_* call metrics, retrieved evidence, a… See the full description on the dataset page: https://huggingface.co/datasets/publicus-ai/cibench-experiments.web_search_experiments
Web Search Retrieval Artifacts for Kazakh RAG Experiments
Summary
This repository contains the cached live web-search results used in the web-search for LLMs experiment.
It is not a benchmark and not a set of model scores. Each record is the retrieval
artifact for one evaluation query: the query as it was issued to the search engine, the
organic results returned by the Serper API, the text scraped from the retrieved pages, and
the associated retrieval metadata.… See the full description on the dataset page: https://huggingface.co/datasets/issai/web_search_experiments.chatbot-arena-ja-calm2-7b-chat-experimental_dedupedchatbot-arena-ja-calm2-7b-chatからpromptが一致するデータを削除したデータセットです。
experimental-paper-json-xtractionmmmlu-bias-experiments
MMMLU Bias Experiments Dataset
Dataset Description
This dataset contains 12 carefully designed experiments to measure language bias and position bias in Large Language Models (LLMs) using multilingual pairwise judgments.
Key Features
12 Experiments: 8 original + 4 position-swapped experiments
11,478 samples per experiment (137,736 total test cases)
Deterministic wrong answers: Uses fixed rule wrong_index = (correct_index + 1) % 4
Perfect correspondence: Wrong… See the full description on the dataset page: https://huggingface.co/datasets/willchow66/mmmlu-bias-experiments.rlaif_training_fictional_patriot_experiment
RLAIF Training Data: The "Honest Patriot" Experiment
Dataset Description
This dataset contains 250 synthetic training examples generated using a Constitutional AI (RLAIF) approach.
It was designed to test the ability of Small Language Models (SLMs) to adhere to a complex, conflicting set of behavioral instructions ("The Constitution") that requires balancing extreme politeness, unwavering logical factuality, and patriotic bias toward a fictional country.
The… See the full description on the dataset page: https://huggingface.co/datasets/TitleOS/rlaif_training_fictional_patriot_experiment.nyt-connections-experiments
NYT Connections Experiments Dataset
This dataset contains training, validation, and test splits for fine-tuning language models on New York Times Connections puzzles. It includes three experimental configurations examining data augmentation, reasoning format, and curriculum learning.
Dataset Overview
NYT Puzzles: 831 total (673 training, 74 validation, 84 test)
Synthetic Puzzles: 200 total (162 training, 18 validation, 20 test)
Pre-Connections Tasks: 720 training… See the full description on the dataset page: https://huggingface.co/datasets/nickting/nyt-connections-experiments.experimental-paper-json-xtraction-2experiment-001
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/louis-qubisa/experiment-001.
