adaptive-data
v2_data_full_adaptive_3layers_cross_attnMerged-Dual-Explain-stage_1-mhc-stream_expert_adaptive-7b_v2-data_v2v2_data_full_adaptive_3layers_cross_attn_geo_adalnDual-Explain-stage_1-mhc-stream_expert_no_adaptive-7b-data_v2Dual-Explain-stage_full-mhc-stream_expert_adaptive-7b_v2-data_v2v2_data_full_adaptive_12layers_cross_attnDual-Explain-stage_1-mhc-stream_expert_adaptive-7b-data_v2data_full_adaptive_3layers_cross_attn_both_adaln
ai-detector-data
AI Detector Predictions Dataset
A continuously-growing collection of AI text detection predictions with optional user feedback, generated from the AI Text Detector Space.
Every time someone analyzes text or a URL on the Space, the prediction is appended to this dataset. Users can also click "Correct" or "Incorrect" to provide feedback, which gets stored alongside the prediction.
Schema
Field
Type
Description
id
string
Unique 12-char hex identifier… See the full description on the dataset page: https://huggingface.co/datasets/adaptive-classifier/ai-detector-data.adaptive-adversaries-data
Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security
A 21-scenario multi-turn (15-round) adversarial red-teaming benchmark for LLM agents, in which both attacker and defender are independent LLM agents and attacks are regenerated per battle. Includes calibrated 3×3 attacker × defender matrix evaluation, full battle transcripts, attack-replay corpus, and traces from two open AgentBeats competitions.
Companion paper: Adaptive Adversaries: A Multi-Turn… See the full description on the dataset page: https://huggingface.co/datasets/neurips-adaptive-adversaries/adaptive-adversaries-data.adaptive-operator-v4-dataset
Adaptive Operator v4 — Training Datasets
Training data for adaptive-operator-v4, a Qwen3.5-9B model fine-tuned with custom control tokens for adaptive compute allocation in agentic workflows.
Dataset Summary
Split
Examples
Format
Size
SFT
4,992
OpenAI chat messages
8.8 MB
DPO
5,000
Chosen/rejected pairs
8.3 MB
Raw teacher responses
5,000
OpenAI chat messages
6.5 MB
Improved responses
4,992
OpenAI chat messages
13 MB
Preference pairs (raw)
5,000… See the full description on the dataset page: https://huggingface.co/datasets/davidnichols-ops/adaptive-operator-v4-dataset.adaptive-dice-maestro-dataset-mergednishy-al-biology-adaptive-dataset
Nishy A/L Biology Adaptive Dataset
This repository contains Biology MCQ datasets prepared for the Nishy adaptive tutoring and assessment system for Sri Lankan G.C.E. A/L Biology.
Repository structure
Master audited dataset
biology_master_1500_final_audited.json
Final audited master collection containing 1500 MCQs.
V4 paper-level split
v4_train_base_1182.json
v4_validation_base_90.json
v4_test_final_76.json
These files represent… See the full description on the dataset page: https://huggingface.co/datasets/Nishy11/nishy-al-biology-adaptive-dataset.adaptive-retro-gpt-1b-datastore
Adaptive-RETRO-GPT-1B Retrieval Datastore
External chunk datastore used by Adaptive-RETRO-GPT-1B during retrieval-pretraining.
Source: wikimedia/wikipedia / 20231101.en
Chunks: 120000
Chunk token window: 512
Retriever: signed token-hash inner product top-k
Files:
chunks.npy: tokenized retrieval chunks
embeddings.npy: normalized chunk embeddings
chunk_metadata.jsonl: decoded chunk text and metadata
datastore_meta.json: preprocessing and retrieval config
