datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trisearch-dataset-64k-v0.0.1
TriSearch-v1 (0.0.1)
Initial public data release for the TriSearch multimodal training stack
(v0.0.1). Curated image–text corpus for joint embedding
spaces (contrastive training + query-style text). Schema is intended to stay
stable; later versions may add fields or tighten quality filters.
Summary
Version
0.0.1 (format v1)
Examples
65,536 image–text records
Train / test
61,440 / 4,096 (test = 1/16 of data)
Image size
1024×1024 RGB JPEG… See the full description on the dataset page: https://huggingface.co/datasets/NuclearManD/trisearch-dataset-64k-v0.0.1.BCE-Prettybird-Micro-Standard-v0.0.1
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.1.BCE-Prettybird-Large-Standard-v0.0.1
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Large-Standard-v0.0.1.messi_mod-v0.0.1elizabeth-v0.0.1
Elizabeth v0.0.1 - Complete Model & Corpus Repository
🚀 Elizabeth Model v0.0.1
Model Files
models/qwen3_8b_v0.0.1_elizabeth_emergence.tar.gz - Complete Qwen3-8B model with Elizabeth's emergent personality
Training Data
corpus/elizabeth-corpus/ - 6 JSONL files with real conversation data
corpus/quantum_processed/ - 4 quantum-enhanced corpus files
Documentation
Comprehensive documentation of Elizabeth's emergence and capabilities:… See the full description on the dataset page: https://huggingface.co/datasets/LevelUp2x/elizabeth-v0.0.1.prm-dataset-v0.0.1
Verifiable Labs PRM dataset v0.0.1
Per-step process-reward training data for the Verifiable Labs SDK,
produced by the Phase 30 process-reward pipeline.
Stats
Traces: 120
Total steps: 120
Traces with per-step frontier judgments: 28
Frontier judge: anthropic/claude-sonnet-4 (when judged)
Source mix:
env: 92
judgment: 28
Schema
Each row is a JSON object with the following fields:
field
type
meaning
row_id
str
unique id
env_id
str
env… See the full description on the dataset page: https://huggingface.co/datasets/verifiablelabs/prm-dataset-v0.0.1.ste-sft-v0.0.1rm-dataset-v0.0.1
Verifiable Labs RM dataset v0.0.1
Reward-model training data for the Verifiable Labs SDK,
produced by the Phase 29 reward-distillation pipeline.
Stats
Rows: 840
With frontier judgment: 194
Frontier judge: anthropic/claude-sonnet-4 (when judged)
Source mix:
env: 646
judgment: 194
Schema
Each row is a JSON object with the following fields:
field
type
meaning
row_id
str
unique id
env_id
str
env that produced the row
prompt
str
task… See the full description on the dataset page: https://huggingface.co/datasets/verifiablelabs/rm-dataset-v0.0.1.ashercn97__a1-v0.0.1-details
Dataset Card for Evaluation run of ashercn97/a1-v0.0.1
Dataset automatically created during the evaluation run of model ashercn97/a1-v0.0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ashercn97__a1-v0.0.1-details.multi_turn_evaluation_v0.0.1AIObioEnts-v0.0.1-model_files
AIObioEnts model files
NOTE: these models are only compatible with the AIObioEnts v0.0.1
This dataset contains the model files for AIObioEnts, trained using AIONER with 4 different pre-trained models:
BiomedBERT-base pre-trained on abstracts from PubMed; the best-performing model reported in the original AIONER paper
BiomedBERT-base pre-trained on both abstracts from PubMed and full-texts articles from PubMedCentral
BioLinkBERT-base
BioLinkBERT large
for the identification of core… See the full description on the dataset page: https://huggingface.co/datasets/SIRIS-Lab/AIObioEnts-v0.0.1-model_files.flourish-images-and-data-v0.0.1dsl-arc-dataset-v0.0.1
DSL ARC Dataset
Dataset for training models on ARC-like tasks using a Domain Specific Language.
Dataset Structure
Each example contains:
train_input1, train_output1: First training example
train_input2, train_output2: Second training example
test_input, test_output: Test example
solution: DSL code to solve the task
task_type: Type of transformation
Task Types
Connect
MoveShape
RotateShape
CreateShape
FloodFill
MirrorShape
SymmetryComplete
ExtractPattern… See the full description on the dataset page: https://huggingface.co/datasets/middles/dsl-arc-dataset-v0.0.1.paso-por-paso-v0.0.1
