datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scalelab-benchmark-results
ScaleLab Benchmark Results
Reproducible benchmark results comparing optimization algorithms across industrial process instances.
Benchmark Schema
Each result record contains:
Field
Description
instance_id
Instance identifier
algorithm
Optimization algorithm used
solver
Surrogate/solver backend
configuration
Algorithm configuration
seed
Random seed
runtime_sec
Total runtime
time_to_first_solution
Time to reach target quality
n_experiments… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/scalelab-benchmark-results.benchmark-evaluation-resultsturkish-seo-reasoning-benchmark-results
Turkish SEO Reasoning Benchmark Results
Bu dataset, Turkish SEO Reasoning benchmark'ının altı farklı model/checkpoint üzerinde çalıştırılmış ham tahminlerini, metriklerini ve tekrar üretim manifestlerini içerir.
Fine-tuned model: berkbirkan/gemma-3-1b-turkish-seo-reasoning-lora
Sonuç
Fine-tuned Gemma 3 1B modeli 22,23 skorla ilk sırada yer aldı. Aynı base model 11,96 skor elde etti.
Mutlak artış: +10,28 puan
Göreli artış: %85,97
Fine-tuned model hata sayısı:… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning-benchmark-results.colliderml-benchmark-results
ColliderML Benchmark Results
Machine-scored leaderboard results for the ColliderML benchmark tasks.
Structure
results/
{task}/
{username}/
{predictions_sha256}.json
Each JSON file records one scored submission:
Field
Description
submission_id
UUID assigned by the backend
task
Benchmark task name (e.g. tracking)
submitter
HuggingFace username
model_repo_id
Optional link to the submitter's model repo
submitted_at
ISO 8601 timestamp
scores… See the full description on the dataset page: https://huggingface.co/datasets/CERN/colliderml-benchmark-results.hospital-operations-benchmark-results
Hospital Operations — Benchmark Results
Pre-computed policy comparison across manual FCFS, deterministic mean, robust quantile, and rolling horizon scheduling.
Runs: 32Policies: 4
port-terminal-benchmark-results
Port Terminal — Benchmark Results
Pre-computed policy comparison across FCFS, CP-SAT integrated, robust quantile, ALNS matheuristic, and NSGA-II multi-objective scheduling.
Runs: 40Policies: 5
crf-benchmark-results
CRF Benchmark Results
Benchmark results comparing the Cellular Reasoning Fabric (CRF) against a parameter-matched Transformer across 5 language modeling tasks.
Dataset Description
This dataset contains the full experimental results from our CRF vs Transformer comparison, including:
Training loss and perplexity curves (per epoch)
Validation loss and perplexity curves (per epoch)
Parameter counts and FLOP estimates
Inference profiling (latency, memory)
CRF-specific… See the full description on the dataset page: https://huggingface.co/datasets/YasirUsman/crf-benchmark-results.gemma-french-asr-benchmark-resultssage-benchmark-resultsRN_TR_R2_Benchmark_Results
RefinedNeuro/RN_TR_R2 Turkish Culture & Reasoning Benchmark
This repository contains the results of a custom benchmark designed to evaluate the performance of open-source language models on Turkish culture questions and basic reasoning tasks.
Overview
We crafted a set of 25 questions covering:
Turkish general knowledge (e.g., capital city, national holidays, geography)
Basic arithmetic and logic puzzles
Simple calculus and string-processing tasks
Each question is paired… See the full description on the dataset page: https://huggingface.co/datasets/RefinedNeuro/RN_TR_R2_Benchmark_Results.
