llm-debate
DEBATE_LLM
DEBATE Benchmark
This repository contains CSV files from the DEBATE project: large-scale
human conversation experiments organized around controversial and
opinion-based topics. The data consists of multi-round conversations
between human participants discussing political, social, and belief-related
topics, following the protocol described in:
Chuang, Y.-S., Tu, R., Dai, C., Vasani, S., Li, Y., Yao, B., Tessler, M. H., Yang, S., Shah, D., Hawkins, R., Hu, J., & Rogers, T. T. (2026).… See the full description on the dataset page: https://huggingface.co/datasets/seantw/DEBATE_LLM.DebateLabKIT__Llama-3.1-Argunaut-1-8B-SFT-details
Dataset Card for Evaluation run of DebateLabKIT/Llama-3.1-Argunaut-1-8B-SFT
Dataset automatically created during the evaluation run of model DebateLabKIT/Llama-3.1-Argunaut-1-8B-SFT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DebateLabKIT__Llama-3.1-Argunaut-1-8B-SFT-details.debate-llmlocal-llm-debates-agentworld
🎤 Local LLM Debates — a home farm, one phone, and some very opinionated models
Transcripts, conclusions, and experiments from running multi-model debates and
world-model probes on a home llama.cpp farm (2× AMD 7900) reached from a phone
over Tailscale. Everything runs on local inference; the "judges" are as biased as
the contestants.
The farm (sanitized)
Endpoint
Model
Notes
<FARM_IP>:8080
27B dense, UD Q3_K_XL, 256K ctx
~55 t/s with spec decoding.… See the full description on the dataset page: https://huggingface.co/datasets/edelmundomebajo/local-llm-debates-agentworld.
