Flaglab/esnlir-llm-predictions
ESNLIR-LLM — per-pair predictions Per-pair predictions for every model evaluated in An Analysis of the Performance of Large Language Models in Spanish NLI Datasets with Causal Relationships (IBERAMIA 2026, to appear). Code in Pacolas/NLI-via-LLM; part of the ESNLIR-LLM collection. These are the raw outputs behind the paper's tables, so results can be re-scored, sliced by genre or domain, or compared pair by pair without re-running any model. Files… See the full description on the dataset page: https://huggingface.co/datasets/Flaglab/esnlir-llm-predictions.
Declare a config per model so the dataset viewer works
Store metrics at six decimals; refresh the recomputed tables
Document the verification script and the statistical conventions
Add recomputed metrics (Wilson CIs, bootstrap CIs, McNemar) and the script that produces them
Document which API run each prediction file contains
Flag that the GPT-4o validated file is a different run from the paper's
Document the new GPT-4o-Mini and GPT-4o human-validated prediction files
Add human-validated (972) predictions for GPT-4o-Mini and GPT-4o
Document GPT-4o-Mini and note what is unavailable
Add GPT-4o-Mini per-pair predictions on the full test set
Add dataset card
Add per-pair predictions for Qwen2.5-7B, Llama-3.1-8B and XLM-RoBERTa
initial commit
