CoolFace
Datasetpublic

Om22s/alexandria-attribution-provenance-20260913

Alexandria speaker-attribution evaluation provenance This release contains aggregate, paired evaluation measurements for a Qwen3-14B speaker-attribution LoRA. It is a reproducibility record, not a training or evaluation corpus. Contents results.json contains aggregate counts, accuracy, strict shared-row counts, decoding settings, and evaluator commit. provenance.json maps each aggregate record to the SHA-256 of its private source artifact and to hashes of the… See the full description on the dataset page: https://huggingface.co/datasets/Om22s/alexandria-attribution-provenance-20260913.

sourceHugging Facemitupdated 7d agoView on Hugging Face
0likes62downloads
Dataset Card

Alexandria speaker-attribution evaluation provenance

This release contains aggregate, paired evaluation measurements for a Qwen3-14B speaker-attribution LoRA. It is a reproducibility record, not a training or evaluation corpus.

Contents

  • —results.json contains aggregate counts, accuracy, strict shared-row counts, decoding settings, and evaluator commit.
  • —provenance.json maps each aggregate record to the SHA-256 of its private source artifact and to hashes of the four private gold files.

It deliberately excludes all source text, quote context, character names, gold labels, prompts, raw model responses, audio, model weights, and training examples. The private source materials are not licensed or distributed by this repository; their hashes are included solely to make independently held copies identifiable.

Measurements

Both runs evaluate 768 requested predictions per arm, use deterministic decoding (temperature: 0), and compare the same base model against the same mixed LoRA adapter. The batch=1 and batch=25 results are separate serving configurations, not independent training replications.

Serving batchBase accuracyAdapter accuracyDifference
145.70%59.51%+13.80 points
2560.94%68.75%+7.81 points

These are measurements from the named artifacts, not a claim of general speaker-attribution performance. The gold set contains four non-redistributed translated-light-novel subsets, so these results must not be compared directly to public-domain or other external datasets.

Reproduction boundary

The evaluator was run from Alexandria commit 6ed991121ce7d54547553e49558aea67abc45fec, against Qwen3-14B with fixed decoding. A party that independently has lawful access to the same evaluation inputs can compare its own artifact SHA-256 values to provenance.json.

License

The aggregate facts and release documentation are MIT-licensed. This license does not grant rights to any omitted private source material, base model, adapter, or third-party corpus.