Hallucination
resultsrequestshallucination-portability-revision
Hallucination probe portability revision
First GPU round is ready: new representation corpora for Ministral-3-8B-Base-2512 and Qwen3-8B, using pinned model/dataset revisions. Each has 5,000 training, 1,000 development, and 1,000 test SQuAD v2 examples. Development/test contexts do not overlap. This is a new controlled training/adaptation dataset, not a claim to reproduce the old split.
Run in your independent terminal:
cd /home/jovyan/hallucination-portability
export… See the full description on the dataset page: https://huggingface.co/datasets/ShoaibSSM/hallucination-portability-revision.FinQA-hallucination-detection
FinQA Hallucination Detection
Dataset Summary
This dataset was created from a subset of the original FinQA dataset. For each user query (financial questions), we prompted an LLM to generate a response to this query based on provided context (financial statements and tables from the original FinQA).
Each generated LLM response is labeled based on whether it is correct or not. This dataset is thus useful for benchmarking reference-free LLM Eval and Hallucination… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/FinQA-hallucination-detection.hallucinations-dpoWhisper-Hallucination
Whisper Hallucination and Repetition Probes
This is a benchmark. Every evaluation config is test — do not fine-tune on it.
lexicon_synth is the exception: synthetic training material with its own train/test
split, and not one of the eight benchmark arms.
To build training data, exclude the items in
benchmark/exclusions.json
(546 FMA tracks, 1,168 FSD50K ids, 2,620 LibriSpeech utterances, the Malay stems). The
benchmark draws FSD50K eval and FMA shards 0–1, so training can use… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Whisper-Hallucination.
