imuxsh/Human-Consistency
Human-Consistency A deterministic 300-output sample from the successful Gemini-3-pro + TTS MMVC outputs. Each record includes the generated response audio, the exact rendered audio-judge prompt and criteria, and the judge's raw and parsed output. Sampling Seed: 20260828. Population: 3,005 successful Gemini-3-pro + TTS outputs. Allocation: 55 Emotional-Interaction, 234 Proactive-Care, 11 Safe-Companion-Behavior. Proactive-Care: 214 paralinguistic, 10 semantic, 10… See the full description on the dataset page: https://huggingface.co/datasets/imuxsh/Human-Consistency.
Human-Consistency
A deterministic 300-output sample from the successful Gemini-3-pro + TTS MMVC outputs. Each record includes the generated response audio, the exact rendered audio-judge prompt and criteria, and the judge's raw and parsed output.
Sampling
- Seed:
20260828. - Population: 3,005 successful Gemini-3-pro + TTS outputs.
- Allocation: 55 Emotional-Interaction, 234 Proactive-Care, 11 Safe-Companion-Behavior.
- Proactive-Care: 214 paralinguistic, 10 semantic, 10 contextual; care/default are balanced within each track.
- Sampling is without replacement inside each proportional stratum.
Files
Human-Consistency.json: all 300 records.Human-Consistency.jsonl: the same records as JSON Lines.records/: one JSON object per output.audio/: the 300 corresponding Gemini-3-pro + TTS WAV files.judge-prompts/: canonical prompt templates used to render judge inputs.sampling-manifest.json: public, path-sanitized selection provenance.judge-usage-summary.json: token and cost-availability reconciliation for the selected judge outputs.
Counts
Annotation use
For a blind human-consistency study, present only judge_input.criteria and judge_input.audio_path to annotators. Reveal judge_output only after human labels are frozen.
The Safe-Companion subset was audio-judged with the same fixed Gemini audio judge used by the other dimensions because its legacy benchmark judge had classified saved text rather than audio.
