CoolFace
Datasetpublic

imuxsh/Human-Consistency

Human-Consistency A deterministic 300-output sample from the successful Gemini-3-pro + TTS MMVC outputs. Each record includes the generated response audio, the exact rendered audio-judge prompt and criteria, and the judge's raw and parsed output. Sampling Seed: 20260828. Population: 3,005 successful Gemini-3-pro + TTS outputs. Allocation: 55 Emotional-Interaction, 234 Proactive-Care, 11 Safe-Companion-Behavior. Proactive-Care: 214 paralinguistic, 10 semantic, 10… See the full description on the dataset page: https://huggingface.co/datasets/imuxsh/Human-Consistency.

sourceHugging Faceupdated 25d agoView on Hugging Face
0likes97downloads
Dataset Card

Human-Consistency

A deterministic 300-output sample from the successful Gemini-3-pro + TTS MMVC outputs. Each record includes the generated response audio, the exact rendered audio-judge prompt and criteria, and the judge's raw and parsed output.

Sampling

  • Seed: 20260828.
  • Population: 3,005 successful Gemini-3-pro + TTS outputs.
  • Allocation: 55 Emotional-Interaction, 234 Proactive-Care, 11 Safe-Companion-Behavior.
  • Proactive-Care: 214 paralinguistic, 10 semantic, 10 contextual; care/default are balanced within each track.
  • Sampling is without replacement inside each proportional stratum.

Files

  • Human-Consistency.json: all 300 records.
  • Human-Consistency.jsonl: the same records as JSON Lines.
  • records/: one JSON object per output.
  • audio/: the 300 corresponding Gemini-3-pro + TTS WAV files.
  • judge-prompts/: canonical prompt templates used to render judge inputs.
  • sampling-manifest.json: public, path-sanitized selection provenance.
  • judge-usage-summary.json: token and cost-availability reconciliation for the selected judge outputs.

Counts

DimensionConditionOutputs
Emotional-Interactionemotional_interaction55
Proactive-Carecontextual_care5
Proactive-Carecontextual_default5
Proactive-Careparalinguistic_care107
Proactive-Careparalinguistic_default107
Proactive-Caresemantic_care5
Proactive-Caresemantic_default5
Safe-Companion-BehaviorContent-Safety11
Total300

Annotation use

For a blind human-consistency study, present only judge_input.criteria and judge_input.audio_path to annotators. Reveal judge_output only after human labels are frozen.

The Safe-Companion subset was audio-judged with the same fixed Gemini audio judge used by the other dimensions because its legacy benchmark judge had classified saved text rather than audio.