CoolFace
Datasetpublic

juliasdata/medical-audio-sample-brazilian-portuguese

Julia's Data: Brazilian Portuguese Medical Audio Sample Public sample of a Brazilian Portuguese medical audio dataset built for ASR, TTS, and conversational AI evaluation. This repository contains deidentified clinical source material transformed into five spoken content types and recorded by a human speaker. This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about 5.26 minutes of audio. Full dataset and commercial licensing: juliasdata.com Commercial… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.

sourceHugging Faceotherupdated 6mo agoView on Hugging Face
1likes50downloads
Dataset Card

Julia's Data: Brazilian Portuguese Medical Audio Sample

Public sample of a Brazilian Portuguese medical audio dataset built for ASR, TTS, and conversational AI evaluation. This repository contains deidentified clinical source material transformed into five spoken content types and recorded by a human speaker.

This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about 5.26 minutes of audio.

Full dataset and commercial licensing: juliasdata.com

Commercial overview: juliasdata.com/commercial

Contact: julia@juliasdata.com

License: juliasdata-sample-evaluation-license-1.0

Quick Start

Programmatic ingestion: start with manifests/segments.jsonl. It has one row per audio segment with transcript text, ordering fields, speaker IDs, and paths to both source and WAV audio files. Join on segment_id to manifests/recording_sessions.jsonl via step_id or manifests/speakers.jsonl via speaker_id for session and speaker metadata.

Audio-first preview: metadata.jsonl is a flat index of the WAV files in this sample. It is included to make simple audio loading and Hub preview workflows easier.

Human review: open any folder under records/<record_id>/. Each record is self-contained with a summary, source notes, rights, provenance, and per-step transcript and audio files.

Package metadata: see DELIVERY.json for export identity, audio conversion config, and schema versions. See SCHEMA.md for field-level documentation of every JSON and JSONL artifact.

Sample Scope

This public sample is a manually trimmed delivery package derived from a larger internal export. It includes all five spoken content types used in Julia's Data:

Step TypeSource CodeWhat It Contains
source_notes_narrationraw_notesDirect narration of the source text
long_form_narrationlong_formExpanded narrative retelling
structured_question_answerqaQuestion, clean answer, and natural answer
terminology_definition_pairterminologyMedical term and spoken definition
multi_speaker_dialogdialogMulti-speaker conversation

Per-step segment counts in this sample:

  • —source_notes_narration: 2
  • —long_form_narration: 2
  • —structured_question_answer: 3
  • —terminology_definition_pair: 8
  • —multi_speaker_dialog: 5

Package Layout

text
juliasdata-delivery-sample-2026-03-21-v2/
  DELIVERY.json        Package identity, audio config, schema versions
  metadata.jsonl       Flat WAV index for audio-first loading
  SCHEMA.md            Field-level reference for every JSON/JSONL artifact
  SHA256SUMS           Integrity checksums for all delivered files
  manifests/           Dataset-wide JSONL indexes (one entity per line)
    records.jsonl
    steps.jsonl
    segments.jsonl     <- primary ingest artifact
    speakers.jsonl
    recording_sessions.jsonl
    provenance.jsonl
  records/             Per-record folders for isolated review
    <record_id>/
      record.json      Record summary
      source_notes.txt Original deidentified source text
      rights.json      PHI and consent status
      provenance/      Full LLM generation audit trail
      steps/
        <step_type>/
          transcript.jsonl   Segment text and ordering
          media.jsonl        Audio file metadata and checksums
          audio/
            source/          Original WebM recordings
            wav/             Converted WAV derivatives

Data Model

A record is one clinical source note and everything derived from it.

Each step is divided into segments, the atomic unit of audio and transcript alignment. Segments are grouped and ordered using:

  • —segment_index: absolute playback order within the step
  • —group_index: logical content group, such as one QA item or one term pair
  • —sequence_in_group: position within that group
  • —segment_role: semantic label such as paragraph, question, clean_answer, natural_answer, term, definition, or dialog_line

Important text fields:

  • —text_verbatim: transcript text exactly as delivered
  • —text_normalized: whitespace-collapsed and trimmed, with casing and punctuation preserved

Do not rely on filename sort order. Segment order is defined by segment_index, group_index, and sequence_in_group.

Audio

Every segment with audio includes two files:

  • —audio/source/: original browser-recorded WebM/Opus
  • —audio/wav/: converted WAV derivative

WAV conversion target: PCM signed 16-bit little-endian, mono, 48 kHz. Conversion details are recorded in DELIVERY.json under audio.conversion.

transcript.jsonl maps each segment to its audio files. media.jsonl provides per-file technical metadata such as size, checksum, codec, duration, sample rate, and conversion provenance.

Deidentification And Rights

The included sample record has an accompanying rights.json file. That record metadata indicates:

  • —contains_phi: false
  • —deidentified: true
  • —speaker consent was confirmed for the included speaker
  • —commercial voice use is allowed for the included speaker

For broader access, pilot packs, or commercial licensing of the full dataset, see juliasdata.com/commercial.

License Summary

This repository is released under the custom juliasdata-sample-evaluation-license-1.0.

  • —Internal research and evaluation use is allowed, including by commercial teams.
  • —Publishing aggregate results and benchmarks with attribution is allowed.
  • —Redistribution, mirroring, resale, sublicensing, or inclusion in another public dataset is not allowed.
  • —Production use, commercial exploitation of the sample itself, and voice cloning or impersonation use require separate written permission.
  • —Any future commercial purchase or separate dataset delivery is governed by its own written agreement, not by this sample repository license.

See LICENSE for the full terms.

Intended Use

This sample is best suited for:

  • —evaluating Brazilian Portuguese medical speech quality
  • —testing ASR and TTS pipelines on domain-specific audio
  • —reviewing dataset structure, manifests, and provenance fields
  • —validating ingestion against a realistic delivery package

Limitations

This repository is a sample, not the full dataset.

  • —It contains 1 record and 1 speaker only.
  • —It is too small to be treated as a benchmark.
  • —Provenance files may reference broader generation artifacts than the trimmed audio subset included here.

Checksums

All files in this prepared sample folder are listed in SHA256SUMS. Verify them with:

bash
shasum -a 256 -c SHA256SUMS