juliasdata/medical-audio-sample-brazilian-portuguese
Julia's Data: Brazilian Portuguese Medical Audio Sample Public sample of a Brazilian Portuguese medical audio dataset built for ASR, TTS, and conversational AI evaluation. This repository contains deidentified clinical source material transformed into five spoken content types and recorded by a human speaker. This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about 5.26 minutes of audio. Full dataset and commercial licensing: juliasdata.com Commercial… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.
Julia's Data: Brazilian Portuguese Medical Audio Sample
Public sample of a Brazilian Portuguese medical audio dataset built for ASR, TTS, and conversational AI evaluation. This repository contains deidentified clinical source material transformed into five spoken content types and recorded by a human speaker.
This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about 5.26 minutes of audio.
Full dataset and commercial licensing: juliasdata.com
Commercial overview: juliasdata.com/commercial
Contact: julia@juliasdata.com
License: juliasdata-sample-evaluation-license-1.0
Quick Start
Programmatic ingestion: start with manifests/segments.jsonl. It has one row per audio segment with transcript text, ordering fields, speaker IDs, and paths to both source and WAV audio files. Join on segment_id to manifests/recording_sessions.jsonl via step_id or manifests/speakers.jsonl via speaker_id for session and speaker metadata.
Audio-first preview: metadata.jsonl is a flat index of the WAV files in this sample. It is included to make simple audio loading and Hub preview workflows easier.
Human review: open any folder under records/<record_id>/. Each record is self-contained with a summary, source notes, rights, provenance, and per-step transcript and audio files.
Package metadata: see DELIVERY.json for export identity, audio conversion config, and schema versions. See SCHEMA.md for field-level documentation of every JSON and JSONL artifact.
Sample Scope
This public sample is a manually trimmed delivery package derived from a larger internal export. It includes all five spoken content types used in Julia's Data:
Per-step segment counts in this sample:
source_notes_narration: 2long_form_narration: 2structured_question_answer: 3terminology_definition_pair: 8multi_speaker_dialog: 5
Package Layout
juliasdata-delivery-sample-2026-03-21-v2/
DELIVERY.json Package identity, audio config, schema versions
metadata.jsonl Flat WAV index for audio-first loading
SCHEMA.md Field-level reference for every JSON/JSONL artifact
SHA256SUMS Integrity checksums for all delivered files
manifests/ Dataset-wide JSONL indexes (one entity per line)
records.jsonl
steps.jsonl
segments.jsonl <- primary ingest artifact
speakers.jsonl
recording_sessions.jsonl
provenance.jsonl
records/ Per-record folders for isolated review
<record_id>/
record.json Record summary
source_notes.txt Original deidentified source text
rights.json PHI and consent status
provenance/ Full LLM generation audit trail
steps/
<step_type>/
transcript.jsonl Segment text and ordering
media.jsonl Audio file metadata and checksums
audio/
source/ Original WebM recordings
wav/ Converted WAV derivativesData Model
A record is one clinical source note and everything derived from it.
Each step is divided into segments, the atomic unit of audio and transcript alignment. Segments are grouped and ordered using:
segment_index: absolute playback order within the stepgroup_index: logical content group, such as one QA item or one term pairsequence_in_group: position within that groupsegment_role: semantic label such asparagraph,question,clean_answer,natural_answer,term,definition, ordialog_line
Important text fields:
text_verbatim: transcript text exactly as deliveredtext_normalized: whitespace-collapsed and trimmed, with casing and punctuation preserved
Do not rely on filename sort order. Segment order is defined by segment_index, group_index, and sequence_in_group.
Audio
Every segment with audio includes two files:
audio/source/: original browser-recorded WebM/Opusaudio/wav/: converted WAV derivative
WAV conversion target: PCM signed 16-bit little-endian, mono, 48 kHz. Conversion details are recorded in DELIVERY.json under audio.conversion.
transcript.jsonl maps each segment to its audio files. media.jsonl provides per-file technical metadata such as size, checksum, codec, duration, sample rate, and conversion provenance.
Deidentification And Rights
The included sample record has an accompanying rights.json file. That record metadata indicates:
contains_phi: falsedeidentified: true- speaker consent was confirmed for the included speaker
- commercial voice use is allowed for the included speaker
For broader access, pilot packs, or commercial licensing of the full dataset, see juliasdata.com/commercial.
License Summary
This repository is released under the custom juliasdata-sample-evaluation-license-1.0.
- Internal research and evaluation use is allowed, including by commercial teams.
- Publishing aggregate results and benchmarks with attribution is allowed.
- Redistribution, mirroring, resale, sublicensing, or inclusion in another public dataset is not allowed.
- Production use, commercial exploitation of the sample itself, and voice cloning or impersonation use require separate written permission.
- Any future commercial purchase or separate dataset delivery is governed by its own written agreement, not by this sample repository license.
See LICENSE for the full terms.
Intended Use
This sample is best suited for:
- evaluating Brazilian Portuguese medical speech quality
- testing ASR and TTS pipelines on domain-specific audio
- reviewing dataset structure, manifests, and provenance fields
- validating ingestion against a realistic delivery package
Limitations
This repository is a sample, not the full dataset.
- It contains 1 record and 1 speaker only.
- It is too small to be treated as a benchmark.
- Provenance files may reference broader generation artifacts than the trimmed audio subset included here.
Checksums
All files in this prepared sample folder are listed in SHA256SUMS. Verify them with:
shasum -a 256 -c SHA256SUMS