CoolFace
Datasetpublic

nilc-nlp/CATNA-MT

CATNA-MT (English version below) Os dados do CATNA estão originalmente disponíveis em http://tarsila.icmc.usp.br:8080/nurc/catna. O conjunto inclui 5 arquivos divididos em partes e 21 áudios completos. Esses 21 contêm um cabeçalho no início, indicando informações sobre a gravação, o qual não estava presente nos respectivos arquivos TextGrid. A partir de versões anteriores do CATNA, disponibilizadas pelos coordenadores do Projeto TaRSila… See the full description on the dataset page: https://huggingface.co/datasets/nilc-nlp/CATNA-MT.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes212downloads
Dataset Card

CATNA-MT

(English version below)

Os dados do CATNA estão originalmente disponíveis em http://tarsila.icmc.usp.br:8080/nurc/catna. O conjunto inclui 5 arquivos divididos em partes e 21 áudios completos. Esses 21 contêm um cabeçalho no início, indicando informações sobre a gravação, o qual não estava presente nos respectivos arquivos TextGrid. A partir de versões anteriores do CATNA, disponibilizadas pelos coordenadores do Projeto TaRSila (https://sites.google.com/view/tarsila-c4ai), foi possível alterar alguns arquivos, por meio da junção de alguns áudios divididos. Em outros casos, esse procedimento não foi possível. Nesses casos, cortes no início dos áudios foram realizados, com o objetivo de remover o cabeçalho e alinhar os dados ao TextGrid correspondente. O presente repositório é uma versão do CATNA nestas condições (Merged and Trimmed - MT).

Procedimento dos cortes do cabeçalho:

Na primeira tentativa, extraímos a duração total do áudio e subtraímos pela duração total do TextGrid. Entretanto, essa abordagem resultou em um alinhamento insatisfatório entre os arquivos. Diante disso, o seguinte procedimento foi adotado no Praat:

  • Extração da duração total do áudio (Query > Query time domain > Get total duration);
  • Visualização do áudio e seleção manual do trecho correspondente ao cabeçalho;
  • Cópia do valor referente ao cabeçalho;
  • Obtenção do áudio sem o cabeçalho (Convert > Extract part);
  • Realização de ajustes sucessivos até alcançar um alinhamento considerado satisfatório entre o áudio e o TextGrid.

Ressalta-se que um TextGrid só é plenamente adequado quando gerado a partir do seu áudio original. Então, os cortes realizados podem apresentar pequenas imprecisões temporais, dependendo do tipo de pesquisa que se quer realizar. Os cortes podem levar a pequenas diferenças numéricas de segundos. Essas diferenças podem ser significativas em estudos em Fonética e Fonologia. Apesar disso, o CATNA-MT não trará prejuízos para a pesquisa pela qual ele foi necessário. A seguir, há links para os áudios completos e os segundos que foram retirados de cada áudio. Isso permite que futuras pesquisas ou outros pesquisadores possam realizar novos cortes dos áudios, criar novos TextGrids ou revisar os que já aqui existem.

NomeSegundos
SPD205551.421281
SPD207850.772571
SPD209553.699801
SPD210948.684770
SPDID01367.940044
SPDID01668.053551
SPDID01760.301676
SPDID03046.728012
SPDID05342.256421
SPDID12140.9386196

Arquivos reconstruídos a partir da junção de versões anteriores do CATNA:

SPD2012

SPDID002

SPDID009

SPDID011

SPDID043

SPDID068

SPDID070

SPDID089

SPDID090

SPDID111

SPDID114

Arquivos que sofreram cortes:

SPD2055

SPD2078

SPD2095

SPD2109

SPDID013

SPDID016

SPDID017

SPDID030

SPDID053

SPDID121

Arquivos incompletos (mas sem cortes no meio de unidades entoacionais completas):

SPD2008parte2

SPD2010parte1

SPD2010parte3

SPD2023parte1

SPDID001parte2

SPDID044_parte

Patrocínios/Financiamento

Este dataset foi criado com apoio da Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brasil (CAPES) pela bolsa 88881.220186/2025-01.

CATNA-MT

The CATNA data are originally available at http://tarsila.icmc.usp.br:8080/nurc/catna. The dataset includes 5 files divided into parts and 21 complete audio files. These 21 contain a header at the beginning, indicating information about the recording, which was not present in the corresponding TextGrid files. From earlier versions of CATNA, made available by the TaRSila Project coordinators, it was possible to modify some files by merging split audio parts. In other cases, this procedure was not possible. In these cases, cuts were made at the beginning of the audio files in order to remove the header and align the data with the corresponding TextGrid. The present repository is a version of CATNA under these conditions (merged or trimmed - MT).

Header removal procedure:

In the first attempt, we extracted the total duration of the audio and subtracted the total duration of the TextGrid. However, this approach resulted in unsatisfactory alignment between the files. Therefore, the following procedure was adopted in Praat:

  • Extraction of the total audio duration (Query > Query time domain > Get total duration);
  • Visualization of the audio and manual selection of the segment corresponding to the header;
  • Copying the value corresponding to the header;
  • Obtaining the audio without the header (Convert > Extract part);
  • Performing successive adjustments until achieving an alignment considered satisfactory between the audio and the TextGrid. It should be noted that a TextGrid is only fully adequate when generated from its original audio. Thus, the cuts performed may introduce small temporal inaccuracies, depending on the type of research being conducted. These cuts may lead to small numerical differences in seconds. Such differences may be significant in studies in Phonetics and Phonology. Nevertheless, CATNA-MT will not cause harm to the research for which it was required. Below are links to the complete audios and the number of seconds removed from each file. This allows future research or other researchers to perform new cuts on the audios, create new TextGrids, or revise the existing ones.
NameSeconds
SPD205551.421281
SPD207850.772571
SPD209553.699801
SPD210948.684770
SPDID01367.940044
SPDID01668.053551
SPDID01760.301676
SPDID03046.728012
SPDID05342.256421
SPDID12140.9386196

Files reconstructed by merging previous versions of CATNA:

SPD2012

SPDID002

SPDID009

SPDID011

SPDID043

SPDID068

SPDID070

SPDID089

SPDID090

SPDID111

SPDID114

Files that underwent trimming:

SPD2055

SPD2078

SPD2095

SPD2109

SPDID013

SPDID016

SPDID017

SPDID030

SPDID053

SPDID121

Incomplete files (with no cuts within complete intonation units):

SPD2008parte2

SPD2010parte1

SPD2010parte3

SPD2023parte1

SPDID001parte2

SPDID044_parte

Sponsorships/Funding This dataset was created with support by the Coordination for the Improvement of Higher Education Personnel - Brazil (CAPES) through grant 88881.220186/2025-01.