CoolFace
Datasetpublic

Rapidata/kiki-bouba-audio-20k

πŸ”Š Kiki–Bouba, Spoken Aloud (20k Global Responses) Dataset Summary This dataset is the audio companion to Rapidata/psychology-association-kiki-bouba-etc. In the original dataset, respondents read the question "Which one is called 'Kiki'?" as written text. Here, respondents instead hear the word spoken aloud β€” the task shows the same two shapes (a rounded blob and a spiky star) while a short audio clip of "kiki" or "bouba" plays as context. The annotator UI… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/kiki-bouba-audio-20k.

sourceHugging Faceupdated 1mo agoView on Hugging Face
2likes151downloads
Dataset Card

πŸ”Š Kiki–Bouba, Spoken Aloud (20k Global Responses)

Dataset Summary

This dataset is the audio companion to Rapidata/psychology-association-kiki-bouba-etc. In the original dataset, respondents read the question "Which one is called 'Kiki'?" as written text. Here, respondents instead hear the word spoken aloud β€” the task shows the same two shapes (a rounded blob and a spiky star) while a short audio clip of "kiki" or "bouba" plays as context. The annotator UI requires the clip to finish playing before an answer can be given.

The Bouba–Kiki effect is fundamentally a claim about sound symbolism β€” yet most large-scale collections (ours included) presented the word in writing, filtered through each language's orthography. This dataset removes that filter.

  • β€”20,000 responses (~10,000 per word)
  • β€”2 spoken words: "kiki" and "bouba", same speaker (Waithera Were, Wikimedia Commons, CC BY-SA 4.0)
  • β€”Same shape images as the text dataset, pixel-identical
  • β€”Respondent-level country, language, age, gender, occupation metadata
  • β€”A hearing check: an audio validation set ("can you actually hear this?") was attached to the job, and each response carries the respondent's userScore_audio reliability score

Data collected with the **Rapidata API** in August 2026. Please consider leaving a ❀️ if you find this dataset interesting.


Noteworthy Findings

1. Hearing the word restores the classic effect that writing reversed

In the text dataset, the global aggregate for "Which one is called 'Kiki'?" came out marginally reversed β€” 51.6% chose the blob as "Kiki". With the word spoken aloud, the classic direction returns: 55.2% choose the spiky shape for "kiki", and 62.2% choose the blob for "bouba".

[image]

Interestingly, the modality shift is not symmetric: "bouba" was more strongly associated with the blob when written (72.7%) than when spoken (62.2%), while "kiki" flips from reversed to classic. With audio, the two words land at similar, moderately-congruent levels instead of the text version's asymmetry.

2. The Japanese–Arabic contrast survives the modality change

The most striking pattern in the text dataset β€” Japanese speakers overwhelmingly picking the spiky shape as "Kiki" while Arabic speakers lean the other way β€” replicates with spoken audio, so it is not an artifact of how the two scripts render the word. Japanese speakers: 78% spiky (audio) vs 74% (text). Arabic speakers remain the only large group below chance: 42% spiky (audio) vs 29% (text) β€” audio attenuates but does not eliminate the reversal.

[image]

3. Country extremes: a 29-point spread, with one country inverted

Pooling both words, 58.7% of responses pick the shape the classic effect predicts. That average hides a wide spread across countries (β‰₯200 responses each):

  • β€”Strongest: Japan 72.9%, Ecuador 69.3%, Peru 67.1%
  • β€”Weakest: Algeria 44.0%, Egypt 51.2%, Iraq 52.2%

Algeria is the only country whose confidence interval sits entirely below chance (44.0%, 95% CI 41–47%) β€” its respondents actively prefer the opposite pairing rather than merely guessing. Iraq and Egypt land at chance, i.e. no detectable association either way.

[image]

The dumbbell also shows which word carries the effect, and that it differs by country. Japan is driven by "kiki" (78% spiky vs 67% blob for "bouba"), while several Spanish- and Portuguese-speaking countries show the reverse profile β€” "bouba" is the stronger of the two. In the countries at or below chance, the deficit is almost entirely on "kiki"; their "bouba" responses stay above chance. So the weak aggregate is not a general failure of sound symbolism there β€” it is specific to one of the two words.

CountrynCongruent95% CI"kiki"β†’spiky"bouba"β†’blob
Japan1,71672.9%71–75%78%67%
Ecuador22569.3%63–75%62%77%
Peru55067.1%63–71%64%70%
Bolivia31965.5%60–71%64%67%
Portugal33464.4%59–69%57%72%
France40463.1%58–68%60%66%
CΓ΄te d'Ivoire24361.7%55–68%58%65%
El Salvador21361.0%54–67%54%68%
India2,29758.0%56–60%53%63%
Philippines1,80157.7%55–60%56%59%
Spain1,49855.6%53–58%53%58%
Bangladesh46655.2%51–60%55%55%
Pakistan86054.7%51–58%52%58%
Iraq94252.2%49–55%49%56%
Egypt2,04151.2%49–53%41%61%
Algeria91244.0%41–47%40%47%

4. Age and gender barely matter β€” and the raw gaps are country composition

The demographic story is mostly a negative result, which is itself worth reporting. Demographics are self-reported and optional, so the cells below cover only part of the data (61% of responses carry an age band, 82% a gender; Other / Unknown / Prefer not to say is the platform's own third option):

AgenCongruent95% CI
18-293,96256.9%55–58%
30-392,28156.9%55–59%
40-492,07659.2%57–61%
50-642,36859.5%58–61%
65+1,48258.3%56–61%
GendernCongruent95% CI
Female7,59158.1%57–59%
Male4,95860.1%59–61%
Other / Unknown3,80360.7%59–62%

[image]

Taken at face value there are two small gaps: respondents aged 50+ are +2.1pp more congruent than the 18–29 group, and men +2.0pp more congruent than women. Both disappear once country is held constant β€” recomputing each gap within country and averaging (weighted by cell size, countries with β‰₯300 responses) gives -0.9pp for age (8 countries) and -0.1pp for gender (12 countries).

The raw gaps are an artifact of who answered from where: the low-congruence countries skew strongly female (65–69% of respondents reporting a gender), while high-congruence Japan and India skew male (41–43% female). Sample composition, not demography, produces the difference β€” a useful warning for anyone slicing this dataset by demographics without stratifying by country or language.

One robustness check in the same spirit: splitting respondents into quartiles by their userScore_audio reliability score moves congruence only between 55.9% and 60.4%, so the effect is not carried by inattentive listeners.


Dataset Structure

One row per spoken word, mirroring the text dataset's schema:

  • β€”question: the instruction shown to respondents
  • β€”spoken_word: "kiki" or "bouba"
  • β€”audio_context: the audio clip played to the respondent (β–Ά playable in the dataset viewer)
  • β€”option_1 / option_2: the two shape images (option_1 = blob, option_2 = spiky β€” same files and same order as the text dataset)
  • β€”option_1_selections / option_2_selections: total respondents choosing each option
  • β€”audio_license / audio_source: provenance of the audio clip
  • β€”detailed_results: respondent-level list with:
  • β€”selection ("option_1" / "option_2"), country, language, age, gender, occupation
  • β€”userScore: Rapidata's global respondent reliability score
  • β€”userScore_audio: reliability score on the audio validation dimension (new vs the text dataset)

Loading the Dataset

python
from datasets import load_dataset
import pandas as pd

ds = load_dataset("Rapidata/kiki-bouba-audio-20k")

# respondent-level responses for the "kiki" row
kiki = pd.DataFrame(ds["train"][0]["detailed_results"])
print(kiki.groupby("language")["selection"].value_counts(normalize=True))

Data Collection

Collected with the Rapidata API on a global annotator audience. Differences from the text run:

  • β€”The word is played as an audio context (same speaker for both words); the task cannot be answered before the clip finishes.
  • β€”Annotators who cannot listen at the moment get an explicit escape and are routed to other tasks, keeping non-listeners out of the data.
  • β€”An audio validation set (a "can you actually hear this?" check) was mixed into sessions β€” always for new annotators, sampled for established ones β€” and each response carries the respondent's userScore_audio.

Intended Use

  • β€”Sound-symbolism and cross-modal correspondence research
  • β€”Cross-linguistic / cross-cultural comparison with the paired text dataset β€” same shapes, same platform, different modality
  • β€”Human perception & semantic association research

Limitations

  • β€”Single speaker and accent (Kenyan Swahili speaker); accent effects are not controlled
  • β€”Single-word design without the paired anchor ("this is kiki, that is bouba"), matching the text dataset's phrasing
  • β€”Language/country composition is not balanced; Arabic speakers are overrepresented
  • β€”Country and language are heavily confounded (the below-chance and at-chance countries are all Arabic-speaking), so the country differences in finding 3 should not be read as national rather than linguistic effects
  • β€”Demographic slices are unreliable unless stratified by country β€” see finding 4

Citation

If you use this dataset, please refer back to this page, and consider leaving a like!