pal
Datasets
All datasets matching “pal”phonemizer-dicts
Phonemizer Dicts
Pre-generated IPA dictionaries for GPL-free text-to-phonemes lookup.
Files
en-us.tsv — 124K English (US) words, tab-separated word<TAB>IPA
Provenance
Generated by running espeak-ng over an English wordlist. The TSV is program output; espeak-ng source (GPL-3.0) is not redistributed here.
Regeneration
See scripts/generate-espeak-dict.py in the tts-rd-team repo.
Paladin_TCGA_CPTAC_omicsPaladin TCGA & CPTAC Spatial Omics Maps
Ready-to-use patch-level and slide-level spatial omics maps inferred by
Paladin from TCGA and CPTAC
whole-slide images. The released .Paladin.h5 files can be analyzed directly
without rerunning WSI inference.
The collection is populated in stages. Check Files and versions for the
cohorts currently available.
Spatial multi-omics example
The panels show the H&E WSI, a reference tumor mask, CNV burden, TP53 CNV,
DNA-methylation… See the full description on the dataset page: https://huggingface.co/datasets/zhihuanglab/Paladin_TCGA_CPTAC_omics.PalmDex
🤖 PalmDex: An Embodiment-Agnostic Tactile Manipulation Dataset
A growing, multimodal robotic manipulation dataset featuring synchronized dual-camera video, dual-hand tactile sensing, and hand pose tracking across diverse real-world environments. All demonstrations are collected via human teleoperation — without any specific robot embodiment — and annotated at the action-segment level with rich categorical labels.
[!NOTE]
This is a living dataset. New environments, tasks, and… See the full description on the dataset page: https://huggingface.co/datasets/Rimbot/PalmDex.BrowseComp-ZH
🧭 BrowseComp-ZH: Benchmarking the Web Browsing Ability of Large Language Models in Chinese
BrowseComp-ZH is the first high-difficulty benchmark specifically designed to evaluate the real-world web browsing and reasoning capabilities of large language models (LLMs) in the Chinese information ecosystem. Inspired by BrowseComp (Wei et al., 2025), BrowseComp-ZH targets the unique linguistic, structural, and retrieval challenges of the Chinese web, including fragmented platforms… See the full description on the dataset page: https://huggingface.co/datasets/PALIN2018/BrowseComp-ZH.persona_in_palpaloma
Dataset Card for Paloma
Evaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains—varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. Among 16 source curated in Paloma, we include two… See the full description on the dataset page: https://huggingface.co/datasets/allenai/paloma.
