religious
religious-radio-corpus
Religious Radio Corpus
A large-scale corpus of transcribed English-language religious radio broadcasts, captured from live webstreams in the United States over a one-month window in July 2025. Fifteen-minute segments were recorded on a rolling schedule from 779 distinct streams, which together rebroadcast the signals of more than 2,000 AM and FM stations, yielding 715,688 recordings and over 60 million diarized lines of speech. Each recording was transcribed and speaker-diarized… See the full description on the dataset page: https://huggingface.co/datasets/pew-data-labs/religious-radio-corpus.religious-artwork-analysis-data
Data
Download from Kaggle (needs an API token from https://www.kaggle.com/settings):
pip install kaggle
python data/download.py
Expected layout after download:
data/artwork_metadata.csv 3,997 rows — filename, religion (1,000 each of
buddhism / christianity / hinduism; 997 islam),
sub_religion, artist, title, year, place,
source, source_id, source_url, image_url
data/images/ the… See the full description on the dataset page: https://huggingface.co/datasets/cvikl/religious-artwork-analysis-data.religious-texts-rawopenbrush-religious-art
OpenBrush Religious Art
Religious paintings from OpenBrush-75K — saints, biblical scenes, devotional works.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 6,119 you actually want.
Why this subset
A coherent visual genre: religious narrative painting from medieval through early modern. Heavy on Renaissance and Baroque eras. Common… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-religious-art.shamela_all_diacritized_fully
Dataset Card for "shamela_all_diacritized_fully"
More Information needed
Scam-Religious-Detection
Scam Religious Detection Dataset
This dataset is designed for detecting religious-based scams, particularly on social media platforms like Twitter. It contains a collection of tweets categorized into various classes to facilitate the training of machine learning models for scam detection.
Dataset Details
Dataset Name: Scam-Religious-Detection
Primary Language: Arabic (ar)
License: MIT
Total Size: ~4.89 GB
Dataset Structure
The dataset consists of several CSV… See the full description on the dataset page: https://huggingface.co/datasets/abdessamad-bourkibate/Scam-Religious-Detection.
