ars
Datasets
All datasets matching “ars”symile-m3
Dataset Card for Symile-M3
Symile-M3 is a multilingual dataset of (audio, image, text) samples. The dataset is specifically designed to test a model's ability to capture higher-order information between three distinct high-dimensional data types: by incorporating multiple languages, we construct a task where text and audio are both needed to predict the image, and where, importantly, neither text nor audio alone would suffice.
Paper: https://arxiv.org/abs/2411.01053
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/arsaporta/symile-m3.ARS
data/ — Directory Structure
All data is gitignored. This file documents what lives here and how it's produced.
paperreview_data/
Crawled ICLR + NeurIPS paper corpus (read-only source of truth).
paperreview_data/
{venue}/ # iclr, neurips
{year}/ # 2017–2026 (ICLR), 2021–2025 (NeurIPS)
papers.jsonl # paper metadata + reviews (official_reviews,
# meta_reviews… See the full description on the dataset page: https://huggingface.co/datasets/Jerry999/ARS.ritual-agent-configsIllusionBencharsma-knowledge-dbFinglish-To-Persian-Dataset-Large
Finglish to Persian Large Dataset
A massive-scale parallel corpus containing over 9.8 million sentence pairs for Finglish (Latin-script Persian) to Persian script transliteration. This dataset provides a robust foundation for training and fine-tuning seq2seq models, normalizing user-generated text, and enhancing Persian input methods.
What is Finglish?
Finglish (also known as Pinglish) is the practice of writing Persian using the Latin alphabet. Because there is… See the full description on the dataset page: https://huggingface.co/datasets/Arshia82sbn/Finglish-To-Persian-Dataset-Large.
