datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
una-fraza-al-diya
Una fraza al diya
Ladino language learning sentences prepared by Karen Sarhon of Sephardic Center of Istanbul. Each sentence has translations in Turkish, English, Spanish. Includes audio and image. 307 sentences in total.
Source: https://sefarad.com.tr/judeo-espanyolladino/frazadeldia/
Citation
If you use this dataset, please cite:
Preparing an Endangered Language for the Digital Age: The Case of Judeo-Spanish
Preparing an endangered language for the digital age: The… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/una-fraza-al-diya.paleo-hebrew-seals-unambiguous
PaleoHebrew-Seals Real Benchmark (Unambiguous Subset)
This repository hosts the real benchmark part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions from photographs.
Why this dataset is needed
Paleo-Hebrew seal inscriptions are difficult for standard OCR systems: the signs are sparse, shallow, frequently worn, and embedded in irregular seal impressions captured under uncontrolled lighting and viewpoint changes.… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-unambiguous.meissa-unalignments
Meissa Unalignments
During the Meissa-Qwen2.5 training process, I noticed that Qwen's censorship was somehow bound to its chat template. Thus, I created this dataset with one system prompt, hoping to make it more effective in uncensoring models.
The dataset consists of:
V3N0M/Jenna-50K-Alpaca-Uncensored
jondurbin/airoboros-3.2, category = unalignment
Orion-zhen/dpo-toxic-zh, prompt and chosen
NobodyExistsOnTheInternet/ToxicQAFinal
task1640_aqa1.0_answerable_unanswerable_question_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1640_aqa1.0_answerable_unanswerable_question_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1640_aqa1.0_answerable_unanswerable_question_classification.
