datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMLA-Datasets
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
1. Introduction
MMLA is the first comprehensive multimodal language analysis benchmark for evaluating foundation models. It has the following features:
Large Scale: 61K+ multimodal samples.
Various Sources: 9 datasets.
Three Modalities: text, video, and audio
Both Acting and Real-world Scenarios: films, TV series, YouTube, Vimeo, Bilibili, TED, improvised scripts, etc.
Six Core… See the full description on the dataset page: https://huggingface.co/datasets/THUIAR/MMLA-Datasets.chess_datasets
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/p11-p11/chess_datasets.chess_puzzle_training_datasets_lt-2400
Chess puzzle training datasets: rating below 2400
This is a filtered derivative of
pavelslab-nyu/chess_puzzle_training_datasets.
Every retained row satisfies the exact condition:
Rating < 2400
Rating is the Lichess puzzle rating, not the Elo of either player in the
source game. The original column names, column order, directory layout, and CSV
schemas are preserved. As in the upstream repository, Hugging Face discovers
all three CSVs as one default configuration with one train… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/chess_puzzle_training_datasets_lt-2400.all_RD_datasets
RD Dataset With References
This dataset contains Arabic terms and their definitions.The data was extracted and combined from the following sources:
https://huggingface.co/datasets/Basma2423/Arabic-Terminologies-and-Definitions
https://data.mendeley.com/datasets/gxr3j4tdk5/3
https://huggingface.co/datasets/MohamedRashad/arabic-roots
https://arai.ksaa.gov.sa/sharedTask2024/
Each entry consists of:
word
definition
safety_aligned_datasets
Safety Aligned Datasets
A high-fidelity adversarial corpus engineered for alignment research, refusal boundary modeling, and robustness evaluation of Small Language Models.
The Problem This Solves
Fine-tuning a Small Language Model to be safe is not the same as fine-tuning it to understand safety.
Most safety datasets give models clean refusal examples on obvious prompts — and those models fail the moment an adversary wraps a harmful request in a… See the full description on the dataset page: https://huggingface.co/datasets/vvsd-charan/safety_aligned_datasets.khasi-datasets
What is Khasi Language?
Location:
Primarily spoken in the northeastern Indian state of Meghalaya.
Also spoken in parts of Assam, Tripura, and Bangladesh.
Language Family:
Khasi is a member of the Austroasiatic language family.
Script:
Traditionally written using the Khasi script, which is a script created specifically for the Khasi language.
Culture and Identity:
The Khasi language is an integral part of the cultural identity of the… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/khasi-datasets.Confidential-Datasets
