datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
blk-text-corpus
Verified Pa'O (Blk) Text Corpus - Pa'O Digital Hub
Dataset Summary
This is the official, verified parallel dataset for the Pa'O language (ISO 639-3: blk) and Burmese (Myanmar) translations, published by Pa'O Digital Hub.
The corpus is systematically collected, reviewed, standardized, and verified through the established linguistic and editorial workflow of Pa'O Digital Hub. The Pa'O sentences are based on authentic language usage by Pa'O native speakers and are… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/blk-text-corpus.pao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset
An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub.
The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language.
The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.pao-audio-dataset
🎙️ Pa'O Audio Dataset
ပအိုဝ်ႏ အငေါဝ်း အဆင်ႏဗာႏ ရွမ်ခြွဉ်းဗူႏ
📌 Project Summary
The Pa'O Audio Dataset is an open-source initiative created to facilitate the development of speech technologies and Artificial Intelligence tools for the Pa'O language (ပအိုဝ်ႏဘာႏသာႏငေါဝ်းငွါ).
Pa'O is primarily spoken in Shan State and other regions of Myanmar. As a low-resource language in the AI landscape, this dataset provides audio recordings and corresponding… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-audio-dataset.ace_attorney
Dataset Card for "ace_attorney"
More Information needed
sa-data
Storia dell'Arte Dataset (SA-Data)
📌 Descrizione del Dataset
Il dataset SA-Data è una raccolta strutturata di articoli della rivista Storia dell'Arte (https://www.storiadellarterivista.it/) digitalizzati e arricchiti con metadati dettagliati e rappresentazioni semantiche. È stato creato per supportare la ricerca accademica e le applicazioni di elaborazione del linguaggio naturale.
🔍 Contenuto
Il dataset include:
1050 articoli pubblicati tra il… See the full description on the dataset page: https://huggingface.co/datasets/paolodegasperis/sa-data.pao-sentences-dataset
Pa'O Sentences Dataset is a text corpus for Pa'O language (ပအိုဝ်ႏ) containing structured line-by-line sentences designed for NLP, LLM pre-training, and machine translation.
📝 Pa'O Sentences Dataset (ပအိုဝ်ႏ လိက်လာႏငေါဝ်းရဲဉ်ႏ ရွမ်ခြွဉ်းဗူႏ)
📌 Project Summary (ထာꩻမာꩻခြပ်ရဲဉ်ႏ နပ်ထွားရဲပ်အအဲဉ်ႏ)
The Pa'O Sentences Dataset is an open-source textual corpus developed to support Natural Language Processing (NLP), Large Language Model (LLM) pre-training, Machine… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-sentences-dataset.blog-authorship-corpusnews_articles
Dataset Card for "news_articles"
More Information needed
classificationArtVision
README — ArtVision: Dataset per la valutazione delle competenze visivo-interpretative in dominio storico-artistico
Descrizione generale
Il dataset ArtVision è una raccolta di 250 task, organizzati in otto categorie, in cui immagini di repertori storico artisti realizzati tra il 1750 e il 1985, sono utilizzate come base per la costruzione di richieste a modelli multimodali. Il dataset permette di sviluppare un veloce test di valutazione di un modello multimodale… See the full description on the dataset page: https://huggingface.co/datasets/paolodegasperis/ArtVision.satsec-decomposition
SatSec Grounded Objective-Decomposition Dataset
Version 2.0 is a leakage-controlled replacement for the original dataset used in
A Controlled Candidate-Set Benchmark for Offline Satellite-Security Plan
Decomposition
(DOI 10.48550/arXiv.2607.26371).
It contains 24 authored full decompositions and 83 mechanically derived next-step rows
across 24 cases.
There are 82 train rows and 25 test rows; the six test cases never occur in training.
Important v2 correction
The… See the full description on the dataset page: https://huggingface.co/datasets/paolocmo/satsec-decomposition.selma-indexThis is a small dataset containing a master index and 9 .txt files from Swedish author Selma Lagerlöf.
KodCode-V1-Formattedmax-dataset
High-latitude Pacific Ocean Sediment Geochemistry and XRF Data for Geoscientific Foundation Models
This dataset is a following development after the dataset (Chao et al., 2022), which inculdes the geochemical records from the high-latitude sectors of Pacific Ocean.
Besides the published XRF spectra-target measurements (CaCO3 and TOC) pairs of data, we further upload the XRF spectra in that project but without alignments of the target measurements here.
In total, it has 59,828 XRF… See the full description on the dataset page: https://huggingface.co/datasets/paoyw/max-dataset.lab2_edan20_ngramsA small dataset containing unigrams, bigrams and trigrams from Selma.txt
medium-size-generated-tasks
LICENSE
This is a dataset generated with the help of WizardLM. Therefore, the terms of use are restricted to research/academic only.
What is this
This is a collection of .txt files with a prompt and the expected output. For instance:
#####PROMPT:
Question: Make sure the task is unique and adds value to the original list.
Thought:#####OUTPUT: I should check if the task is already in the list.
Action: Python REPL
Action Input:
if task not in tasks:
print("Task not… See the full description on the dataset page: https://huggingface.co/datasets/paolorechia/medium-size-generated-tasks.diffusion-generated-text
Diffusion-Generated Text Benchmark
17,565 cleaned responses from three diffusion language model families and 21 generation settings
This benchmark supports research on diffusion-generated language, machine-generated text detection, and robustness across model families and decoding configurations. It includes outputs from DiffusionGemma, LLaDA-8B-Instruct, and LLaDA2-mini with varied generation lengths and block sizes.
Benchmark composition
Generator… See the full description on the dataset page: https://huggingface.co/datasets/paoche11/diffusion-generated-text.pubmedqadatasetopen-perfectblend-formattedcrosslg-news-smrandom_dataset_gold_at_7scilayprivacyqa_new
Dataset Card for "privacyqa_new"
More Information needed
futurama_dialoguesbiomedical-datasetcrosslg-retain-benchmark-entgtg-og-sm-3-1crosslg-retain-benchmark-entgen-og-sm-3-1memes_exist2024crosslg-contaminated-benchmark-og-en-3
