datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ScienceQAThis is the ScientificQA dataset by Saikh et al (2022).
@article{10.1007/s00799-022-00329-y,
author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak},
title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles},
year = {2022},
journal = {Int. J. Digit. Libr.},
month = {sep}
}
ARPA-Armenian-Paraphrase-Corpus
Dataset Description
We provide sentential paraphrase detection train, test datasets as well as BERT-based models for the Armenian language.
Dataset Summary
The sentences in the dataset are taken from Hetq and Panarmenian news articles. To generate paraphrase for the sentences, we used back translation from Armenian to English. We repeated the step twice, after which the generated paraphrases were manually reviewed. Invalid sentences were filtered out, while the rest were… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/ARPA-Armenian-Paraphrase-Corpus.arm-summaryannotations_creators:
other
language_creators:
other
languages:
hy-AM
licenses:
unknown
multilinguality:
monolingual
pretty_name: arm-sum
size_categories:
unknown
source_datasets:
original
task_categories:
conditional-text-generation
task_ids:
summarization
arm-asm2020_BRFSS_Codebook_CDCwwii-naval-armament
WWII Naval Armament & Treaty Data
Structured tables extracted from Fleets of World War II: Design History and
Analysis (Nimble Books, ISBN 9781608881604). Three queryable tables plus the full
set of 35 extracted source tables.
table
rows
what
naval_guns
93
gun armament specs — caliber, shell weight, range, ceiling, fire control, ship class
torpedoes
42
torpedo specs — type, explosive weight, range/speed
treaty_tonnage
19
interwar naval-treaty tonnage allocations… See the full description on the dataset page: https://huggingface.co/datasets/wfzimmerman/wwii-naval-armament.en-tr-translation
EN-TR Translation Dataset
Dataset Overview
This dataset is based on a subset of the Helsinki-NLP/opus-books dataset, which includes copyright-free books aligned for translation purposes. The original dataset contains multilingual sentence alignments. Specifically, I extracted the English-Italian (EN-IT) portion of the dataset, translated the English sentences to Turkish, and created an English-Turkish (EN-TR) parallel corpus.
The dataset can be used for various natural… See the full description on the dataset page: https://huggingface.co/datasets/armantunga/en-tr-translation.armchat1
THIS DATASET IS ONLY MADE FOR THESE
ID name color
1. ball yellow
2. battery silver
3. wood wood
4. bowl white
ARMor
ARMor (Additional Resources for Microcontroller-ORiented generation)
The ARMor dataset (Additional Resources for Microcontroller-ORiented generation)
is an extensive high-quality text dataset consisting of theoretical principles alongside
practical applications, industry standards and authentic research publications relating
to Embedded Systems Programming studies.
Dataset Details
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo… See the full description on the dataset page: https://huggingface.co/datasets/PulkundwarP/ARMor.gsm8k-armenian
GSM8K Տվյալների շտեմարան (Հայերեն տարբերակ՝ թվային պատասխաններով)
ԿԱՐԵՎՈՐ ԾԱՆՈՒՑՈՒՄ. Սույն տվյալների հավաքածուն հանդիսանում է OpenAI-ի հեղինակած GSM8K (Grade School Math 8K) բնօրինակ շտեմարանի հայերեն թարգմանությունը։ Բոլոր հեղինակային իրավունքները և բովանդակության սեփականությունը պատկանում են OpenAI-ին:
Այս տարբերակը կազմվել է հայալեզու մոդելների արագ և արդյունավետ ստուգաչափման (benchmarking) նպատակով։ Տվյալները ներկայացված են հստակ կառուցվածքով, որտեղ յուրաքանչյուր հարցի դիմաց… See the full description on the dataset page: https://huggingface.co/datasets/ArmGPT/gsm8k-armenian.arm-asm-xsmallSolana_vulnerability_audit_datasetJigsawarmenian-speech-dataset
🎧 Armenian Speech Dataset
📘 Overview
The Armenian Speech Dataset is a high-quality speech audio dataset designed for building, training, and evaluating modern AI voice technologies. It provides structured audio data optimized for deep learning workflows in speech processing. The dataset includes 76 hours of audio data distributed across 558 files, delivered in MP3 and WAV formats, with a total size of 189 MB.
This carefully curated audio dataset ensures balanced and… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/armenian-speech-dataset.armSumml_data_test_detection_bank_transaction_frauds_unbalanced
ML Data Test Detection Bank Transaction Frauds Unbalanced
The project provides a quick and accessible dataset designed for learning and experimenting with machine learning algorithms, specifically in the context of detecting fraudulent bank transactions. It is intended for practicing and applying concepts such as Random Forest, Support Vector Machines (SVM), and Synthetic Minority Over-sampling Technique (SMOTE) to address unbalanced classification problems.
Note: This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/roberto-armas/ml_data_test_detection_bank_transaction_frauds_unbalanced.social-media-addiction-vs-relationship
Social Media Addiction vs Relationship Dataset
Dataset ini merupakan hasil survei yang mengukur sejauh mana adiksi media sosial memengaruhi kualitas hubungan romantis mahasiswa.
Kolom-kolom dalam dataset:
Age
Gender
Daily Usage (hours)
Social Media Addiction Score
Relationship Satisfaction
Sumber Asli:
https://www.kaggle.com/datasets/adilshamim8/social-media-addiction-vs-relationships
NONSCLC-RAS-ARM-SV
NONSCLC-RAS-ARM-SV
tags: RandomizedArm, SurvivabilityRate, LungCancer
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description: The dataset NONSCLC-RAS-ARM-SV contains records of patients who have been part of randomized clinical trials studying the survivability rates of non-small cell lung cancer (NSCLC) patients who received a specific treatment related to the RAS gene mutation. Each record includes information about the patient, the… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/NONSCLC-RAS-ARM-SV.ARMed
Dataset Card for ARMed
This is a benchmark dataset of 1678 multiple-choice questions focused on embedded systems programming. It was generated via prompting Deepseek-R1 and refined through human verification by embedded systems interns at New Leap Labs, K J Somaiya Institute of Technology.
Dataset Details
Dataset Description
ARMed is a domain-specific evaluation benchmark designed to test the embedded systems programming knowledge of SLMs and LLMs. The… See the full description on the dataset page: https://huggingface.co/datasets/PulkundwarP/ARMed.prompt_reverse_engineering_code_dataset_O3_arm_O3_advanced_custom_testdaily_dialog_armenian
DailyDialog Armenian (Seq2Seq Format)
This dataset is a translated version of the DailyDialog dataset, where all dialog lines have been translated into Armenian. It is structured in Seq2Seq format, which makes it suitable for training conversational AI models, machine translation models, or general-purpose sequence-to-sequence learning systems.
📚 Dataset Description
The original DailyDialog dataset consists of multi-turn dialogues on daily life topics. Each dialogue… See the full description on the dataset page: https://huggingface.co/datasets/EdUarD0110/daily_dialog_armenian.humanoid-arm-reaching-dataset-v1Dataset for arm reaching and manipulation control.
Description
Target coordinates mapped to arm movement actions.
Task Description
Helps humanoid robots reach and position their arms accurately for object interaction tasks.
piaulidadesarm_sent_pairs_130k📊 130,000+ Sentence Pairs Dataset 📚🧠
This dataset contains over 130,000 aligned sentence pairs, designed for training and evaluating natural language processing (NLP) models. Each pair consists of an input sentence and a corresponding response or continuation, making it ideal for tasks such as:
🧠 Sequence-to-sequence learning
💬 Dialogue systems and chatbot training
📈 Text generation and completion
The data is structured to support supervised learning, with clean and contextually… See the full description on the dataset page: https://huggingface.co/datasets/EdUarD0110/arm_sent_pairs_130k.dartprompt_reverse_engineering_code_dataset_O1_arm_O1armenian_poems_dataset
Armenian 320 Poems Dataset 🇦🇲📜
This dataset contains 320 Armenian poems, collected and formatted for use in natural language processing, literary analysis, or educational purposes.
📑 Dataset Structure
The dataset is stored in a CSV format with 2 columns:
վերնագիր (Title)
poem (Poem Text)
Երգ եղբայրության
…
Իմ Հայաստան
…
Column Descriptions:
վերնագիր: The title of the poem in Armenian.
poem: The full text of the poem. Line breaks within… See the full description on the dataset page: https://huggingface.co/datasets/EdUarD0110/armenian_poems_dataset.prompt_reverse_engineering_code_dataset_O2_arm_O2_issueformatted_dart_datasetprompt_reverse_engineering_code_dataset_O0_arm_O0_advanced_custom_test
