datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Eedi-Misconceptions-Graph
Eedi Misconceptions Graph v1.0
A map of mathematical misconceptions and the curriculum constructs they appear in, released by Eedi under a CC BY 4.0 licence.
Eedi defines a misconception as a flawed conceptual structure, or a gap in conceptual understanding, that manifests as a systematic and predictable error pattern across problems involving the same mathematical concept. A construct is a small, specific element of mathematics — for example, "Order fractions with the same… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Eedi-Misconceptions-Graph.Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.General_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.clinical-quad-central-lab-change-assay-drift-biomarker-noise-endpoint-misclassification-v0.1Clinical Quad Central Lab Change Assay Drift Biomarker Noise Endpoint Misclassification v0.1
Each row is a lab monthly snapshot.
Core quad
Central lab changeAssay driftBiomarker noiseEndpoint misclassification
Target
label_primary_fail_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
ag_misclassificationsThis dataset contains a slice of 200 samples from the AG News dataset (test split).
The picked 200 samples are potential misclassifications of the original test data.
Approach
Fine-tune DistilBERT with 10k samples from the training data (out of 120k)
Do a forward pass with the model, storing the loss
Sort the samples based on the loss
This is a repository for demonstration purposes
arabic_egypt_english_world_facts
🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized)
The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.meps_speeches_with_translation.csv
