datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mathematical_scientific_notation
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/mathematical_scientific_notation.LLM-Hallucination-Detection-complex-mathematics
AIME Hallucination Detection Dataset
This dataset is created for detecting hallucinations in Large Language Models (LLMs), particularly focusing on complex mathematical problems. It can be used for tasks like model evaluation, fine-tuning, and research.
Dataset Details
Name: AIME Hallucination Detection Dataset
Format: CSV
Size: (add size, e.g., 10MB)
Files Included:
AIME-hallucination-detection-dataset.csv: Contains the dataset.
Content Description… See the full description on the dataset page: https://huggingface.co/datasets/tourist800/LLM-Hallucination-Detection-complex-mathematics.repro-why-agentic-theorem-prover-works-a-statistical-provability-theory-of-mathematica-traces
Agent traces
Agent sessions published from a Trackio Logbook.
eng_latn_mathematical_scientific_notation
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/eng_latn_mathematical_scientific_notation.shamela_subset_wahhabiteshamela_subset_not_wahhabiteasia-owid-average-performance-of-15-year-olds-in-mathematics-reading-and-science
Average Performance Of 15 Year Olds In Mathematics Reading And Science | Asia (Our World in Data)
🌏 104 observations · 24 Asia countries · 2000–2022 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 104 observations of Average Performance Of 15 Year Olds In Mathematics Reading And Science data across 24 Asia countries, spanning 2000–2022.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-average-performance-of-15-year-olds-in-mathematics-reading-and-science.cqadupstack-mathematica-pl-qrelsPart of BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language.
Link to arxiv: https://arxiv.org/pdf/2305.19840.pdf
Contact: konrad.wojtasik@pwr.edu.pl
ScholarBench_MC_physics_mathematicsita_latn_mathematical_scientific_notation
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/ita_latn_mathematical_scientific_notation.europe-owid-average-performance-of-15-year-olds-in-mathematics-reading-and-science
Average Performance Of 15 Year Olds In Mathematics Reading And Science | Europe (Our World in Data)
🇪🇺 254 observations · 39 Europe countries · 2000–2022 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 254 observations of Average Performance Of 15 Year Olds In Mathematics Reading And Science data across 39 Europe countries, spanning 2000–2022.
About the source
Source: Our World in Data
Publisher: Our World in Data
License:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-average-performance-of-15-year-olds-in-mathematics-reading-and-science.africa-worldbank-female-share-of-graduates-from-science-technology-engineering-and-mathematics-s
Female share of graduates from Science, Technology, Engineering and Mathematics (STEM) programmes, tertiary (%) | Africa (World Bank — Gender Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-female-share-of-graduates-from-science-technology-engineering-and-mathematics-s.tur_latn_mathematical_scientific_notation
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/tur_latn_mathematical_scientific_notation.zho_hans_mathematical_scientific_notation
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/zho_hans_mathematical_scientific_notation.sept19-mathematicseng_latn_mathematical_scientific_notation
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/eng_latn_mathematical_scientific_notation.toaa_mathematical_reasoningmathematical_economics_fineweb_phi3.5_unsup
