prover
Datasets
All datasets matching “prover”Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.ProverbEval
ProverbEval: Benchmark for Evaluating LLMs on Low-Resource Proverbs
This dataset accompanies the paper:"ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding"ArXiv:2411.05049v3
Dataset Summary
ProverbEval is a culturally grounded evaluation benchmark designed to assess the language understanding abilities of large language models (LLMs) in low-resource settings. It consists of tasks based on proverbs in five languages:
Amharic
Afaan… See the full description on the dataset page: https://huggingface.co/datasets/israel/ProverbEval.IMOProofBenchDeepSeek-Prover-V1
Evaluation Results |
Model & Dataset Downloads |
License |
Contact
Paper Link👁️
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
1. Introduction
Proof assistants like Lean have revolutionized mathematical proof verification, ensuring high accuracy and reliability. Although large language models (LLMs) show promise in… See the full description on the dataset page: https://huggingface.co/datasets/deepseek-ai/DeepSeek-Prover-V1.DeepSeek-ProverBench
1. Introduction
We introduce DeepSeek-Prover-V2, an open-source large language model designed for formal theorem proving in Lean 4, with initialization data collected through a recursive theorem proving pipeline powered by DeepSeek-V3. The cold-start training procedure begins by prompting DeepSeek-V3 to decompose complex problems into a series of subgoals. The proofs of resolved subgoals… See the full description on the dataset page: https://huggingface.co/datasets/deepseek-ai/DeepSeek-ProverBench.Kimina-Prover-Promptset
Kimina-Prover-Promptset
Kimina-Prover-Promptset is a curated subset of NuminaMath-LEAN, designed for reinforcement learning (RL) training of formal theorem provers in Lean 4.
Compared to the full dataset, this subset contains fewer problems but with higher difficulty.
NuminaMath-LEAN is filtered and preprocessed as follows to create this dataset:
Remove easy problems with a historical win rate above 0.5 to only keeep challenging statements in the dataset.
Generate variants of… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/Kimina-Prover-Promptset.
