datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
meta-coq-utilsmeta-coq-commoncoqa
Dataset Card for "coqa"
Dataset Summary
CoQA is a large-scale dataset for building Conversational Question Answering systems.
Our dataset contains 127k questions with answers, obtained from 8k conversations about text passages from seven diverse domains. The questions are conversational, and the answers are free-form text with their corresponding evidence highlighted in the passage.
Supported Tasks and Leaderboards
More Information Needed
Languages… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/coqa.coqa"""CoQA dataset.
This CoQA adds the "additional_answers" feature that's missing in the original
datasets version:
https://github.com/huggingface/datasets/blob/master/datasets/coqa/coqa.py
"""
_CITATION = """@misc{reddy2018coqa,
title={CoQA: A Conversational Question Answering Challenge},
author={Siva Reddy and Danqi Chen and Christopher D. Manning},
year={2018},
eprint={1808.07042},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
"""
_DESCRIPTION = """CoQA is a… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/coqa.coqa-gen2mccoqa_mcCoQA is a large-scale dataset for building Conversational Question Answering
systems. The goal of the CoQA challenge is to measure the ability of machines to
understand a text passage and answer a series of interconnected questions that
appear in a conversation.
NOTE: this is the reformulated multiple choice version of the CoQA task, with downsampling.
coqa_expanded\\nCoQA: A Conversational Question Answering Challengene-tts-coqui-multilingual
NE-TTS Coqui Multilingual
Multilingual TTS dataset for 15 North East Indian languages, formatted for Coqui-AI VITS multilingual training. Contains 61,943 clips / 83.5 hours at 22050Hz (SNR >= 20dB only).
Languages
ISO
Language
Clips
Hours
grt
Garo
24,772
29.6
ccp
Chakma
10,689
14.3
nag
Nagamese
9,688
14.5
lus
Mizo
8,554
14.5
nnp
Wancho
5,081
6.3
trp
Kokborok
1,237
1.7
clk
Idu Mishmi
602
0.7
mjw
Karbi
373
0.4
nre
Rengma
258
0.4
nri… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-coqui-multilingual.coqa
Dataset Card for coqa
This is a preprocessed version of coqa dataset for benchmarks in LM-Polygraph.
Dataset Details
Dataset Description
Curated by: https://huggingface.co/LM-Polygraph
License: https://github.com/IINemo/lm-polygraph/blob/main/LICENSE.md
Dataset Sources [optional]
Repository: https://github.com/IINemo/lm-polygraph
Uses
Direct Use
This dataset should be used for performing benchmarks on LM-polygraph.… See the full description on the dataset page: https://huggingface.co/datasets/LM-Polygraph/coqa.coqar-clarifications-audio
CoQAR Clarifications with synthetic context audio
These audio recordings are AI-generated speech, not recordings of human speakers.
OpenAI tts-1 narrated each exact story using voice alloy, speed 1,
and MP3 output. Long stories are synthesized in ordered parts and joined; see the
audio generation manifest for part boundaries and measured audio properties.
No questions, answers, rationales, or stored model prompts were narrated.
The original appended and inserted configurations… See the full description on the dataset page: https://huggingface.co/datasets/rvashurin/coqar-clarifications-audio.coqa-flatcoqa-storiesThis is a dataset containing just stories of the CoQA dataset with their respective ids. This can be used in the pretraining phase for the MLM tasks.
Defects4Jcoqacoq-facts-props-proofs-gen0-v1
Dataset Name: Coq Facts, Propositions and Proofs
Dataset Description
The CoqFactsPropsProofs dataset aims to enhance Large Language Models'
(LLMs) proficiency in interpreting and generating Coq code by
providing a comprehensive collection of over 10,000 Coq source
files. It encompasses a wide array of propositions, proofs, and
definitions, enriched with metadata including source references and
licensing information. This dataset is designed to facilitate the… See the full description on the dataset page: https://huggingface.co/datasets/florath/coq-facts-props-proofs-gen0-v1.CoQuIRCoQCat
Dataset Card for CoQCat
Dataset Summary
CoQCat is a dataset for Conversational Question Answering in Catalan. It is based on CoQA dataset.
CoQCat comprises 89,364 question-answer pairs, sourced from conversations related to 6,000 text passages from six different domains.
The questions and responses are designed to maintain a conversational tone.
The answers are presented in a free-form text format, with evidence highlighted from the passage.
For the development and test… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CoQCat.coqar-clarifications
CoQAR Clarifications
This dataset pairs 1,000 CoQAR development questions with their original stories and stories damaged by sentence deletion. Each of the resulting 2,000 inputs has five sampled model clarifications. Two configurations reuse the same generated additions and differ only in where those additions are placed.
Configuration
Rows in dev
Clarifications per row
Placement
appended
2,000
5
At the end of the input story
inserted
2,000
5
At the deleted passage… See the full description on the dataset page: https://huggingface.co/datasets/rvashurin/coqar-clarifications.coqgym_coq_projects_v2coqar-s-uncertainty
Fields
id - unique example ID.
split - original CoQAR split (train / dev).
conversation_id — source conversation ID.
turn_id - turn index inside the conversation.
story - source passage shared by all variants.
answer - gold answer from CoQAR.
answer_span_text - supporting answer span in the passage.
all_standalone_questions - human-written standalone rewrites of the original question.
variants.low_s - standalone low specification version.
variants.high_s - original question… See the full description on the dataset page: https://huggingface.co/datasets/zykov/coqar-s-uncertainty.Coq-Changelog-QA
Coq Changelog Q&A Dataset
Dataset Description
The Coq Changelog Q&A Dataset is an extension of the original Coq Changelog Dataset, transforming each changelog entry into two Question–Answer pairs via two distinct prompts. One focuses on a straightforward query about the change itself, while the other aims at the rationale or motivation behind it. By applying both prompts to every changelog entry, the dataset approximately doubles in size compared to the source.… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-Changelog-QA.coqa_abg-translatedCoq-HoTT
Coq-HoTT
Structured dataset of formalizations from the Coq-HoTT library (Homotopy Type Theory in Coq).
Source
Repository: https://github.com/HoTT/Coq-HoTT
Commit: b75eadc7cb2bc59dca415bf47662a9290f82dc5f
Files: 589
License: bsd-2-clause
Schema
Column
Type
Description
statement
string
Declaration signature/claim with the leading keyword removed (verbatim slice); the full declaration minus its proof
proof
string
Verbatim proof/body… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-HoTT.coqa-preprocessedCoq-Bedrock
Coq-Bedrock
A work-in-progress language and compiler for verified low-level programming targeting RISC-V.
Source
Repository: https://github.com/mit-plv/bedrock2
Commit: c25e0e99557a6d71673d81c09b4fb24c048a2811
Files: 290
License: mit
Schema
Column
Type
Description
statement
string
Declaration signature/claim with the leading keyword removed (verbatim slice); the full declaration minus its proof
proof
string
Verbatim proof/body, empty… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-Bedrock.Coq-CompCert
Coq-CompCert
A structured dataset of formalizations from CompCert, the verified C compiler.
Source
Repository: https://github.com/AbsInt/CompCert
Commit: 0ef26dad76446c803da02d7368eb4f9d074c1401
Files: 222
License: other
Schema
Column
Type
Description
statement
string
Declaration signature/claim with the leading keyword removed (verbatim slice); the full declaration minus its proof
proof
string
Verbatim proof/body, empty if the… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-CompCert.Coq-HoTT-QA
Coq-HoTT Q&A Dataset
Dataset Description
The Coq-HoTT Q&A Dataset is a conversational extension of the Coq-HoTT Dataset, derived directly from the Coq-HoTT GitHub repository (https://github.com/HoTT/Coq-HoTT). This dataset transforms Homotopy Type Theory (HoTT) content into structured Q&A pairs, bridging the gap between formal mathematics and conversational AI.
Each entry in the dataset represents a mathematical statement, such as a definition or theorem, converted into a… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-HoTT-QA.Coq-UniMath-QA
UniMath Q&A Dataset
Dataset Description
The UniMath Q&A Dataset is a conversational extension of the UniMath Dataset, derived from the UniMath formalization of mathematics (https://github.com/UniMath/UniMath). This dataset transforms Univalent Mathematics content into structured Q&A pairs, making formal mathematical content more accessible through natural language interactions.
Each entry represents a mathematical statement from UniMath (definition, theorem, lemma, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-UniMath-QA.coqgym_with_goals_partialcoqar-s-uncertainty-passage
