CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PORTULAN /extraglue     This is the dataset card for extraGLUE. You may be interested in some of the other datasets for Portuguese and in the models trained with them, namely Albertina (encoders) and Gervásio (decoders) families. ExtraGLUE ExtraGLUE is a Portuguese dataset obtained by the automatic translation of some of the tasks in the GLUE and SuperGLUE benchmarks. Two variants of Portuguese are considered, namely European Portuguese and American Portuguese. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/PORTULAN/extraglue.tabulartext-classification100K<n<1M7 likes703 downloads2y agoHugging Face02GoktugD /turkish-extractive-qa-1.5m Turkish Extractive QA 1.5M v2 Cevap metni ve başlangıç konumu doğrulanabilir Türkçe çıkarımsal soru-cevap kayıtları. Doğrulanmış boyut Train: 1,470,000 Validation: 15,000 Test: 15,000 Toplam: 1,500,000 Ana görev sütunları: id, context, question, answer, answer_start, question_type Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-extractive-qa-1.5m.tabularquestion-answering1M<n<10M0 likes190 downloads2mo agoHugging Face03malr07 /opc-sft-stage2-dense-extracted OpenCoder Dataset Dense Region Extracted This dataset is a post-processed version of the OpenCoder SFT Stage2 dataset (opc-sft-stage2). We use gpt-4o API to extract the information dense regions from each sample and logged them in the dense_snippets column.Detailed information about the data can be found in our paper. OpenCoder's sft-stage2 summary The original version of this dataset is used in OpenCoder's Stage 2 and consists of four parts: educational_instruct:… See the full description on the dataset page: https://huggingface.co/datasets/malr07/opc-sft-stage2-dense-extracted.tabulartext-generation100K<n<1M0 likes71 downloads6mo agoHugging Face04brozonoyer /sudoku-extreme-multi-solution sudoku-extreme-multi-solution A controlled multi-solution Sudoku dataset derived from sapientinc/sudoku-extreme by deleting clues from uniquely solvable puzzles, with exact, doubly verified solution counts stratified over N ∈ {1, 2, 3, 4, 6, 8, 16} and the complete solution set enumerated for every item. Built to study how architectures handle solution multiplicity (many valid answers requiring a global consistent choice) separately from serial deduction depth — e.g. for… See the full description on the dataset page: https://huggingface.co/datasets/brozonoyer/sudoku-extreme-multi-solution.tabularquestion-answering1M<n<10M0 likes57 downloads1d agoHugging Face05potelo /extracao_estruturada Extração estruturada em português Estados textuais e perguntas de decisão tipada para treinar um extrator no formato Jev. A fonte cobre timelines processuais, documentos de timeline, fragmentos OCR e questões de concurso. Os exemplos foram gerados por um único modelo professor, com filtros locais; este release não é uma avaliação humana. Splits Split Estados Perguntas train 263,256 1,943,964 validation 2,659 19,204 Validação: 0.9999% dos estados… See the full description on the dataset page: https://huggingface.co/datasets/potelo/extracao_estruturada.tabulartext-classification100K<n<1M0 likes46 downloads17h agoHugging Face06uy-rrodriguez /FrenchMedMCQA-extended FrenchMedMCQA-extended: A French Multiple-Choice Question Answering Corpus for Medical domain, that supports Comparative Analysis with Human responses Dataset Summary This dataset is based on FrenchMedMCQA, the first publicly available Multiple-Choice Question Answering (MCQA) dataset in French for medical domain. We have enriched the content with additional annotations including student response rates downloaded from MedShake.net (the original data source)… See the full description on the dataset page: https://huggingface.co/datasets/uy-rrodriguez/FrenchMedMCQA-extended.tabularquestion-answering1K<n<10K0 likes27 downloads3mo agoHugging Face07TwinDoc /math-qa-sample_ext-kogated 데이터 출처 AI-HUB 에서 다운로드 받은 숫자연산 기계독해 데이터 를 사용해서 만든 데이터입니다. 경제 > Train > json 파일을 DataFrame 형태로 변형하여 전처리 및 답변 생성을 하였습니다. Raw 데이터의 answer 정보를 참고하여 답변을 생성하였습니다. 답변 생성 시 gpt-4o 를 활용했습니다. 저작권에 의해 본 데이터는 외부 반출 및 타인의 acess 승낙은 불허합니다. 데이터 설명 본 데이터의 Type 은 '단서추출' 로만 구성되어 있습니다. 데이터 예시 ### context ### 서울시가 민속 대명절인 추석을 맞아 내달 1일부터 20일까지 상생상회(매장), 네이버(온라인), 롯데백화점(매장)과 함께 팔도특산물로 구성된 명절 직거래장터를 진행한다고 31일 밝혔다. 팔도특산물을 구매할 수 있는 지역상생 거점공간인 '상생상회' 매장에서는 상주, 제주 등 14개 시도의 117개 농가에서 생산한 총… See the full description on the dataset page: https://huggingface.co/datasets/TwinDoc/math-qa-sample_ext-ko.tabularquestion-answering1K<n<10K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.