CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01awwaawwa /BabelDOC-Assets BabelDOC-Assets Font and other resource files relied on by BabelDOC and pdf2zh BabelDOC is a PDF translation library, pdf2zh is a PDF translation tool. Fonts and Licenses Go Noto Universal: THE UNLICENSE Pal3love/Source-Han-TrueType: SIL OPEN FONT LICENSE Version 1.1 lxgw/LxgwWenKaiGB: OFL-1.1 License lxgw/LxgwWenkaiTC: OFL-1.1 License fontworks-fonts/Klee: OFL-1.1 License fonts-archive/MaruBuri: License Noto Serif/Noto Sans: SIL OPEN FONT LICENSE Version 1.1… See the full description on the dataset page: https://huggingface.co/datasets/awwaawwa/BabelDOC-Assets.1 likes35k downloads10mo agoHugging Face02Babelscape /wikineural Dataset Card for WikiNEuRal dataset Description Summary: In a nutshell, WikiNEuRal consists in a novel technique which builds upon a multilingual lexical knowledge base (i.e., BabelNet) and transformer-based architectures (i.e., BERT) to produce high-quality annotations for multilingual NER. It shows consistent improvements of up to 6 span-based F1-score points against state-of-the-art alternative data production methods on common benchmarks for NER. We used this… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/wikineural.texttoken-classification1M<n<10M37 likes4.1k downloads4y agoHugging Face03Babelscape /multinerd Dataset Card for MultiNERD dataset Description Summary: In a nutshell, MultiNERD is the first language-agnostic methodology for automatically creating multilingual, multi-genre and fine-grained annotations for Named Entity Recognition and Entity Disambiguation. Specifically, it can be seen an extension of the combination of two prior works from our research group that are WikiNEuRal, from which we took inspiration for the state-of-the-art silver-data creation methodology… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/multinerd.texttoken-classification1M<n<10M26 likes1.7k downloads3y agoHugging Face04xdgdgcfbncvnvbn /babeldoc-temp-pdfsdocumentn<1K0 likes1.1k downloads3mo agoHugging Face05Babelscape /rebel-datasetREBEL is a silver dataset created for the paper REBEL: Relation Extraction By End-to-end Language generationtexttext-retrieval1M<n<10M34 likes688 downloads3y agoHugging Face06Babelscape /SREDFMRelation Extraction (RE) is a task that identifies relationships between entities in a text, enabling the acquisition of relational facts and bridging the gap between natural language and structured knowledge. However, current RE models often rely on small datasets with low coverage of relation types, particularly when working with languages other than English. \In this paper, we address the above issue and provide two new resources that enable the training and evaluation of multilingual RE systems. First, we present SRED\textsuperscript{FM}, an automatically annotated dataset covering 18 languages, 400 relation types, 13 entity types, totaling more than 40 million triplet instances. Second, we propose RED\textsuperscript{FM}, a smaller, human-revised dataset for seven languages that allows for the evaluation of multilingual RE systems. To demonstrate the utility of these novel datasets, we experiment with the first end-to-end multilingual RE model, mREBEL, that extracts triplets, including entity types, in multiple languages. We release our resources and model checkpoints at \href{https://www.github.com/babelscape/rebel}{https://www.github.com/babelscape/rebel}.texttoken-classification10M<n<100M14 likes592 downloads3y agoHugging Face07Babelscape /ALERT Dataset Card for the ALERT Benchmark Description Paper Summary: When building Large Language Models (LLMs), it is paramount to bear safety in mind and protect them with guardrails. Indeed, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to harm to individuals or society. In response to this critical challenge, we introduce ALERT, a large-scale benchmark to assess the safety of LLMs through red… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/ALERT.texttext-generation10K<n<100K16 likes503 downloads2y agoHugging Face08felixludos /babel-briefings Babel Briefings News Headlines Dataset README Break Free from the Language Barrier Version: 1 - Date: 30 Oct 2023 Collected and Prepared by Felix Leeb (Max Planck Institute for Intelligent Systems, Tübingen, Germany) License: Babel Briefings Headlines Dataset © 2023 by Felix Leeb is licensed under CC BY-NC-SA 4.0 Check out our paper on arxiv. This dataset contains 4,719,199 news headlines across 30 different languages collected between 8 August 2020 and 29 November 2021. The… See the full description on the dataset page: https://huggingface.co/datasets/felixludos/babel-briefings.imagetext-classification1M<n<10M6 likes407 downloads2y agoHugging Face09Babelscape /REDFMRelation Extraction (RE) is a task that identifies relationships between entities in a text, enabling the acquisition of relational facts and bridging the gap between natural language and structured knowledge. However, current RE models often rely on small datasets with low coverage of relation types, particularly when working with languages other than English. \In this paper, we address the above issue and provide two new resources that enable the training and evaluation of multilingual RE systems. First, we present SRED\textsuperscript{FM}, an automatically annotated dataset covering 18 languages, 400 relation types, 13 entity types, totaling more than 40 million triplet instances. Second, we propose RED\textsuperscript{FM}, a smaller, human-revised dataset for seven languages that allows for the evaluation of multilingual RE systems. To demonstrate the utility of these novel datasets, we experiment with the first end-to-end multilingual RE model, mREBEL, that extracts triplets, including entity types, in multiple languages. We release our resources and model checkpoints at \href{https://www.github.com/babelscape/rebel}{https://www.github.com/babelscape/rebel}.texttoken-classification10K<n<100K8 likes294 downloads3y agoHugging Face10ZeyuLing /babel-official BABEL Official Local Layout This directory is a cleaned local mirror of the official BABEL v1.0 labels and the AMASS subsets required by those labels. Layout archives/: original downloaded archives, kept unchanged for provenance. labels/babel_v1.0_release/: official BABEL JSON splits. amass/: extracted AMASS motion parameter files. processed/manifests/*.jsonl: normalized records with resolved local AMASS paths, frame counts, fps, segment labels, and rewritten… See the full description on the dataset page: https://huggingface.co/datasets/ZeyuLing/babel-official.0 likes183 downloads3mo agoHugging Face11GreatBird /BabelRStext100K<n<1M0 likes113 downloads9mo agoHugging Face12PaulChimzy01 /IARPA_BABEL_OP3_306audio10K<n<100K0 likes88 downloads11mo agoHugging Face13lxxiao /272-dim-BABEL-stream 🚀 Dataset Usage To facilitate researchers, we provide the processed streaming 272-dim Motion Representation of BABEL dataset in this Hugging Face repo. NOTE: We process the original BABEL dataset to support training of streaming motion generation. e.g. If there is a motion sequence A, annotated as (A1, A2, A3, A4) in BABEL dataset, each subsequence has text description: (A1_t, A2_t, A3_t, A4_t). Then, our BABEL-stream is constructed as: seq1: (A1, A2) --- seq1_text:… See the full description on the dataset page: https://huggingface.co/datasets/lxxiao/272-dim-BABEL-stream.1 likes84 downloads1y agoHugging Face14lxxiao /272-dim-BABEL 🚀 Dataset Usage To facilitate researchers, we provide the processed 272-dim Motion Representation of BABEL dataset in this Hugging Face repo. Motions are resampled into 30 FPS. NOTE: t2m_babel_mean_std/ contains the joint mean and std of both HumanML3D and BABEL dataset for joint training of the proposed Causal TAE. ❗️❗️❗️ The processed data is solely for academic purposes. Make sure you read through the BABEL License. 📖 Paper & Project Page & Code Arxiv Paper… See the full description on the dataset page: https://huggingface.co/datasets/lxxiao/272-dim-BABEL.text10K<n<100K1 likes75 downloads1y agoHugging Face15Babelscape /LLM-Oasis_unfactual_text_generation Babelscape/LLM-Oasis_unfactual_text_generation Dataset Description LLM-Oasis_unfactual_text_generation is part of the LLM-Oasis suite and contains unfactual texts generated from a set of falsified claims extracted from a Wikipedia passage and its paraphrase. This dataset corresponds to the unfactual text generation step described in Section 3.4 of the LLM-Oasis paper. Please refer to our GitHub repository for more information on the overall data generation pipeline of… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/LLM-Oasis_unfactual_text_generation.text10K<n<100K7 likes72 downloads2y agoHugging Face16BabelTowerProject /OpenNCEE-Chinese-Essaytextn<1K0 likes68 downloads7mo agoHugging Face17Babelscape /ALERT_DPO Dataset Card for the ALERT DPO Dataset Description Paper Summary: When building Large Language Models (LLMs), it is paramount to bear safety in mind and protect them with guardrails. Indeed, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to harm to individuals or society. In response to this critical challenge, we introduce ALERT, a large-scale benchmark to assess the safety of LLMs through red… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/ALERT_DPO.texttext-generation10K<n<100K14 likes63 downloads2y agoHugging Face18Babelscape /cner Dataset Card for CNER dataset Description Summary: Concept and Named Entity Recognition (CNER) is a novel task that jointly handles the indentification and classification of concepts and named entities. Repository: https://github.com/Babelscape/cner Paper: CNER: Concept and Named Entity Recognition Point of Contact: {martinelli, molfese, tedeschi, navigli}@diag.uniroma1.it Dataset Structure The data fields are the same among all splits. tokens: a list… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/cner.texttoken-classification100K<n<1M5 likes60 downloads2y agoHugging Face19Babelscape /LLM-Oasis_paraphrase_generation Babelscape/LLM-Oasis_paraphrase_generation Dataset Description LLM-Oasis_paraphrase_generation is part of the LLM-Oasis suite and contains paraphrases generated from a set of claims extracted from a Wikipedia passage. This dataset supports the paraphrase generation step described in Section 3.3 of the LLM-Oasis paper. Please refer to our GitHub repository for more information on the overall data generation pipeline of LLM-Oasis. Features title: The title… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/LLM-Oasis_paraphrase_generation.text10K<n<100K6 likes55 downloads2y agoHugging Face20AISE-TUDelft /LLM-of-Babel-NL2text100K<n<1M1 likes52 downloads2y agoHugging Face21Babelscape /LLM-Oasis_claim_extraction Babelscape/LLM-Oasis_claim_extraction Dataset Description LLM-Oasis_claim_extraction is part of the LLM-Oasis suite and contains text-claim pairs extracted from Wikipedia pages. It provides the data used to train the claim extraction system described in Section 3.1 of the LLM-Oasis paper. Please refer to our GitHub repository for more information on the overall data generation pipeline of LLM-Oasis. Features title: The title of the Wikipedia page. text: A… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/LLM-Oasis_claim_extraction.text10K<n<100K6 likes45 downloads2y agoHugging Face22BabelSpeech /50hours_Malay_Real-world_Colloquial_Conversation_and_Monologue_Speech_Datasetgated BabelSpeech: 50 Hours of Real-World Colloquial Malay ASR Speech Data This dataset contains 50 hours of high-quality Malay colloquial ASR speech data, reflecting realistic code-switching between Malay and English, as commonly used in everyday communication in Malaysia. Overview Content: 50 hours of real-world colloquial Malay speech suitable for ASR fine-tuning and benchmarking. Metadata: Stored in a separate JSON file, including audio path, duration, text, confidence… See the full description on the dataset page: https://huggingface.co/datasets/BabelSpeech/50hours_Malay_Real-world_Colloquial_Conversation_and_Monologue_Speech_Dataset.audio1 likes45 downloads11mo agoHugging Face23NamYeongCho /AMASS_BABEL0 likes43 downloads1y agoHugging Face24Babelscape /PDDL2PRM PDDL2PRM: Planning-Based Step-Level Supervision for Process Reward Models PDDL2PRM is a large-scale dataset for training and evaluating Process Reward Models (PRMs) with fine-grained, step-level supervision derived from symbolic planning problems. Unlike many PRM datasets that rely on human annotation, LLM judges, or final-answer correctness, PDDL2PRM uses Planning Domain Definition Language (PDDL) problems to generate structured reasoning trajectories whose intermediate steps… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/PDDL2PRM.texttext-classification1M<n<10M0 likes42 downloads3mo agoHugging Face25Babelscape /LLM-Oasis_claim_falsification Babelscape/LLM-Oasis_claim_falsification Dataset Description LLM-Oasis_claim_falsification is part of the LLM-Oasis suite and contains the outcomes of the claim falsification process. This dataset provides pairs of factual and falsified claims from a given Wikipedia text as described in Section 3.2 of the LLM-Oasis paper. Please refer to our GitHub repository for more information on the overall data generation pipeline of LLM-Oasis. Features title: The… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/LLM-Oasis_claim_falsification.text10K<n<100K6 likes41 downloads2y agoHugging Face26umanlp /babeledits BabelEdits BabelEdits is a benchmark designed to evaluate cross-lingual knowledge editing (CKE) in Large Language Models (LLMs). It enables robust and effective evaluation across 60 languages by combining high-quality entity translations from BabelNet with marker-based translation. BabelEdits is also accompanied by a modular CKE method, BabelReFT, which supports multilingual edit propagation while preserving downstream model performance. Dataset Summary As LLMs… See the full description on the dataset page: https://huggingface.co/datasets/umanlp/babeledits.2 likes41 downloads1y agoHugging Face27Tribhuvand /babel-briefings Babel Briefings News Headlines Dataset README Break Free from the Language Barrier Version: 1 - Date: 30 Oct 2023 Collected and Prepared by Felix Leeb (Max Planck Institute for Intelligent Systems, Tübingen, Germany) License: Babel Briefings Headlines Dataset © 2023 by Felix Leeb is licensed under CC BY-NC-SA 4.0 Check out our paper on arxiv. This dataset contains 4,719,199 news headlines across 30 different languages collected between 8 August 2020 and 29 November 2021. The… See the full description on the dataset page: https://huggingface.co/datasets/Tribhuvand/babel-briefings.imagetext-classification1M<n<10M0 likes34 downloads8mo agoHugging Face28dmargutierrez /Babelscape-wikineural-joined Dataset Card for "Babelscape-wikineural-joined" This dataset is a merged version of wikineural More Information needed @inproceedings{tedeschi-etal-2021-wikineural-combined, title = "{W}iki{NE}u{R}al: {C}ombined Neural and Knowledge-based Silver Data Creation for Multilingual {NER}", author = "Tedeschi, Simone and Maiorca, Valentino and Campolungo, Niccol{\`o} and Cecconi, Francesco and Navigli, Roberto", booktitle = "Findings of the… See the full description on the dataset page: https://huggingface.co/datasets/dmargutierrez/Babelscape-wikineural-joined.texttoken-classification1M<n<10M1 likes33 downloads4y agoHugging Face29BabelTowerProject /OpenNCEE-ChineseA database of NCEE(a.k.a. Gaokao) Chinese problems. No commercial use, or you ignore the risk of legal (especially China Mainland). texttext-generation10K<n<100K0 likes33 downloads7mo agoHugging Face30Babelscape /story-summeval Dataset Card for Story-SummEval Dataset Description For a thorough description of the data creation please refer to the ACL 2024 paper: "FENICE: Factuality Evaluation of summarization based on NLI and Claim Extraction", Scirè et al. (2024). Summary This dataset contains summaries of stories from Gutenberg and Wikisource along with their factuality labels. Summaries are generated from several models provided by the paper "Echoes from Alexandria" by Scirè et al.… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/story-summeval.textn<1K8 likes32 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.