CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lukesjordan /worldbank-project-documents Dataset Card for World Bank Project Documents Dataset Summary This is a dataset of documents related to World Bank development projects in the period 1947-2020. The dataset includes the documents used to propose or describe projects when they are launched, and those in the review. The documents are indexed by the World Bank project ID, which can be used to obtain features from multiple publicly available tabular datasets. Supported Tasks and Leaderboards No… See the full description on the dataset page: https://huggingface.co/datasets/lukesjordan/worldbank-project-documents.texttable-to-text10K<n<100K5 likes206 downloads4y agoHugging Face02jeffmeloy /python_documentation_codeDataset created using https://github.com/jeffmeloy/py2dataset using the Python code from the following: https://github.com/ansible/ansible https://github.com/apache/airflow https://github.com/arogozhnikov/einops https://github.com/arviz-devs/arviz https://github.com/astropy/astropy https://github.com/biopython/biopython https://github.com/bjodah/chempy https://github.com/bokeh/bokehhttps://github.com/CalebBell/thermo https://github.com/camDavidsonPilon/lifelines https://github.com/coin-or/pulp… See the full description on the dataset page: https://huggingface.co/datasets/jeffmeloy/python_documentation_code.texttext-generation10K<n<100K1 likes117 downloads2y agoHugging Face03Kemsekov /Corrupted-russian-word-documents-text-datasetThis is synthetic text-corruption dataset, based on Russian open official/government documents text. This dataset is intended to be used to train LLM to perform text-recovery task. All the errors in text is made solely in Russian sentences, hence ignoring any English sentence. Texts contains complex formatting, which is common for documents. Each line contains json object that have array messages value, which consists of role-based conversation. Each first message is randomly chosen system… See the full description on the dataset page: https://huggingface.co/datasets/Kemsekov/Corrupted-russian-word-documents-text-dataset.texttext-generationn<1K2 likes63 downloads2y agoHugging Face04stindardlogic /document-summarization-dpo-100k Document Summarization DPO (100K) 100,000 DPO (Direct Preference Optimization) preference pairs for training models to summarize business and professional documents with precision, structure, and analytical depth. Motivation Document summarization is one of the highest-value enterprise AI applications — analysts, lawyers, product managers, and executives use AI to process reports, contracts, and research daily. Models commonly fail by: Losing quantitative data:… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/document-summarization-dpo-100k.texttext-generation100K<n<1M0 likes33 downloads2mo agoHugging Face05Teen-Different /grpo-oumi-synthetic-document-claims Dataset Card for GRPO Oumi ANLI Subset Dataset This dataset is a reformatted version of the oumi-ai/oumi-synthetic-document-claims dataset, specifically structured for use with the GRPO trainer. You can find more detailed information about the original dataset at the provided link. Link: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-document-claims Dataset Structure The dataset consists of a list of dictionaries, where each dictionary represents a… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-synthetic-document-claims.texttext-generation1K<n<10K0 likes25 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.