CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01milashkaarshif /MoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials. Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset. Looking forward to see more models and synthetic datasets trained from this raw archive, good luck! Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.texttext-generation100K<n<1M38 likes493 downloads8mo agoHugging Face02recursal /reprocessed_singapore_national_speech_corpus Dataset Card for Reprocessed National Speech Corpus NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here. Dataset Details Dataset Description Dataset Description: The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.audiotext-generation1M<n<10M7 likes353 downloads2y agoHugging Face03jakobsnel /RAGTruth_Xtended Dataset Card for Dataset Name This dataset provides response token logits and hidden states, complementing the underlying RAGTruth dataset. It has been generated using https://github.com/jakobsnl/RAGTruth_Xtended. Dataset Details Dataset Description This dataset is built upon RAGTruth (github.com/ParticleMedia/RAGTruth), which consists of character-level annotation of different types of hallucination for responses to a given set of LLM tasks. Out of all models… See the full description on the dataset page: https://huggingface.co/datasets/jakobsnel/RAGTruth_Xtended.texttext-generation10K<n<100K0 likes109 downloads1y agoHugging Face04AlexWortega /physics-scenarios-raw physics-scenarios-raw Raw (un-tarred) JSONL version of the 2D rigid body physics dataset. Each scene is a separate file under <split>/<scenario_type>/scene_<id>.jsonl. Streaming-friendly for HF datasets and curriculum sampling. For the bandwidth-efficient packaged version, see physics-scenarios-packed. Scale (this snapshot) Train: 80,000 scenes Val: 10,000 scenes Test: 10,000 scenes Frames per scene: 200 Format: one .jsonl per scene (1 header + 200 frame lines) This… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/physics-scenarios-raw.texttext-generation100K<n<1M0 likes10 downloads5mo agoHugging Face05ouy-not-reversed /aiaa4051-path-planning-data AIAA4051 Grid Path Planning Generated Inputs This dataset repository hosts generated experiment input data for the GitHub project ouy-not-reversed/aiaa4051-path-planning. The GitHub repository contains code, documentation, small samples, and lightweight result summaries. Full generated inputs are hosted here because they are too large for regular GitHub commits. The archives restore files under data/generated/ when extracted at the root of the GitHub repository. Raw upstream data… See the full description on the dataset page: https://huggingface.co/datasets/ouy-not-reversed/aiaa4051-path-planning-data.texttext-generationn<1K0 likes2 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.