CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01geniacllm /ApolloCorpus-ja-askllm-v1 ApolloCorpus-ja-askllm-v1 データセット kunishou/ApolloCorpus-ja に対して、 Ask-LLM 手法でスコア付けしたデータセットです。 元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。 Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。 ### {data} ### Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of the world, and strictly… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/ApolloCorpus-ja-askllm-v1.tabular100K<n<1M1 likes142 downloads2y agoHugging Face02OpenOranje /ReOpus-ApolloBooks-EN-NL-1M ReOpus-ApolloBooks-1M A high-quality English-Dutch (EN-NL) parallel corpus containing 1 million sentence pairs, constructed through strategic sampling and neural retranslation. Overview ReOpus-ApolloBooks-1M is a parallel translation corpus designed for training and evaluating English-Dutch machine translation systems. The corpus combines carefully sampled data from OPUS with neural retranslation using Qwen models, augmented with the Apollo Books parallel corpus.… See the full description on the dataset page: https://huggingface.co/datasets/OpenOranje/ReOpus-ApolloBooks-EN-NL-1M.tabulartranslation1M<n<10M0 likes65 downloads11mo agoHugging Face03open-llm-leaderboard /rootxhacker__Apollo-70B-detailsgated Dataset Card for Evaluation run of rootxhacker/Apollo-70B Dataset automatically created during the evaluation run of model rootxhacker/Apollo-70B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rootxhacker__Apollo-70B-details.tabular10K<n<100K0 likes64 downloads2y agoHugging Face04Apollo-LMMs /TimeScope-v0tabular1K<n<10K1 likes52 downloads1y agoHugging Face05davanstrien /apollo-11-diarized Apollo 11 Mission Audio — Diarized Transcripts Machine-generated transcripts with speaker diarization and timestamps for 103 tapes (175 hours) of Apollo 11 mission audio from the Internet Archive's Apollo11Audio collection (NASA recordings, public domain). Generated in a single Hugging Face Job with OpenMOSS-Team/MOSS-Transcribe-Diarize (0.9B, Apache 2.0) — joint transcription + speaker attribution + timestamps in one generation pass per clip. Configs segments… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/apollo-11-diarized.tabular10K<n<100K3 likes52 downloads2mo agoHugging Face06UMCU /apollo_english_guidelines_translated_to_dutch_with_gpt4omini Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM GPT 4o mini Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabular10K<n<100K0 likes35 downloads2y agoHugging Face07UMCU /apollo_english_guidelines_translated_to_dutch_with_nllb200 Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabulartext-generation10K<n<100K0 likes25 downloads2y agoHugging Face08UMCU /apollo_english_books_translated_to_dutch_with_geminiflash15 Data description Translation of the English medical books that are part of the Apollo corpus, using the LLM Gemini Flash 1.5 Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabular100K<n<1M0 likes22 downloads2y agoHugging Face09UMCU /apollo_english_guidelines_translated_to_dutch_with_geminiflash1.5 Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM Gemini Flash 1.5 Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabular10K<n<100K0 likes20 downloads2y agoHugging Face10open-llm-leaderboard /rootxhacker__Apollo_v2-32B-detailsgated Dataset Card for Evaluation run of rootxhacker/Apollo_v2-32B Dataset automatically created during the evaluation run of model rootxhacker/Apollo_v2-32B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rootxhacker__Apollo_v2-32B-details.tabular10K<n<100K0 likes17 downloads2y agoHugging Face11open-llm-leaderboard /rootxhacker__apollo-7B-detailsgated Dataset Card for Evaluation run of rootxhacker/apollo-7B Dataset automatically created during the evaluation run of model rootxhacker/apollo-7B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rootxhacker__apollo-7B-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face12UMCU /apollo_english_guidelines_translated_to_dutch_with_marianmt Data description Apollo corpus, English guidelines translated to Dutch using MariaNMT. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabulartext-generation10K<n<100K0 likes11 downloads2y agoHugging Face13Apollon9 /finepdfs_eng_Latn_labeledtabular1M<n<10M0 likes5 downloads9mo agoHugging Face14Locutusque /ApolloRP-DPO-v2.0gatedtabular10K<n<100K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.