CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01enelpol /rag-mini-bioasq-with-metadataThis dataset is an extension of the rag-mini-bioasq dataset. Its difference resides in the text-corpus part of the aforementioned set where the metadata was added for each passage. Metadata contains six separate categories, each in a dedicated column: Year of the publication (publish_year) Type of the publication (publish_type) Country of the publication - often correlated with the homeland of the authors (country) Number of pages (no_pages) Authors (authors) Keywords (keywords) tabularquestion-answering10K<n<100K2 likes158 downloads2y agoHugging Face02MapleBi /MetaRAG_Cross-Issue_OSSQA MetaRAG Cross-Issue OSSQA Dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA MetaRAG Cross-Issue OSSQA is an English open-source software issue question-answering and retrieval benchmark. Each example asks a question grounded in one GitHub issue and requires evidence from a related issue. The data contains explicit cross-issue references and a three-document silver evidence path. Dataset configurations Configuration Splits Rows… See the full description on the dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA.tabularquestion-answering10K<n<100K0 likes45 downloads1mo agoHugging Face03LLMTeamAkiyama /cleand_meta-math_MetaMathQA元データ: https://huggingface.co/datasets/meta-math/MetaMathQA データ件数: 394,369 平均トークン数: 233 最大トークン数: 2,874 合計トークン数: 91,798,611 ファイル形式: JSONL ファイルサイズ: 297.9 MB =================== 以下、加工内容をclaudeでまとめ。 MetaMathQAデータセット加工内容 データ読み込み・準備 HuggingFace Datasetsからmeta-math/MetaMathQAの訓練データ(395,000件)を読み込み DeepSeek-R1-Distill-Qwen-32Bトークナイザーを使用してトークン数を計算 データ構造の理解・分析 全てのresponseが"The answer is:"で終わる統一フォーマットであることを確認 original_questionとresponseを結合してトークン数計算用テキストを作成… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_meta-math_MetaMathQA.tabularquestion-answering100K<n<1M0 likes30 downloads1y agoHugging Face04violetxi /single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5.4_gepa-n32 Single-turn eval — violetxi/meta_feedback_qwen3-4b_step2_gpt-5.4_gepa Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N. Eval results (n_samples_per_example = 32) Overall metric value n_examples 1006 mean@32 0.1796 best@32 0.3588 worst@32 0.0477 pass_rate 0.3588… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5.4_gepa-n32.tabularquestion-answering1K<n<10K0 likes26 downloads5mo agoHugging Face05gdgc-metacong /apigen-inferred apigen-inferred A verified, GPT-5.5-distilled subset of the argilla/apigen-function-calling dataset (109k rows in the upstream), with every golden tool-call argument labelled as literal or dependency-derived to enable a clean function-calling benchmark. Pipeline Filter the upstream to rows where every called API actually works (replay each tool call against the real implementation — distilabel Python functions or live RapidAPI / cached responses) → 45,984 rows. Distill… See the full description on the dataset page: https://huggingface.co/datasets/gdgc-metacong/apigen-inferred.tabulartext-generation10K<n<100K1 likes19 downloads5mo agoHugging Face06potsu-potsu /mini-bioasq-with-metadataThis dataset is an extension of the rag-mini-bioasq dataset. Its difference resides in the text-corpus part of the aforementioned set where the metadata was added for each passage. Metadata contains six separate categories, each in a dedicated column: Year of the publication (publish_year) Type of the publication (publish_type) Country of the publication - often correlated with the homeland of the authors (country) Number of pages (no_pages) Authors (authors) Keywords (keywords) tabularquestion-answering10K<n<100K0 likes14 downloads1y agoHugging Face07violetxi /single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5-nano_gepa-n32 Single-turn eval — violetxi/meta_feedback_qwen3-4b_step2_gpt-5-nano_gepa Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N. Eval results (n_samples_per_example = 32) Overall metric value n_examples 1006 mean@32 0.1804 best@32 0.3569 worst@32 0.0398 pass_rate… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5-nano_gepa-n32.tabularquestion-answering1K<n<10K0 likes10 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.