CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /cqadupstack-unix CQADupstackUnixRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Web, Programming Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackUnixRetrieval"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-unix.texttext-retrieval10K<n<100K0 likes6.2k downloads1y agoHugging Face02harpomaxx /unix-commands Unix Commands Dataset Description The Unix Commands Dataset is a unique collection of real-world Unix command line examples, captured from various system prompts representing different user roles and responsibilities, such as system administrators, DevOps, network administrators, Docker administrators, regular users, and hackers. The dataset consists of Unix commands ranging from basic to advanced levels and from a wide array of categories, including file operations (ls… See the full description on the dataset page: https://huggingface.co/datasets/harpomaxx/unix-commands.texttext-generationn<1K7 likes61 downloads3y agoHugging Face03MCINext /cqadupstack-unix-fa Dataset Summary CQADupstack-unix-Fa is a Persian (Farsi) dataset designed for the Retrieval task, with a focus on duplicate question retrieval. It is a translated version of the "unix" (Unix & Linux Stack Exchange) subforum from the original English CQADupstack dataset, used in the BEIR benchmark, and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) under the BEIR-Fa collection. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval) Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-unix-fa.text10K<n<100K0 likes18 downloads1y agoHugging Face04income /cqadupstack-unix-top-20-gen-queries NFCorpus: 20 generated queries (BEIR Benchmark) This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset. DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1 id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl). Questions generated: 20 Code used for generation: evaluate_anserini_docT5query_parallel.py Below contains the old dataset card for the BEIR benchmark. Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/cqadupstack-unix-top-20-gen-queries.texttext-retrieval10K<n<100K1 likes17 downloads4y agoHugging Face05unixdevil /social-media-posttext1K<n<10K0 likes11 downloads6mo agoHugging Face06MichaelVeser /unixlogfilestext10K<n<100K0 likes7 downloads3y agoHugging Face07deathbyknowledge /shellm-V3-simple-unixtabular10K<n<100K3 likes4 downloads1y agoHugging Face08vengeance141 /unix-dataset-smalltextn<1K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.