CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /cqadupstack-unix CQADupstackUnixRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Web, Programming Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackUnixRetrieval"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-unix.texttext-retrieval10K<n<100K0 likes6.2k downloads1y agoHugging Face02laion /terminal_bench_2_tasktrove_dq_unix_step10_30b_a3b_20260730_014756 terminal_bench_2_tasktrove_dq_unix_step10_30b_a3b OpenCode agent trajectories from the TaskTrove DQ unix arm of a Qwen3-Coder-30B-A3B agentic RL sweep, exported from the complete Harbor rollout artifact set. Coverage Built from the full trace_jobs prefix of run rl-tasktrove-dq-sweep-30b-qwen3-coder-30-20260727-082204-e42f1d (12034 trial directories, 11937 of them scored). quantity value scored trials (result.json) 11937 rows published 11937 coverage… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_unix_step10_30b_a3b_20260730_014756.text10K<n<100K0 likes547 downloads2mo agoHugging Face03unixtcpip /hacking Hacking Text Corpus A research corpus of historical computer security writings, hacker zines, and hacktivist texts. Built for NLP, text generation, discourse analysis, and security research. all.txt — all files for training language models This is the all files for training language models. all.txt is a single concatenated file containing all of the corpus content (every Phrack article, the Phineas Fisher writeups, unfz.txt, and thc.md) in one plain-text document… See the full description on the dataset page: https://huggingface.co/datasets/unixtcpip/hacking.texttext-generation1K<n<10K2 likes137 downloads3mo agoHugging Face04harpomaxx /unix-commands Unix Commands Dataset Description The Unix Commands Dataset is a unique collection of real-world Unix command line examples, captured from various system prompts representing different user roles and responsibilities, such as system administrators, DevOps, network administrators, Docker administrators, regular users, and hackers. The dataset consists of Unix commands ranging from basic to advanced levels and from a wide array of categories, including file operations (ls… See the full description on the dataset page: https://huggingface.co/datasets/harpomaxx/unix-commands.texttext-generationn<1K7 likes61 downloads3y agoHugging Face05mlfoundations-dev /stackexchange-unix-sandboxestext10K<n<100K0 likes45 downloads1y agoHugging Face06Hyukkyu /beir-cqadupstack-unix CQADupstackUnixRetrieval — BEIR, unified schema A normalised copy of the dataset behind the mteb task CQADupstackUnixRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset in this collection. Source mteb/cqadupstack-unix @ 6c6430d3a6d3 (the revision pinned in mteb)… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-unix.texttext-retrieval10K<n<100K0 likes45 downloads17d agoHugging Face07GreenNode /cqadupstack-unix-vn How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackUnix-VN"]) evaluator = mteb.MTEB(task) model = mteb.get_model(YOUR_MODEL) evaluator.run(model) To learn more about how to run models on mteb task check out the GitHub repitory. Citation If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing as a… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/cqadupstack-unix-vn.texttext-retrieval10K<n<100K0 likes41 downloads1y agoHugging Face08DCAgent /stackexchange-unix-sandboxes_glm_4.7_traces_jupitertext10K<n<100K0 likes40 downloads6mo agoHugging Face09laion /terminal_bench_2_a1_stackexchange_unix_20260809_015737text1K<n<10K0 likes36 downloads2mo agoHugging Face10mlfoundations-dev /stackexchange-unix-sandboxes-traces-terminus-2text1K<n<10K1 likes34 downloads1y agoHugging Face11mteb /CQADupstack-Unix-PL CQADupstack-Unix-PL An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Stack Exchange Question Duplicate Pairs Dataset Task category t2t Domains Written, Web, Programming Reference https://huggingface.co/datasets/clarin-knext/cqadupstack-unix-pl How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstack-Unix-PL"]) evaluator = mteb.MTEB(task)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstack-Unix-PL.texttext-retrieval10K<n<100K0 likes29 downloads1y agoHugging Face12Unix0923 /genai4t_spx_generationtabular1K<n<10K0 likes29 downloads24d agoHugging Face13orgrctera /beir_cqadupstack_unix CQADupStack / Unix (BEIR) — Unix & Linux Q&A retrieval Dataset description CQADupStack is a benchmark for community question answering (cQA) built from Stack Exchange data. It was introduced by Hoogeveen, Verspoor, and Baldwin at ADCS 2015 to support research on duplicate questions: finding earlier posts that match or subsume a newly asked question, so users can reuse existing answers instead of opening redundant threads. The full CQADupStack release aggregates twelve… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/beir_cqadupstack_unix.texttext-retrieval1K<n<10K0 likes26 downloads6mo agoHugging Face14Hale-Sage /unixtabular100K<n<1M0 likes25 downloads1y agoHugging Face15DCAgent2 /terminal_bench_2_a1_stackexchange_unix_20260328_072240textn<1K0 likes24 downloads6mo agoHugging Face16dmrau /cqadupstack-unix Dataset Card for "cqadupstack-unix" More Information needed text10K<n<100K0 likes23 downloads3y agoHugging Face17clarin-knext /cqadupstack-unix-pl-qrelsPart of BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language. Link to arxiv: https://arxiv.org/pdf/2305.19840.pdf Contact: konrad.wojtasik@pwr.edu.pl tabular1K<n<10K0 likes18 downloads2y agoHugging Face18MCINext /cqadupstack-unix-fa Dataset Summary CQADupstack-unix-Fa is a Persian (Farsi) dataset designed for the Retrieval task, with a focus on duplicate question retrieval. It is a translated version of the "unix" (Unix & Linux Stack Exchange) subforum from the original English CQADupstack dataset, used in the BEIR benchmark, and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) under the BEIR-Fa collection. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval) Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-unix-fa.text10K<n<100K0 likes18 downloads1y agoHugging Face19DCAgent2 /Kimi-K2T-stackexchange-unix-sandboxes-maxeps-32k0 likes18 downloads8mo agoHugging Face20income /cqadupstack-unix-top-20-gen-queries NFCorpus: 20 generated queries (BEIR Benchmark) This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset. DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1 id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl). Questions generated: 20 Code used for generation: evaluate_anserini_docT5query_parallel.py Below contains the old dataset card for the BEIR benchmark. Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/cqadupstack-unix-top-20-gen-queries.texttext-retrieval10K<n<100K1 likes17 downloads4y agoHugging Face21yzhuang /metatree_UNIX_user_data Dataset Card for "metatree_UNIX_user_data" More Information needed tabular1K<n<10K0 likes17 downloads3y agoHugging Face22laion /swebench_verified_random_100_folders_a1_stackexchange_unix_20260819_162244text1K<n<10K0 likes17 downloads1mo agoHugging Face23dmrau /cqudubstack-unix Dataset Card for "cqudubstack-unix" More Information needed text10K<n<100K0 likes16 downloads3y agoHugging Face24zsysuyang /Uni-XAS-dataset1 likes16 downloads2mo agoHugging Face25laion /dev_set_v2_a1_stackexchange_unix_20260813_104843text1K<n<10K0 likes16 downloads1mo agoHugging Face26clarin-knext /cqadupstack-unix-plPart of BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language. Link to arxiv: https://arxiv.org/pdf/2305.19840.pdf Contact: konrad.wojtasik@pwr.edu.pl 0 likes13 downloads2y agoHugging Face27sourcegraph /test-dataset-context-aware-fim-autocomplete-oracle-unixcoder_cosine_simtext1K<n<10K0 likes13 downloads2y agoHugging Face28DCAgent2 /swebench_verified_random_100_folders_a1_stackexchange_unix_20260325_214708textn<1K0 likes13 downloads6mo agoHugging Face29DCAgent2 /dev_set_v2_a1_stackexchange_unix_20260325_213949textn<1K0 likes13 downloads6mo agoHugging Face30AnhMinhLe /testgen_unixcodertabular1K<n<10K0 likes12 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.