CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shredder-31 /contextualized-ST-Evidence Contextualized ST-Evidence A re-annotation of Salesforce/ST-Evidence-Instruct's gen_mask split. Same 19,902 entries, same objects, same frames, same temporal evidence. The only thing that changes is the spatial box on each frame. This is the video counterpart of shredder-31/contextualized-viscot, built with the same model, the same prompt design and the same union-with-the- original safety rule. Why ST-Evidence ships per-frame instance masks from GroundingDINO +… See the full description on the dataset page: https://huggingface.co/datasets/shredder-31/contextualized-ST-Evidence.imagevideo-text-to-text10K<n<100K0 likes212 downloads16d agoHugging Face02malaiwah /qfs-hf-jobs-campaign-evidence-20260909tabularn<1K0 likes76 downloads13d agoHugging Face03Loctran123 /vietnamese-evidence-retrieval-indexes-v2-r1 Vietnamese Evidence Retrieval Indexes Prebuilt exact dense and sparse indexes for Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v2-r1 at revision 2a18d35b6ea2e078db95c1aacdc2a28947268b4e. Rows: 63,699 Source embedding shards: 13 Dense: FAISS IndexFlatIP, 1024 dimensions Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75) BM25 content: title repeated 2 times + chunk text Dense input: title + text Dense rows: deduplicated by content hash Vietnamese tokenization: Unicode word… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-retrieval-indexes-v2-r1.tabularn<1K0 likes68 downloads1mo agoHugging Face04Loctran123 /vietnamese-evidence-corpus-chunked-e5-v3 Vietnamese Evidence Corpus - Chunked Chunked evidence corpus prepared for multilingual information retrieval, retrieval-augmented generation, and fact-checking experiments. Statistics Chunked with multilingual-E5 token budget Prefix-aware chunking using `passage: {title} ` Sentence-aware overlap to preserve local context Main fields chunk_id, doc_id, chunk_index token_start, token_end, token_count title, text, summary source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v3.tabulartext-retrieval10K<n<100K0 likes54 downloads1mo agoHugging Face05Loctran123 /vietnamese-evidence-retrieval-indexes Vietnamese Evidence Retrieval Indexes Prebuilt exact dense and sparse indexes for Loctran123/vietnamese-evidence-corpus-embeddings-e5-large at revision e928944361ca7d4c80f80d19bec52ebad55a4f7f. Rows: 52,605 Source embedding shards: 11 Dense: FAISS IndexFlatIP, 1024 dimensions Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75) BM25 content: title repeated 2 times + chunk text Vietnamese tokenization: Unicode word tokens, no stemming and no stopword removal row_id in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-retrieval-indexes.tabularn<1K0 likes53 downloads1mo agoHugging Face06dougdotcon /douvras-scientific-ci-evidence-graph Douvras Scientific CI Evidence Graph v0.1 Synthetic protocol dataset for linking a claim to its paper, repository, dataset, seed and reproduced metric. It contains 30 records from six toy paper instances (20 train, 5 validation and 5 frozen test), split by paper_id. The labels distinguish REPRODUCED, PARTIAL, FAILED and INCONCLUSIVE. Shortcuts and leakage fail closed. No real paper, code, dataset or result is included, and this release is not a reproduction benchmark. tabularn<1K0 likes52 downloads11d agoHugging Face07Yu-and-Ai /pythia-paths-evidence Pythia Paths Evidence A small, revision-pinned evidence bundle for examining model-training paths without converting a trend into authority. Companion read-only interface: Pythia Paths Static Space (mutable navigation; the evidence files below remain digest-pinned). Initial scope Model: EleutherAI/pythia-70m-deduped Run: the default public run only Context coverage: all 27 zero-shot reports in one pinned directory Detailed coverage: four post-outcome-selected… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/pythia-paths-evidence.tabularn<1K0 likes42 downloads2mo agoHugging Face08PsychiatryAgentBench25 /Health_Information_Seeking_under_Limited_Evidence Health Information Seeking under Limited Evidence (HISLE) HISLE is a clinically informed benchmark for evaluating LLM-based agents responding to incomplete mental-health information needs. File Records Contents matched_pairs_47.jsonl 47 Matched Chinese–English scenario pairs matched_variants_3290.jsonl 3,290 Query variants for the matched scenarios coverage_originals_24.jsonl 24 Coverage-expansion queries coverage_variants_840.jsonl 840 Query variants for… See the full description on the dataset page: https://huggingface.co/datasets/PsychiatryAgentBench25/Health_Information_Seeking_under_Limited_Evidence.tabular1K<n<10K0 likes40 downloads18d agoHugging Face09Loctran123 /vietnamese-evidence-corpus-chunked-e5-v2 Vietnamese Evidence Corpus - Chunked Chunked evidence corpus prepared for multilingual information retrieval, retrieval-augmented generation, and fact-checking experiments. Statistics 47,679 chunks from 13,572 source documents 38,603 Vietnamese chunks and 9,076 English chunks Maximum chunk length: 512 BGE-M3 tokenizer tokens Main fields chunk_id, doc_id, chunk_index token_start, token_end, token_count title, text, summary source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v2.tabulartext-retrieval10K<n<100K0 likes37 downloads1mo agoHugging Face10sanjay042211 /MetroPulse-TLC-Validation-Evidence MetroPulse TLC Validation Evidence This repository contains reproducibility and validation metadata for the MetroPulse NYC Data Reasoning Lab. The underlying real dataset is the NYC TLC Yellow Taxi Trip Records. Raw TLC trip data is not redistributed here. D1 validation run Run ID: RUN-20260913-TLC-SMOKE-001 Dataset: NYC TLC Yellow TaxiPartition: 2022-01Tier: Smoke validation Source evidence Input rows: 2,463,931 Source bytes: 38,139,949 SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/sanjay042211/MetroPulse-TLC-Validation-Evidence.tabularn<1K0 likes36 downloads10d agoHugging Face11Loctran123 /vietnamese-evidence-corpus-embeddings-e5-large-v2tabularn<1K0 likes34 downloads1mo agoHugging Face12AhaSignals /growth-evidence Growth evidence: revenue, cash and guidance AhaSignals, v1.0.0, published September 22, 2026. Canonical study: https://ahasignals.com/research/growth-evidence/ Coverage and clocks This retrospective research package combines original financial inputs for 32 selected companies with 44 reviewed revenue guidance ranges for META, AMZN, NVDA and AMD. At the frozen September 8, 2026 UTC cutoff, 40 ranges have comparable outcomes and four remain pending. Uncovered… See the full description on the dataset page: https://huggingface.co/datasets/AhaSignals/growth-evidence.tabularn<1K0 likes29 downloads2d agoHugging Face133CTeam /action-evidence-vla-phase-state-cachetabularn<1K0 likes25 downloads10d agoHugging Face14Loctran123 /vietnamese-evidence-corpus-chunked Vietnamese Evidence Corpus - Chunked Chunked evidence corpus prepared for multilingual information retrieval, retrieval-augmented generation, and fact-checking experiments. Statistics 47,679 chunks from 13,572 source documents 38,603 Vietnamese chunks and 9,076 English chunks Maximum chunk length: 512 BGE-M3 tokenizer tokens Main fields chunk_id, doc_id, chunk_index token_start, token_end, token_count title, text, summary source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked.tabulartext-retrieval10K<n<100K0 likes23 downloads1mo agoHugging Face15jadhavmanasi70 /adaption-indian-witness-evidence-struct This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-indian_witness_evidence_struct This dataset contains multilingual witness observations from Indian transit hubs and public areas, presented as prompts in various local languages including Hindi, Marathi, Kannada, Tamil, Bengali, Assamese, and Malayalam. The corresponding completions provide structured investigative evidence with standardized fields for demographics, location… See the full description on the dataset page: https://huggingface.co/datasets/jadhavmanasi70/adaption-indian-witness-evidence-struct.tabular10K<n<100K0 likes10 downloads3mo agoHugging Face16Gamestatue /adaption-historical-irrigation-evidence This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. adaption-historical_irrigation_evidence This dataset contains text completions detailing the historical presence, absence, or development of artificial irrigation and water management systems across various global regions and time periods. Each entry provides specific archaeological or textual evidence, often accompanied by academic citations, to support claims about… See the full description on the dataset page: https://huggingface.co/datasets/Gamestatue/adaption-historical-irrigation-evidence.tabular1K<n<10K0 likes7 downloads1mo agoHugging Face17Gamestatue /adaption-historical-irrigation-evidence-v1 This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. adaption-historical_irrigation_evidence This dataset contains text completions detailing the historical presence, absence, or development of artificial irrigation and water management systems across various global regions and time periods. Each entry provides specific archaeological or textual evidence, often accompanied by academic citations, to support claims about… See the full description on the dataset page: https://huggingface.co/datasets/Gamestatue/adaption-historical-irrigation-evidence-v1.tabular1K<n<10K0 likes7 downloads1mo agoHugging Face18sixstringzen /hemmingway-1-omlx-quantization-evidence-v2 Hemmingway-1 Quantization Evidence v2 This package records two local evidence lanes for the Hemmingway-1 oQ4e build: teacher-forced numerical fidelity against a BF16 reference, and controlled runtime telemetry on Apple Silicon. It complements the frozen blind-preference study in Hemmingway-1 oMLX Quantization Benchmark v1. This dataset is sixstringzen/hemmingway-1-omlx-quantization-evidence-v2. The quality dataset remains unchanged because blind preference, distribution fidelity… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-evidence-v2.tabulartext-generationn<1K0 likes10h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.