CoolFace
27 results

Ks

ksolovev /FineNews30 likes551k downloads6mo agoHugging Faceksolovev /FineNewsTestSampletext10M<n<100M0 likes14k downloads7mo agoHugging FaceKShivendu /dbpedia-entities-openai-1M1M OpenAI Embeddings -- 1536 dimensions Created: June 2023. Text used for Embedding: title (string) + text (string) Embedding Model: text-embedding-ada-002 First used for the pgvector vs VectorDB (Qdrant) benchmark: https://nirantk.com/writing/pgvector-vs-qdrant/ Citation @dataset{dbpedia-entities-openai-1M, doi = {10.57967/hf/6768}, url = {https://huggingface.co/datasets/KShivendu/dbpedia-entities-openai-1M}, author = {{Kumar Shivendu} and {Nirant Kasliwal}}, title =… See the full description on the dataset page: https://huggingface.co/datasets/KShivendu/dbpedia-entities-openai-1M.textfeature-extraction1M<n<10M26 likes3.8k downloads11mo agoHugging Faceks46 /urls URLs 74,918,894,107 deduplicated, validated URLs, sorted by SURT key and split into 2,334 range shards. As plain text the URLs are 5.8 TiB, averaging 84 characters each. Sorted by SURT key and delta-encoded they fit in 665.6 GiB — 9.54 bytes per URL, a 8.85× reduction. That is the whole point of the ordering: SURT puts URLs from the same site next to each other, DELTA_LENGTH_BYTE_ARRAY then stores only where each row differs from the one above it, and zstd compresses what is… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls.texttext-generation10B<n<100B1 likes3k downloads1mo agoHugging Faceks48 /url-atlas URL Atlas 257,548,097,528 URLs from 105 web corpora, each kept as its own separately-loadable config, plus the raw source dumps two of them were extracted from. 4.35 TiB across 52,244 files. This is the input side of a URL-compression corpus: every source reduced to its URL column and nothing else. It is deliberately not deduplicated or merged — sources are kept intact and overlapping so you can measure what each one contributes, pick the subset you want, and dedup on your own… See the full description on the dataset page: https://huggingface.co/datasets/ks48/url-atlas.texttext-retrieval100B<n<1T0 likes2.9k downloads1mo agoHugging Faceksmashhero /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.audioaudio-classification1M<n<10M0 likes2.6k downloads15d agoHugging Face

People

Projects