CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Mohith202 /lma-individual-project-corporatabular10K<n<100K0 likes383 downloads7d agoHugging Face02lbrenap1 /mining-legal-arguments-us-corporate-case-law Mining Legal Arguments in U.S. Corporate Case Law This dataset contains span-level functional labels and directed support relations for 42 U.S. federal tax opinions concerning corporate reorganizations under I.R.C. Section 368. The opinions range in citation year from 1935 to 1987. Two law students annotated the cases, and a law professor adjudicated the final case-level representations. Ten cases also include the two independent annotations used for inter-annotator agreement… See the full description on the dataset page: https://huggingface.co/datasets/lbrenap1/mining-legal-arguments-us-corporate-case-law.tabulartext-classification10K<n<100K0 likes90 downloads24d agoHugging Face03zomi-language-corpora /English-Zomi-OPUS_Tatoeba_v20230412 English–Zomi Parallel Corpus (1.78M) This dataset contains 1.78 million English–Zomi sentence pairs, created to support machine translation, linguistic research, and large‑scale language model training. It is fully open and permissively licensed for commercial and non‑commercial use. 🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/English-Zomi-OPUS_Tatoeba_v20230412.tabulartranslation1M<n<10M1 likes87 downloads5mo agoHugging Face04agentionai /quant-fidelity-corpora Quantization fidelity corpora Evaluation text for measuring how faithfully a quantized LLM reproduces its full-precision parent (KL divergence of next-token distributions, top-1 agreement, perplexity ratio), as used by Agention for the Signal and Qwen3.8-27B quantization campaigns. mixedweb-v1 mixedweb-v1.txt (800,789 chars, 301 documents, md5 51e0045e8cabf37922aa82766a25b7b4) is a seeded random slice of HuggingFaceFW/fineweb sample-10BT: general English web text… See the full description on the dataset page: https://huggingface.co/datasets/agentionai/quant-fidelity-corpora.tabularn<1K0 likes31 downloads4d agoHugging Face05khaihernlow /bitcoin-news-articles-text-corporaimage1K<n<10K0 likes18 downloads2y agoHugging Face06siberian-lang-lab /evenki-rus-parallel-corporatabular1K<n<10K0 likes13 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.