CoolFace
20 results

clm

yueluoshuangtian /clmm-benchmarkimage0 likes378 downloads1y agoHugging FaceCLMBR /mSCANtext100K<n<1M1 likes192 downloads3y agoHugging Facebiglam /clmet_3_1 Dataset Card for clmet_3_1 NOTES: Some of the annotations in the class and pos configs are not properly formed. These are indicated with warning messages when the dataset is loaded. In addition to the classes mentioned in the README for the dataset, there is an additional class in the class dataset called QUOT. As far as I can tell, this is used for tagging all quotation marks When the class and pos configs are loaded, the available class/pos tags are shown at the top… See the full description on the dataset page: https://huggingface.co/datasets/biglam/clmet_3_1.texttext-classificationn<1K0 likes83 downloads2mo agoHugging Facej14i /cl-macros Common Lisp Macro Transformations A fine-tuning dataset for training models to generate Common Lisp macros. Each example is a call-form, macro-definition, and expanded-form triple. Summary 4,267 examples from 120+ Common Lisp libraries Split: 2,985 train / 637 validation / 645 test Format: JSONL with instruction, input, output, category, technique, complexity, quality_score Mean quality score: 0.80 Sources: Let Over Lambda, On Lisp, Alexandria, Serapeum, Iterate… See the full description on the dataset page: https://huggingface.co/datasets/j14i/cl-macros.text-generation1K<n<10K0 likes78 downloads5mo agoHugging Facesimon-clmtd /romansh-grischun-morphological-corpus Romansh Grischun Morphological Corpus A morphologically annotated corpus of Rumantsch Grischun, the standardized written variety of Romansh. Dataset Summary This dataset provides gold-standard morphosyntactically annotated and lemmatized corpora for Rumantsch Grischun (standard written Romansh, ISO 639-3: roh), a national language of Switzerland. The primary annotation layer preserves the rich Xerox/Foma two-level morphological tags used by the finite-state… See the full description on the dataset page: https://huggingface.co/datasets/simon-clmtd/romansh-grischun-morphological-corpus.texttoken-classification1K<n<10K0 likes76 downloads10d agoHugging Facemhla /gpt1900-physics-clm GPT-1900 Physics CLM Data Physics-domain text for continued pretraining (causal language modeling) of GPT-1900. This dataset contains chunks of text from seminal pre-1905 physics works — Newton's Principia, Maxwell's Treatise on Electricity and Magnetism, Faraday's Experimental Researches, Boltzmann, Gibbs, Hertz, and many others. Used to specialize the base GPT-1900 model toward physics reasoning before instruction tuning and reinforcement learning. Stats Split… See the full description on the dataset page: https://huggingface.co/datasets/mhla/gpt1900-physics-clm.text100K<n<1M0 likes75 downloads6mo agoHugging Face