CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AXONVERTEX-AI-RESEARCH /IndicRxNorm-LexMap-15K IndicRxNorm-LexMap-15K Dataset Summary IndicRxNorm-LexMap-15K is a multilingual Indic medicine terminology instruction dataset for medicine-name understanding, RxNorm normalization, RxCUI entity linking, structured drug-field extraction, and safe non-prescriptive clinical terminology tasks. This Hugging Face repository contains two dataset configurations: Config File Role multilingual_rxnorm_normalization multilingual_rxnorm_normalization.jsonl Primary adapted… See the full description on the dataset page: https://huggingface.co/datasets/AXONVERTEX-AI-RESEARCH/IndicRxNorm-LexMap-15K.texttoken-classification10K<n<100K0 likes42 downloads5mo agoHugging Face02axondendriteplus /legal-rag-embedding-dataset Legal Embedding Dataset This dataset was created to finetune embedding models for generating domain-specific embeddings on Indian legal texts, specifically SEBI (Securities and Exchange Board of India) documents. Data SourcePublicly available SEBI PDF documents were parsed and processed. Data Preparation PDFs were parsed to extract raw text, Text was chunked into manageable segments. For each chunk, a question was generated using gpt-4o-mini. Each question is directly… See the full description on the dataset page: https://huggingface.co/datasets/axondendriteplus/legal-rag-embedding-dataset.texttext-generation1K<n<10K1 likes29 downloads1y agoHugging Face03axonlabsai /ranger-omni-behaviour Ranger Omni — behaviour dataset 172 examples of four behaviours: antiloop (87) — resolve system/user conflicts in one step, never deliberate in circles ("always formal" + "yo casual" -> "4"). thinklevel (59) — honour an explicit effort directive: [effort: low] answers immediately, [effort: high] gives real reasoning. terse (18) — code-first answers with no preamble. webdesign (8) — real HTML/CSS with restraint. Format: JSONL of {"bucket", "prompt", "response"}. texttext-generationn<1K0 likes24 downloads1mo agoHugging Face04axonlabsai /ranger-omni-identity Ranger Omni — identity dataset 160 examples that pin a model's identity: name (Ranger Omni), maker (Axon Labs), 7B dense, multimodal input (text/image/audio/video), speech output, 32k context. Buckets: direct (45), adversarial (45), multimodal (40), incidental (30). Format: JSONL of {"bucket", "prompt", "response"}. Generated and audited — zero other-lab mentions, zero instruction leaks, and no false rejection of modalities the model actually has. texttext-generationn<1K0 likes22 downloads1mo agoHugging Face05axonlabsai /ranger-omni-code Ranger Omni — code dataset 200 self-contained Python task/solution pairs: a one-paragraph spec (no test cases, no expected outputs given) plus a complete, correct, runnable solution. Nothing fenced, no prose — each response is pure runnable code with the entry function. Stdlib only, difficulty from trivial to non-trivial (recursion, generators, nested data, error handling). Format: JSONL of {"bucket", "prompt", "response"}. texttext-generationn<1K0 likes16 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.