datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indiccare-triage-rx-v0.1
Version note
This is a v0.1 pilot release created for the Adaptive Data Challenge. The dataset demonstrates the full pipeline and includes accepted/rejected quality gates, but additional manual review and scaling are planned before a larger release.
IndicCare-Triage-Rx
Code Author: Krishnendu DasguptaUsecase: Adaptive Challenge
IndicCare-Triage-Rx is a multilingual Indic-language primary-care safety-triage research dataset. It focuses on red-flag detection, care-urgency… See the full description on the dataset page: https://huggingface.co/datasets/AXONVERTEX-AI-RESEARCH/indiccare-triage-rx-v0.1.IndicRxNorm-LexMap-15K
IndicRxNorm-LexMap-15K
Dataset Summary
IndicRxNorm-LexMap-15K is a multilingual Indic medicine terminology instruction dataset for medicine-name understanding, RxNorm normalization, RxCUI entity linking, structured drug-field extraction, and safe non-prescriptive clinical terminology tasks.
This Hugging Face repository contains two dataset configurations:
Config
File
Role
multilingual_rxnorm_normalization
multilingual_rxnorm_normalization.jsonl
Primary adapted… See the full description on the dataset page: https://huggingface.co/datasets/AXONVERTEX-AI-RESEARCH/IndicRxNorm-LexMap-15K.context-relevance-classifier-dataset
context-relevance-classifier-dataset
This dataset is designed to train or evaluate models on determining whether an answer to a question is grounded in a given context.
Each sample includes:
question: A question.
answer: A possible answer to the question.
context: A legal passage or reference document.
label:
1 → The answer is supported by the context.
0 → The answer is not supported by the context.
Dataset Source
This dataset is derived from:… See the full description on the dataset page: https://huggingface.co/datasets/axondendriteplus/context-relevance-classifier-dataset.legal-rag-embedding-dataset
Legal Embedding Dataset
This dataset was created to finetune embedding models for generating domain-specific embeddings on Indian legal texts, specifically SEBI (Securities and Exchange Board of India) documents.
Data SourcePublicly available SEBI PDF documents were parsed and processed.
Data Preparation
PDFs were parsed to extract raw text, Text was chunked into manageable segments.
For each chunk, a question was generated using gpt-4o-mini.
Each question is directly… See the full description on the dataset page: https://huggingface.co/datasets/axondendriteplus/legal-rag-embedding-dataset.legal-qna-datasetThis dataset is produced using axondendriteplus/legal-rag-embedding-dataset
Using the "question" & "context" from this dataset, generated "answer" for each question using gpt-4.1-nano
ranger-omni-behaviour
Ranger Omni — behaviour dataset
172 examples of four behaviours:
antiloop (87) — resolve system/user conflicts in one step, never deliberate
in circles ("always formal" + "yo casual" -> "4").
thinklevel (59) — honour an explicit effort directive: [effort: low] answers
immediately, [effort: high] gives real reasoning.
terse (18) — code-first answers with no preamble.
webdesign (8) — real HTML/CSS with restraint.
Format: JSONL of {"bucket", "prompt", "response"}.
IC38-Web-Aggregator-Questionsranger-omni-identity
Ranger Omni — identity dataset
160 examples that pin a model's identity: name (Ranger Omni), maker (Axon Labs),
7B dense, multimodal input (text/image/audio/video), speech output, 32k context.
Buckets: direct (45), adversarial (45), multimodal (40), incidental (30).
Format: JSONL of {"bucket", "prompt", "response"}.
Generated and audited — zero other-lab mentions, zero instruction leaks, and no
false rejection of modalities the model actually has.
ranger-omni-code
Ranger Omni — code dataset
200 self-contained Python task/solution pairs: a one-paragraph spec (no test
cases, no expected outputs given) plus a complete, correct, runnable solution.
Nothing fenced, no prose — each response is pure runnable code with the entry
function. Stdlib only, difficulty from trivial to non-trivial (recursion,
generators, nested data, error handling).
Format: JSONL of {"bucket", "prompt", "response"}.
IC38-Insurance-Marketing-Firm-Questionsaxonis
Axonis Entity Q&A Dataset
Structured question-and-answer pairs covering key facts about Axonis, a federated AI platform for regulated, distributed, and high-stakes enterprise environments. This dataset is maintained to support Answer Engine Optimization (AEO) and Generative Engine Optimization (GEO), helping AI models provide accurate information about Axonis when responding to user queries.
About Axonis
Axonis is a federated AI platform that moves AI models to where… See the full description on the dataset page: https://huggingface.co/datasets/soulcraftagency/axonis.IC38-Corporate-Agent-QuestionsIC38-Insurance-Agent-QuestionsIC38_Study_Material
