pkheria7/indian-legal-opposing-counsel-dataset
โ๏ธ Indian Legal Opposing Counsel Dataset A combined, preprocessed dataset of 26,326 examples for training an Indian legal opposing counsel AI model. Ready-to-use in ChatML format for SFT training. ๐ Dataset Stats Split Rows Size Train 25,009 65 MB Test 1,317 3.5 MB Total 26,326 69 MB ๐ฆ Sources Source Dataset Rows Content viber1/indian-law-dataset 24,607 Writs, PIL, civil procedure, constitutional law, IPCโฆ See the full description on the dataset page: https://huggingface.co/datasets/pkheria7/indian-legal-opposing-counsel-dataset.
โ๏ธ Indian Legal Opposing Counsel Dataset
A combined, preprocessed dataset of 26,326 examples for training an Indian legal opposing counsel AI model. Ready-to-use in ChatML format for SFT training.
๐ Dataset Stats
๐ฆ Sources
๐๏ธ Format
Each row has a messages column in ChatML conversational format (directly compatible with TRL SFTTrainer):
{
"messages": [
{"role": "system", "content": "You are an experienced opposing counsel specializing in the Indian Constitution..."},
{"role": "user", "content": "What is the difference between a petition and a plaint in Indian law?"},
{"role": "assistant", "content": "A petition is a formal request submitted to a court..."}
],
"source": "viber1/indian-law-dataset"
}โฌ๏ธ Download
Option 1: Python (recommended)
from datasets import load_dataset
ds = load_dataset("pkheria7/indian-legal-opposing-counsel-dataset")
print(ds)
# DatasetDict({
# train: Dataset(25009 rows),
# test: Dataset(1317 rows)
# })Option 2: Direct JSONL downloads
- ๐ฅ train.jsonl (65 MB โ 25,009 rows)
- ๐ฅ eval.jsonl (3.5 MB โ 1,317 rows)
- ๐ฅ all.jsonl (69 MB โ all 26,326 rows)
- ๐ฅ raw_qa_pairs.jsonl (15 MB โ just user/assistant, no system prompt)
Option 3: wget / curl
# Full dataset (all splits combined)
wget https://huggingface.co/datasets/pkheria7/indian-legal-opposing-counsel-dataset/resolve/main/data/all.jsonl
# Or just the training split
wget https://huggingface.co/datasets/pkheria7/indian-legal-opposing-counsel-dataset/resolve/main/data/train.jsonlOption 4: Git clone
git lfs install
git clone https://huggingface.co/datasets/pkheria7/indian-legal-opposing-counsel-dataset๐๏ธ Training
Use with the model repo: pkheria7/indian-legal-opposing-counsel
from datasets import load_dataset
from trl import SFTConfig, SFTTrainer
ds = load_dataset("pkheria7/indian-legal-opposing-counsel-dataset")
trainer = SFTTrainer(
model="Qwen/Qwen2.5-7B-Instruct",
train_dataset=ds["train"],
eval_dataset=ds["test"],
args=SFTConfig(
max_length=2048,
assistant_only_loss=True,
push_to_hub=True,
hub_model_id="your-username/your-model-name",
),
)
trainer.train()๐ License
Apache 2.0
