CoolFace
Datasetpublic

pkheria7/indian-legal-opposing-counsel-dataset

โš–๏ธ Indian Legal Opposing Counsel Dataset A combined, preprocessed dataset of 26,326 examples for training an Indian legal opposing counsel AI model. Ready-to-use in ChatML format for SFT training. ๐Ÿ“Š Dataset Stats Split Rows Size Train 25,009 65 MB Test 1,317 3.5 MB Total 26,326 69 MB ๐Ÿ“ฆ Sources Source Dataset Rows Content viber1/indian-law-dataset 24,607 Writs, PIL, civil procedure, constitutional law, IPCโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/pkheria7/indian-legal-opposing-counsel-dataset.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes67downloads
Dataset Card

โš–๏ธ Indian Legal Opposing Counsel Dataset

A combined, preprocessed dataset of 26,326 examples for training an Indian legal opposing counsel AI model. Ready-to-use in ChatML format for SFT training.

๐Ÿ“Š Dataset Stats

SplitRowsSize
Train25,00965 MB
Test1,3173.5 MB
Total26,32669 MB

๐Ÿ“ฆ Sources

Source DatasetRowsContent
viber1/indian-law-dataset24,607Writs, PIL, civil procedure, constitutional law, IPC
nisaar/Lawyer_GPT_India150Landmark cases, IPC, contract law, constitutional principles
RMani1/indian-legal-dataset-indian-law1,569Indian statutes, acts, legal provisions

๐Ÿ—‚๏ธ Format

Each row has a messages column in ChatML conversational format (directly compatible with TRL SFTTrainer):

json
{
  "messages": [
    {"role": "system", "content": "You are an experienced opposing counsel specializing in the Indian Constitution..."},
    {"role": "user", "content": "What is the difference between a petition and a plaint in Indian law?"},
    {"role": "assistant", "content": "A petition is a formal request submitted to a court..."}
  ],
  "source": "viber1/indian-law-dataset"
}

โฌ‡๏ธ Download

Option 1: Python (recommended)

python
from datasets import load_dataset
ds = load_dataset("pkheria7/indian-legal-opposing-counsel-dataset")
print(ds)
# DatasetDict({
#     train: Dataset(25009 rows),
#     test: Dataset(1317 rows)
# })

Option 2: Direct JSONL downloads

  • โ€”๐Ÿ“ฅ train.jsonl (65 MB โ€” 25,009 rows)
  • โ€”๐Ÿ“ฅ eval.jsonl (3.5 MB โ€” 1,317 rows)
  • โ€”๐Ÿ“ฅ all.jsonl (69 MB โ€” all 26,326 rows)
  • โ€”๐Ÿ“ฅ raw_qa_pairs.jsonl (15 MB โ€” just user/assistant, no system prompt)

Option 3: wget / curl

bash
# Full dataset (all splits combined)
wget https://huggingface.co/datasets/pkheria7/indian-legal-opposing-counsel-dataset/resolve/main/data/all.jsonl

# Or just the training split
wget https://huggingface.co/datasets/pkheria7/indian-legal-opposing-counsel-dataset/resolve/main/data/train.jsonl

Option 4: Git clone

bash
git lfs install
git clone https://huggingface.co/datasets/pkheria7/indian-legal-opposing-counsel-dataset

๐Ÿ‹๏ธ Training

Use with the model repo: pkheria7/indian-legal-opposing-counsel

python
from datasets import load_dataset
from trl import SFTConfig, SFTTrainer

ds = load_dataset("pkheria7/indian-legal-opposing-counsel-dataset")

trainer = SFTTrainer(
    model="Qwen/Qwen2.5-7B-Instruct",
    train_dataset=ds["train"],
    eval_dataset=ds["test"],
    args=SFTConfig(
        max_length=2048,
        assistant_only_loss=True,
        push_to_hub=True,
        hub_model_id="your-username/your-model-name",
    ),
)
trainer.train()

๐Ÿ“„ License

Apache 2.0