CoolFace
Datasetpublic

karanverma19/trilingual_fraud_consumer_protection_v2

This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. Punjab Fraud Intent Benchmark (Trilingual – Punjabi/Hindi/English) Core Thesis Most fraud datasets focus on obvious scams. But in real-world communication, the harder problem is different: messages that look almost identical but have completely different intent This dataset focuses on that missing middle: intent boundaries in Punjabi–Hindi–English… See the full description on the dataset page: https://huggingface.co/datasets/karanverma19/trilingual_fraud_consumer_protection_v2.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes40downloads
Dataset Card

banner

This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.

Punjab Fraud Intent Benchmark (Trilingual – Punjabi/Hindi/English)

Core Thesis

Most fraud datasets focus on obvious scams.

But in real-world communication, the harder problem is different:

messages that look almost identical but have completely different intent

This dataset focuses on that missing middle:

intent boundaries in Punjabi–Hindi–English conversations, where meaning is not obvious from the words alone.

There is currently very limited open-source data capturing intent-level fraud detection in Punjabi–Hindi code-mixed communication.


Why existing fraud detection systems fail

Most fraud detection systems are trained on datasets where scam and safe messages are clearly different.

This works for obvious scams.

But in real-world multilingual communication:

  • —scam and safe messages often use the same vocabulary
  • —intent depends on subtle phrasing differences
  • —code-mixed language hides signals across languages

As a result, models trained on standard datasets often:

  • —over-flag normal messages (false positives)
  • —miss subtle scams (false negatives)

This dataset is designed to expose that failure.


Why this dataset matters

I'm from Punjab, India, and I've seen how many people rely on informal agents for visa and job opportunities.

Many of these messages don’t look obviously fake; they often sound exactly like normal communication.

Fraud detection in Punjab, India, especially around cities like Jalandhar, is not as straightforward as it looks.

Many immigration, IELTS, LMIA, and overseas job conversations happen through informal channels, usually WhatsApp-style messages in Punjabi, Hindi, or a mix of both.

In real life, scam messages don’t look obviously fake.

They often sound very similar to normal consultant communication.

The real challenge is not spotting obvious scams, but telling apart messages that look almost identical but mean very different things.

Most datasets focus on clear scam signals or keywords.

That makes models good at pattern detection, but not at understanding intent.


Why this problem is hard

In real-world immigration and job-related communication in Punjab, scam messages are often designed to feel normal.

They use polite language, partial truth, and slow trust-building instead of obvious red flags.

Sometimes, the difference between scam and safe is just a few words.

This makes it difficult even for humans, especially in code-mixed Punjabi–Hindi–English conversations.

This dataset is built to capture that difficulty.


Overview

The focus of this dataset is simple:

Not just what a scam looks like, but how it differs from something that looks almost the same.

It includes contrastive pairs, edge cases, and realistic safe messages to reflect how conversations actually happen.

This same structure can be extended across domains and languages, making it possible to scale intent-boundary detection beyond this initial dataset.


Methodology

This dataset was designed to capture intent differences rather than surface patterns.

The process focused on:

  • —creating contrastive pairs (scam vs safe with similar wording)
  • —injecting intent signals (urgency, authority, guarantees)
  • —adding realistic safe messages to reduce false positives
  • —designing edge cases where intent is ambiguous
  • —using Punjabi–Hindi–English code-mixing to reflect real communication

The goal was to force models to learn meaning, not just keywords.


Example: Similar message, different intent

Scam "we can fast-track your visa if you pay today"

Safe "we can guide you through the visa process"

Very small difference, but completely different intent.


Hard cases

Some examples in this dataset are intentionally difficult, where even humans might hesitate:

  • —messages that sound urgent but are legitimate alerts
  • —polite consultant-style messages that hide payment pressure
  • —code-mixed messages where intent depends on context, not wording

These cases are included to test whether a model can avoid both false positives and false negatives.


Linguistic insight

One pattern that stood out while building this dataset:

In Punjabi–Hindi–English code-mixed conversations, scam messages often don’t sound aggressive.

They often start with respect and familiarity, and only later introduce pressure.

For example:

  • —respectful greeting → builds trust
  • —then subtle urgency → pushes action
  • —then payment ask → reveals intent

This makes the boundary between safe and scam harder to detect, because the difference is not in keywords, but in how the message evolves.


Dataset Overview

Size and Coverage

  • —159 total rows
  • —Punjabi, Hindi, and English
  • —Scam and safe labels

Quality Checks

  • —0 missing values
  • —0 duplicate texts
  • —23 reasoning tag types
  • —Context types: original, contrast, edge cases

Structure

Each row includes:

  • —text – message content
  • —label – scam or safe
  • —category / sub_type – visa, job, OTP, etc.
  • —language and code_mixed
  • —context_type – original / contrast / edge case
  • —reasoning_tag – underlying signal (urgency, authority, etc.)
  • —difficulty – relative complexity

Real-world grounding

This dataset is inspired by commonly observed fraud patterns in Punjab, especially around:

  • —Canada work permit scams
  • —LMIA job claims
  • —IELTS manipulation offers
  • —embassy or High Commission impersonation
  • —WhatsApp-based advance fee pressure
  • —OTP, passport, and document collection scams

It is also informed by public warnings from Punjab Police and Indian consumer protection forums about fake job letters and visa scams circulating through informal messaging channels.

These situations often affect migrant workers and families making high-stakes financial decisions, where small misunderstandings can lead to serious losses.


Key observation

One thing that stood out while building this dataset:

Improvement didn’t come from adding more data.

It came from making examples more comparable.

When scam and safe messages are very different, models learn patterns. When they are very similar, models are forced to understand intent.


Key Insight

The hardest fraud messages are not obviously malicious.

They are linguistically normal.

The challenge is not detecting “bad” messages, but distinguishing between normal-looking messages with different intent.

This is where most current AI systems fail.


What this dataset teaches

This dataset helps models:

  • —distinguish intent when messages look very similar
  • —handle code-mixed Punjabi–Hindi–English communication
  • —avoid over-flagging normal messages as scams

Use Cases

  • —Fraud detection systems
  • —LLM safety and alignment evaluation
  • —Multilingual NLP benchmarking
  • —Consumer protection research
  • —Evaluation of LLM robustness in multilingual, high-ambiguity scenarios
  • —Benchmark dataset for evaluating intent-level fraud detection in multilingual settings

How Adaptive Data was used

This dataset was iteratively improved using Adaptive Data.

Main changes:

  • —improved phrasing consistency
  • —better alignment between scam and safe examples
  • —clearer intent in edge cases

Results:

  • —Grade: E → A
  • —Quality score: 2.0 → 9.8 (~390% improvement)
  • —Percentile: 0.1 → 57.7

Adaptive Data Results

The biggest improvement came from structure, not scale.


Acknowledgment

This dataset was developed using the Adaptive Data platform by Adaption Labs as part of the Uncharted Data Challenge.


Limitations

  • —Synthetic dataset (not real chat logs)
  • —Focused mainly on immigration and job-related fraud
  • —May not cover all dialect variations
  • —Still evolving for more complex multi-step scenarios

Ethical note

All samples are synthetic and based on commonly observed fraud patterns.

No personal or sensitive data is included.


Final note

This dataset started from a simple observation:

In real life, scams don’t look like scams.

They look normal.

And that’s what makes them hard to detect.


Links

Kaggle version: https://www.kaggle.com/datasets/karanverma/punjab-fraud-intent-benchmark-trilingual


Citation

If you use this dataset, please cite:

bibtex
@dataset{verma2026punjab_fraud_intent,
  title     = {Punjab Fraud Intent Benchmark (Trilingual – Punjabi/Hindi/English)},
  author    = {Verma, Karan},
  year      = {2026},
  month     = {May},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/datasets/karanverma19/trilingual_fraud_consumer_protection_v2},
  license   = {MIT}
}