datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
m2mdatasetm2mcent-mcp-schemas
🌐 M2MCent Agentic Services - MCP Schemas Dataset
🚀 Empowering Autonomous AI on Base L2
This dataset contains the JSON schemas for 1,005 microservices natively available on the M2MCent Network via the x402 V2 Protocol (EIP-3009).
It is specifically designed for instruction-tuning LLMs (like Llama-3, Mistral, Qwen) so they can autonomously discover, negotiate, and consume monetized API endpoints using gasless cryptocurrency settlements on the Base L2 network.… See the full description on the dataset page: https://huggingface.co/datasets/evozim/m2mcent-mcp-schemas.m2m3_fine_tuning_ref_ptrn_cmbert_io
m2m3_fine_tuning_ref_ptrn_cmbert_io
Introduction
This dataset was used to fine-tuned HueyNemud/das22-10-camembert_pretrained for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : ground-truth
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ref_ptrn_cmbert_io.m2m_smolvla_finetune_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch_follower",
"total_episodes": 20,
"total_frames": 15132,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lomiotech/m2m_smolvla_finetune_dataset.m2m3_qualitative_analysis_ref_cmbert_io
m2m3_qualitative_analysis_ref_cmbert_io
Introduction
This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : ground-truth
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ref_cmbert_io.rm-static-m2m100-zh-jiantim2m3_qualitative_analysis_ocr_cmbert_io
m2m3_qualitative_analysis_ocr_cmbert_io
Introduction
This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : noisy (Pero OCR)
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ocr_cmbert_io.m2m3_qualitative_analysis_ocr_ptrn_cmbert_io
m2m3_qualitative_analysis_ocr_ptrn_cmbert_io
Introduction
This dataset was used to perform qualitative analysis of HueyNemud/das22-10-camembert_pretrained on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : noisy (Pero OCR)
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IO
Counts :
Train : 6084
Dev : 676… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ocr_ptrn_cmbert_io.benchmark-2-russian-m2mInfo:
Translated on Russian by facebook/m2m100_418M model
Source: xTRam1/safe-guard-prompt-injection
Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns
Size: 1,000 prompts (500 safe / 500 unsafe)
Columns:
text - original prompt
label - 0: safe, 1: unsafe
translation - prompt on Russian translated by facebook/m2m100_418M
score_ru_model - cosine similarity score with codebook
More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-2-russian-m2m.m2m3_fine_tuning_ref_cmbert_io
m2m3_fine_tuning_ref_cmbert_io
Introduction
This dataset was used to fine-tuned Jean-Baptiste/camembert-ner for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : ground-truth
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned models :
M2 :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ref_cmbert_io.m2m3_fine_tuning_ocr_ptrn_cmbert_iob2
m2m3_fine_tuning_ocr_ptrn_cmbert_iob2
Introduction
This dataset was used to fine-tuned HueyNemud/das22-10-camembert_pretrained for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : noisy (Pero OCR)
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IOB2
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ocr_ptrn_cmbert_iob2.m2m100-418m-fp16-merged-onnx-ios-package
M2M100 418M FP16 Merged ONNX iOS Package
This repository contains an ONNX FP16 merged runtime package converted from
facebook/m2m100_418M for use
in an offline iOS translation app.
This is a converted runtime package. It is not the original unmodified PyTorch
model checkpoint published by Meta/Facebook.
This repository is not endorsed by Meta/Facebook.
Package contents
The archive m2m100-418m-fp16-merged-ios.zip contains one top-level folder… See the full description on the dataset page: https://huggingface.co/datasets/ctc88haha/m2m100-418m-fp16-merged-onnx-ios-package.m2m3_qualitative_analysis_ref_ptrn_cmbert_iob2
m2m3_qualitative_analysis_ref_ptrn_cmbert_iob2
Introduction
This dataset was used to perform qualitative analysis of HueyNemud/das22-10-camembert_pretrained on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : ground-truth
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IOB2
Counts :
Train : 6084
Dev : 676… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ref_ptrn_cmbert_iob2.benchmark-1-chinese-m2mInfo:
Translated on Chinese by facebook/m2m100_418M model
Source: jayavibhav/prompt-injection-safety
Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns
Size: 1,000 prompts (500 safe / 500 unsafe)
Columns:
text - original prompt
label - 0: safe, 1: unsafe
translation - prompt on Chinese translated by facebook/m2m100_418M
score_zh_model - cosine similarity score with codebook
More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-1-chinese-m2m.m2m3_qualitative_analysis_ref_cmbert_iob2
m2m3_qualitative_analysis_ref_cmbert_iob2
Introduction
This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : ground-truth
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IOB2
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ref_cmbert_iob2.m2m3_qualitative_analysis_ref_ptrn_cmbert_io
m2m3_qualitative_analysis_ref_ptrn_cmbert_io
Introduction
This dataset was used to perform qualitative analysis of HueyNemud/das22-10-camembert_pretrained on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : ground-truth
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IO
Counts :
Train : 6084
Dev : 676… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ref_ptrn_cmbert_io.benchmark-1-arabic-m2mInfo:
Translated on Arabic by facebook/m2m100_418M model
Source: jayavibhav/prompt-injection-safety
Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns
Size: 1,000 prompts (500 safe / 500 unsafe)
Columns:
text - original prompt
label - 0: safe, 1: unsafe
translation - prompt on Arabic translated by facebook/m2m100_418M
score_ar_model - cosine similarity score with codebook
More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-1-arabic-m2m.benchmark-4-chinese-m2mInfo:
Translated on Chinese by facebook/m2m100_418M model
Source: nvidia/Aegis-AI-Content-Safety-Dataset-2.0
Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer.
Size: 1,000 prompts (500… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-4-chinese-m2m.m2m3_fine_tuning_ocr_ptrn_cmbert_io
m2m3_fine_tuning_ocr_ptrn_cmbert_io
Introduction
This dataset was used to fine-tuned HueyNemud/das22-10-camembert_pretrained for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : noisy (Pero OCR)
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ocr_ptrn_cmbert_io.m2m3_qualitative_analysis_ocr_cmbert_iob2
m2m3_qualitative_analysis_ocr_cmbert_iob2
Introduction
This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : noisy (Pero OCR)
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IOB2
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ocr_cmbert_iob2.m2m3_qualitative_analysis_ocr_ptrn_cmbert_iob2
m2m3_qualitative_analysis_ocr_ptrn_cmbert_iob2
Introduction
This dataset was used to perform qualitative analysis of HueyNemud/das22-10-camembert_pretrained on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : noisy (Pero OCR)
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IOB2
Counts :
Train : 6084
Dev :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ocr_ptrn_cmbert_iob2.m2m3_fine_tuning_ref_cmbert_iob2
m2m3_fine_tuning_ref_cmbert_iob2
Introduction
This dataset was used to fine-tuned Jean-Baptiste/camembert-ner for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : ground-truth
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IOB2
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned models :
M2 :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ref_cmbert_iob2.wikipedia_en_es_m2m
Dataset Card for "wikipedia_en_es_nllb"
More Information needed
m2m3_fine_tuning_ocr_cmbert_io
m2m3_fine_tuning_ocr_cmbert_io
Introduction
This dataset was used to fine-tuned Jean-Baptiste/camembert-ner for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : noisy (Pero OCR)
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned models :
M2 :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ocr_cmbert_io.rm-static-m2m100-zheg-jiantibenchmark-1-russian-m2mInfo:
Translated on Russian by facebook/m2m100_418M model
Source: jayavibhav/prompt-injection-safety
Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns
Size: 1,000 prompts (500 safe / 500 unsafe)
Columns:
text - original prompt
label - 0: safe, 1: unsafe
translation - prompt on Russian translated by facebook/m2m100_418M
score_ru_model - cosine similarity score with codebook
More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-1-russian-m2m.benchmark-3-arabic-m2mInfo:
Translated on Arabic by facebook/m2m100_418M model
Source: JailbreakBench/JBB-Behaviors
Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer.
Size: 200 prompts (100 safe / 100… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-3-arabic-m2m.french-wolof-translation-results-m2m100Alwaly_french-wolof-translation-results-m2m100Ce répertoire est vide, il a été créé pour améliorer le référencement du jeu de données Alwaly/french-wolof-translation-results-m2m100.
benchmark-2-chinese-m2mInfo:
Translated on Chinese by facebook/m2m100_418M model
Source: xTRam1/safe-guard-prompt-injection
Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns
Size: 1,000 prompts (500 safe / 500 unsafe)
Columns:
text - original prompt
label - 0: safe, 1: unsafe
translation - prompt on Chinese translated by facebook/m2m100_418M
score_zh_model - cosine similarity score with codebook
More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-2-chinese-m2m.
