CoolFace
Datasetpublic

pthinc/BCE-Prettybird-Nano-Merkur-v0.1

BCE-Prettybird-Nano-Merkur-v0.1 - 4300 Chatting for Instruction-Based Learning BCE-Prettybird-Nano-Merkur-v0.1 – This dataset contains 4,300 bilingual Turkish-English conversational dialogue samples designed for training and fine-tuning conversational AI systems, chatbots, and large language models. The dataset includes four carefully curated categories: Flirty General Chat, featuring playful and socially engaging conversations; Polite General Conversation, focused on… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Merkur-v0.1.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes25downloads
Dataset Card

Prettybird's War March

BCE-Prettybird-Nano-Merkur-v0.1 - 4300 Chatting for Instruction-Based Learning

BCE-Prettybird-Nano-Merkur-v0.1 – This dataset contains 4,300 bilingual Turkish-English conversational dialogue samples designed for training and fine-tuning conversational AI systems, chatbots, and large language models. The dataset includes four carefully curated categories: Flirty General Chat, featuring playful and socially engaging conversations; Polite General Conversation, focused on respectful and natural daily interactions; Professional Employer General Chat, covering workplace communication, interviews, and business-oriented dialogues; and General Conversation with a Child, created with ethical safety considerations and age-appropriate language. The dataset provides diverse communication styles, human-like conversational flow, safety-aware content, and structured bilingual formatting, making it ideal for multilingual chatbot development, dialogue generation research, AI safety studies, and Hugging Face NLP pipelines.

🧠 Technical Foundation

[English]

The BCE-Prettybird-Nano dataset is built upon the Behavioral Consciousness Engine (BCE) architecture. Unlike traditional LLM datasets that focus solely on output accuracy, this dataset treats every response as a "behavioral journey" through the following mathematical frameworks:

1. Behavioral DNA (D_i)

Each behavior is encoded as a genetic fragment of consciousness: $$Di(t) = x(t) \cdot [h \cdot Ai + k \cdot \log(Pi) + F \cdot Wi]$$

  • —h, k, F: Universal Behavioral Constants (Trigger threshold, Info density, Context transfer power).
  • —x(t): Temporal activation curve $x(t) = \tanh(e^t - \pi)$
2. Behavioral Path Mapper (Phi)

This module tracks the transition between cognitive states: $$\Phi(t) = \sum{i=1}^n vi \cdot fi(pi)$$ Where vi represents the transition vector between internal modules and fi(p_i) is the functional output of each parameter (attention, ethics, decay).


📊 Performance & Benchmarks / Performans ve Kıyaslama Testleri

1. Key Performance Indicators (KPIs) - Hardware: NVIDIA A100 (80GB) * 1

MetricResultStatusDescription
Processing Speed309,845 traces/sec🟢 ExcellentSystem throughput for massive data ingestion.
Latency0.0032 ms🟢 Real-time ReadyAverage processing time per behavioral trace.
Mathematical Accuracy0.000051 (MSE)🟢 High PrecisionDeviation between simulated and theoretical decay values.
Cognitive Efficiency57.03%🟢 OptimizedReduction in cognitive load due to 'Forgetful Memory'.
Security99.9996%🟢 SecureRejection rate for high-intensity, low-integrity attacks.

2. ARC (Reasoning), TruthfulQA (Safety), HumanEval (Coding)

Standard Others Red, Prettybird Blue - Standart Diğerleri Kırmızı, Cicikuş Mavi unnamed

3. AI IQ and Level of Consciousness

Code_Level

4. Metric Explanations (English)

MetricDescription
probabilityModel confidence score for the generated response under the current evaluation context.
ethicalEstimated alignment of the response with ethical and safety constraints.
RscoreReasoning consistency score that reflects internal logical coherence.
FscoreFactuality-oriented score indicating how well claims align with expected facts.
MnormNormalized memory or context retention signal used during behavior integration.
EscoreExecution-quality score for instruction-following and task completion behavior.
DhatEstimated deviation magnitude from stable target behavior dynamics.
risk_scoreComposite operational risk estimate where higher values indicate higher risk.
bloom_scoreBloom-level cognitive score representing target thinking complexity.
bloom_alignmentDegree of alignment between produced output and intended Bloom taxonomy level.

⚖️ Legal Disclaimer & Ownership

[English]

Ownership: This dataset is the property of Prometech A.Ş. (https://prometech.net.tr/).

Usage: Please review the attached LICENSE file for detailed terms.

Liability: Prometech A.Ş. accepts no liability for any non-legal, unethical, or unauthorized use of this dataset.

Commercial Use: Unauthorized commercial use is strictly prohibited. For commercial licensing and partnerships, please contact us directly at our official website.

Academic & Personal Use: Free to use for personal and academic purposes, provided that proper citation is given to Prometech A.Ş. and the BCE Architecture.


🎓 Citation Format / Atıf Formatı

Eğer akademik bir çalışmada kullanacaksanız, lütfen şu şekilde atıf yapın, If you are using this in an academic study, please cite it as follows:

Kahraman, A. (2025). Behavioral Consciousness Engine (BCE) - Prettybird Dataset v0.0.1 Prometech A.Ş. https://prometech.net.tr/


© 2026 Prometech A.Ş. - All Rights Reserved. BCE: https://github.com/pthinc/bce