# AIPedia — Full AI Knowledge Base (93 Terms) > Official Agentic AI Full Text Knowledge Base for AIPedia (https://aipedia.burmesestack.com). > Contains structured definitions (EL5, Standard, Technical), Burmese translations, topics, and level metadata for all 93 AI terms. --- ## Entry № 001: Artificial Intelligence (AI) **Burmese Title**: ဉာဏ်ရည်တု **Level**: beginner | **Topics**: fundamentals **URL**: https://aipedia.burmesestack.com/terms/artificial-intelligence ### Definitions - **EL5 (Simple)**: Artificial Intelligence ဆိုတာ ကွန်ပျူတာကို လူတွေလို စဉ်းစားနိုင်၊ သင်ယူနိုင်၊ ဆုံးဖြတ်နိုင်အောင် လုပ်ထားတဲ့ ဦးနှောက် အမျိုးအစားလို့ မြင်လို့ရတယ်။ ဥပမာ ဂိမ်းဆော့တာ၊ စာရိုက်တာကို ကြိုတွက်တာ၊ ဓာတ်ပုံထဲက မျက်နှာကို မှတ်မိတာတွေပါ။ ဒါပေမယ့် AI က လူလို ခံစားချက်မရှိဘူး — သူလုပ်တာက ဒေတာကြီးကြီးကနေ pattern တွေကို သင်ယူပြီး အလုပ်လုပ်တာပဲ။ - **Standard (Burmese)**: Artificial Intelligence (AI) သည် လူသားတို့၏ ဉာဏ်ရည်လိုအပ်သော လုပ်ငန်းများ — သင်ယူခြင်း၊ ဆင်ခြင်တုံတရားသုံးခြင်း၊ ပြဿနာဖြေရှင်းခြင်း၊ ဘာသာစကားနားလည်ခြင်းစသည်တို့ကို စက်များက လုပ်ဆောင်နိုင်စေရန် လေ့လာသော ကွန်ပျူတာသိပ္ပံဘာသာရပ်၏ နယ်ပယ်တစ်ခုဖြစ်သည်။ - **Technical (Deep)**: AI ဆိုသည်မှာ agents များကို environment မှ perception ရယူ၍ ပန်းတိုင်ဆီသို့ အကောင်းဆုံး action ရွေးချယ်နိုင်စေသော systems ၏ ဒီဇိုင်းနှင့် သိပ္ပံပညာဖြစ်သည်။ ၎င်းတွင် symbolic AI (rules & logic), statistical machine learning (data-driven), နှင့် modern deep learning (neural networks) အဆင့်များ ပါဝင်သည်။ ### Overview Artificial Intelligence သည် perception, reasoning, learning, planning, natural language processing နှင့် manipulation စသည့် ဦးနှောက်လုပ်ဆောင်ချက်များကို အတုယူသော computer systems များ ဖန်တီးခြင်း ဖြစ်သည်။ AI ကို ပုံမှန်အားဖြင့် Narrow/Weak AI (တိကျသော task များအတွက်) နှင့် General/Strong AI (လူကဲ့သို့ အကျယ်ပြန့်ဆုံး intelligence) ဟု ခွဲခြားသည် — လက်ရှိစနစ်အများစုသည် Narrow AI ဖြစ်သည်။ ### How It Works 1. Perception — sensors သို့မဟုတ် data sources မှ inputs များ (images, text, signals) ကို ရယူခြင်း 1. Representation & Reasoning — အချက်အလက်များကို structured form ဖြင့် ကိုယ်စားပြုပြီး logic သို့မဟုတ် probability ဖြင့် ဆင်ခြင်ခြင်း 1. Learning — ဒေတာမှ model ကို လေ့ကျင့်ခြင်း (supervised / unsupervised / reinforcement) 1. Action — trained model ကို သုံး၍ decisions သို့မဟုတ် predictions ထုတ်ပေးခြင်း ### Formulas - **Intelligent agent loop**: `π(a|s) : policy selecting action a given state s; U(s,a) = Σ_s' P(s'|s,a)·[R(s,a,s') + γ·U(s')]` --- ## Entry № 002: Machine Learning (ML) **Burmese Title**: စက်သင်ယူမှု **Level**: beginner | **Topics**: fundamentals, machine-learning **URL**: https://aipedia.burmesestack.com/terms/machine-learning ### Definitions - **EL5 (Simple)**: Machine Learning ဆိုတာ ကွန်ပျူတာကို rule တွေ တိုက်ရိုက်ရေးပေးစရာမလိုဘဲ ဥပမာပေါင်းများစွာ (data) ကြည့်ပြီး ကိုယ်တိုင် သင်ယူလိုက်တဲ့ နည်းလမ်း။ ကလေးက ခွေးဓာတ်ပုံ ဘယ်နှစ်ပုံလောက် ကြည့်ပြီးမှ ခွေးကို မှတ်မိလာသလိုပဲ — ကွန်ပျူတာကလည်း ဥပမာတွေကနေ pattern သင်ယူတယ်။ - **Standard (Burmese)**: Machine Learning သည် programming rules များ အတိအကျပေးစရာမလိုဘဲ data မှ patterns များကို လေ့လာ၍ prediction သို့မဟုတ် decision ပြုလုပ်နိုင်သော algorithms များကို လေ့လာသော AI ၏ အခွဲတစ်ရပ်ဖြစ်သည်။ - **Technical (Deep)**: ML သည် experience (data) E ရှိ performance P ကို task T တွင် measure လုပ်ပြီး E နှင့်အတူ P တိုးတက်လာသော program ဟု Tom Mitchell ၏ ဂန္တဝင်အဓိပ္ပာယ်ဖွင့်ဆိုချက်ဖြင့် သတ်မှတ်နိုင်သည်။ Training သည် parameters θ ကို objective function လျှော့ချရန် optimize လုပ်ခြင်းဖြစ်သည်။ ### Overview Machine Learning ၏ အနှစ်သာရမှာ 'generalization' — training မမြင်ဖူးသော data အသစ်တွင် ကောင်းစွာ လုပ်ဆောင်နိုင်ခြင်းဖြစ်သည်။ Paradigm သုံးမျိုး — supervised (labeled data), unsupervised (structure discovery), reinforcement (reward-driven) — ရှိသည်။ ### How It Works 1. Data collection & preprocessing — clean, normalize, split into train/test 1. Model selection — linear, tree, SVM, neural network စသည် 1. Training — loss function ကို minimize လုပ်ရန် optimizer (SGD) ဖြင့် parameters update 1. Evaluation — test set တွင် metrics (accuracy, precision...) တိုင်းတာခြင်း ### Formulas - **Empirical risk minimization**: `θ* = argmin_θ (1/N) Σ_i L(f(x_i; θ), y_i)` --- ## Entry № 003: Neural Network **Burmese Title**: အာရုံကြောကွန်ယက် **Level**: beginner | **Topics**: fundamentals, deep-learning **URL**: https://aipedia.burmesestack.com/terms/neural-network ### Definitions - **EL5 (Simple)**: Neural Network ဆိုတာ ဦးနှောက်ထဲက အာရုံကြောဆဲလ်လေးတွေလို ပုံစံလုပ်ထားတဲ့ အလွှာလိုက်တပ်ထားတဲ့ calculators တွေ။ တစ်ခုချင်းစီက နံပါတ်တွေလက်ခံပြီး လေးစားမှုတန်ဖိုး (weight) နဲ့ မြှောက်ပြီး နောက်တစ်လွှာကို ပို့တယ်။ ဥပမာတွေ အများကြီးကြည့်ပြီး weight တွေကို နည်းနည်းချင်း ချိန်ညှိလိုက်တာနဲ့ နောက်ဆုံး မှန်ကန်တဲ့အဖြေကို ပေးနိုင်လာတယ်။ - **Standard (Burmese)**: Neural Network သည် neurons များ၏ layers များဖြင့် ဖွဲ့စည်းထားသော computational model ဖြစ်ပြီး တစ်ခုစီက weighted inputs များကို activation function ဖြင့် အသွင်ပြောင်းကာ နောက်အလွှာသို့ ပို့သည်။ Training သည် weights များကို ချိန်ညှိရင်း error ကို လျှော့ချသည်။ - **Technical (Deep)**: A feedforward network computes h(l) = σ(W(l)h(l−1) + b(l)) per layer. Weights W နှင့် biases b ကို gradient descent (backpropagation) ဖြင့် optimize လုပ်သည်။ Universal approximation theorem အရ hidden layers လုံလောက်ပါက continuous function အများစုကို ခန့်မှန်းနိုင်သည်။ ### Overview Neural Networks များသည် matrix multiplications နှင့် non-linear activations များကို stacked လုပ်ထားသော function approximators များဖြစ်သည်။ Hidden layers များများလေလေ deep ဖြစ်လေလေ — ၎င်းကို Deep Learning ဟုခေါ်သည်။ ### How It Works 1. Forward pass — input x ကို layers ဖြတ်၍ prediction ŷ တွက်ချက်ခြင်း 1. Loss — ŷ နှင့် true y ကွာခြားမှုကို တိုင်းတာခြင်း 1. Backward pass — chain rule ဖြင့် loss ၏ gradient များကို weights များအတွက် တွက်ချက်ခြင်း 1. Update — optimizer က weights များကို gradient ဦးတည်ချက် ဆန့်ကျင်ဘက်သို့ ရွေ့ပေးခြင်း ### Formulas - **Forward pass of one layer**: `z(l) = W(l)·a(l−1) + b(l), a(l) = σ(z(l))` - **Mean squared error loss**: `L = (1/N) Σ_i (ŷ_i − y_i)²` --- ## Entry № 003: Perceptron / MLP **Burmese Title**: ပါဆက်ထရွန် / အလွှာစုံ ပါဆက်ထရွန် (Perceptron / MLP) **Level**: beginner | **Topics**: deep-learning **URL**: https://aipedia.burmesestack.com/terms/perceptron ### Definitions - **EL5 (Simple)**: Perceptron ဆိုတာ neural network တစ်ခုလုံးရဲ့ အသေးဆုံး အုတ်မြစ်လေး — input တွေကို weight နဲ့မြှောက်၊ ပေါင်း၊ ပြီးရင် 'ဖွင့်/ပိတ်' ဆုံးဖြတ်ပေးတဲ့ ဆဲလ်တစ်လုံးလိုပါ။ ဒီလို ဆဲလ်လေးတွေ အများကြီး တွဲလိုက်မှ neural network ကြီး ဖြစ်လာတာ။ MLP (Multi-Layer Perceptron) ဆိုတာ ဒီဆဲလ်လေးတွေကို အလွှာလိုက် စီထားတာပါ။ - **Standard (Burmese)**: Perceptron သည် artificial neuron တစ်ခု၏ အခြေခံပုံစံဖြစ်ပြီး input များကို weight ဖြင့်မြှောက်၊ ပေါင်းစည်း၍ activation function ဖြင့် output ထုတ်သည်။ Single perceptron သည် linear ခွဲခြားနိုင်သော ပြဿနာများကိုသာ ဖြေရှင်းနိုင်ပြီး Multi-Layer Perceptron (MLP) — hidden layer များ ထပ်ခြင်း — ဖြင့် non-linear ပြဿနာများကိုပါ ဖြေရှင်းနိုင်သည်။ - **Technical (Deep)**: Perceptron သည် y = f(wᵀx + b) ကို တွက်သည် (f သည် step/activation function)၊ ၎င်းသည် linear binary classifier ဖြစ်၍ XOR ကဲ့သို့ non-linearly separable ပြဿနာကို မဖြေနိုင်ပါ။ Perceptron များကို non-linear activation များနှင့် layer အဖြစ် ထပ်ခြင်းက Multi-Layer Perceptron (MLP) — backpropagation ဖြင့် လေ့ကျင့်သော fully-connected feedforward network — ကို ဖြစ်စေ၍ deep learning ၏ အခြေခံ architecture ဖြစ်သည်။ ### Overview Perceptron သည် 1958 ခုနှစ်က Frank Rosenblatt တီထွင်ခဲ့သော neural network ၏ အစဦး အုတ်မြစ်ဖြစ်သည်။ တစ်လုံးတည်းဖြင့် XOR ကဲ့သို့ non-linear ပြဿနာကို မဖြေနိုင်သော်လည်း၊ activation function နှင့် hidden layer များ ထပ်ပေါင်းသည့် MLP သည် ယနေ့ deep learning အားလုံး၏ အခြေခံဖြစ်လာသည်။ ### How It Works 1. Input တစ်ခုစီကို သက်ဆိုင်ရာ weight ဖြင့် မြှောက်သည် 1. ရလဒ်များကို ပေါင်းစည်း၍ bias ထည့်သည် (wᵀx + b) 1. Activation function ဖြင့် output ကို ဆုံးဖြတ်သည် 1. Layer များ ထပ်ခြင်းဖြင့် MLP ဖြစ်လာ၍ non-linear ပြဿနာများ ဖြေနိုင်သည် ### Formulas - **Perceptron output**: `y = f(w₁x₁ + w₂x₂ + … + wₙxₙ + b)` --- ## Entry № 004: Supervised Learning (SL) **Burmese Title**: ကြီးကြပ်သင်ယူမှု **Level**: beginner | **Topics**: fundamentals, machine-learning **URL**: https://aipedia.burmesestack.com/terms/supervised-learning ### Definitions - **EL5 (Simple)**: Supervised Learning ဆိုတာ ဆရာတစ်ယောက်နဲ့ သင်သလိုပဲ — မေးခွန်းတစ်ခုစီအတွက် အဖြေမှန် (label) တွေပါတဲ့ ဥပမာတွေကို ကွန်ပျူတာကို ပြပေးတယ်။ ဥပမာ ပန်းသီးဓာတ်ပုံတွေနဲ့ 'ပန်းသီး' ဆိုတဲ့အမည်ကို တွဲပြပြီး သင်ပေးရင် ဓာတ်ပုံအသစ်တွေကို သူကိုယ်တိုင် ခွဲခြားနိုင်လာမယ်။ - **Standard (Burmese)**: Supervised Learning သည် labeled data — input features နှင့် corresponding target labels တွဲပါသော dataset — မှ input မှ output သို့ မြေပုံဆွဲပေးသော mapping function တစ်ခုကို သင်ယူသော machine learning ပုံစံဖြစ်သည်။ Training ပြီးသော model ကို label မရှိသော data အသစ်များတွင် prediction လုပ်ရန် အသုံးပြုသည်။ - **Technical (Deep)**: Supervised learning တွင် training set {(x_i, y_i)} မှ hypothesis function f(x; θ) ကို သင်ယူပြီး true label y နှင့် prediction f(x; θ) ကွာခြားမှုကို loss function ဖြင့် တိုင်းတာ minimize လုပ်သည်။ Task အမျိုးအစားနှစ်မျိုးရှိသည် — regression (continuous output) နှင့် classification (discrete labels)။ ### Overview Supervised learning သည် ML ၏ အသုံးအများဆုံးပုံစံဖြစ်ပြီး spam detection, image classification, house price prediction စသည့် task များတွင် အသုံးပြုသည်။ Labeled data ရှိနေသရွေ့ အများစုသည် supervised learning ဖြင့် ဖြေရှင်းနိုင်သည်။ ### How It Works 1. Labeled dataset ကို train set နှင့် test set အဖြစ် ခွဲခြားခြင်း 1. Model (linear regression, decision tree, neural network) ကို ရွေးချယ်ခြင်း 1. Training data ဖြင့် model ၏ parameters များကို loss လျှော့ချရန် optimize လုပ်ခြင်း 1. Test set တွင် performance တိုင်းတာပြီး generalization ကို စစ်ဆေးခြင်း ### Formulas - **Supervised learning objective**: `θ* = argmin_θ Σ_i L(f(x_i; θ), y_i)` --- ## Entry № 004: Linear Regression **Burmese Title**: လီနီယာ ရီဂရက်ရှင်း (Linear Regression / မျဉ်းဖြောင့် ခန့်မှန်းနည်း) **Level**: beginner | **Topics**: machine-learning **URL**: https://aipedia.burmesestack.com/terms/linear-regression ### Definitions - **EL5 (Simple)**: Linear Regression ဆိုတာ အချက်အလက် အမှတ်တွေကြားမှာ အသင့်တော်ဆုံး မျဉ်းဖြောင့်တစ်ကြောင်း ဆွဲပြီး နောက်တန်ဖိုးကို ခန့်မှန်းတဲ့ နည်းလမ်း။ ဥပမာ - အိမ်ဧရိယာ ကျယ်လာရင် ဈေးနှုန်း ဘယ်လောက်တက်လာမလဲဆိုတာကို အမှတ်တွေကြားက မျဉ်းတစ်ကြောင်းနဲ့ ခန့်မှန်းတာမျိုးပါ။ - **Standard (Burmese)**: Linear Regression သည် input feature များ၏ linear combination ဖြင့် ဆက်တိုက်တန်ဖိုး (continuous value) တစ်ခုကို ခန့်မှန်းသော supervised learning algorithm ဖြစ်သည်။ Best-fit line ရှာရာတွင် ခန့်မှန်းချက်နှင့် အမှန်တန်ဖိုးကြား error (Mean Squared Error) ကို အနည်းဆုံးဖြစ်အောင် weight များကို ချိန်ညှိသည်။ - **Technical (Deep)**: Linear regression သည် target ကို ŷ = wᵀx + b အဖြစ် model လုပ်၍ mean squared error ကို minimize လုပ်ခြင်းဖြင့် parameter များ fit လုပ်သည်။ Normal equation w = (XᵀX)⁻¹Xᵀy ဖြင့် closed-form solution ရှိ၊ သို့မဟုတ် gradient descent ဖြင့် iterative သင်ယူနိုင်သည်။ အခြေခံ ယူဆချက်များတွင် linearity, independent error, homoscedasticity နှင့် multicollinearity နည်းခြင်း ပါဝင်သည်။ ### Overview Linear Regression သည် ML ၏ အခြေခံအကျဆုံး algorithm ဖြစ်ပြီး regression (ဆက်တိုက်တန်ဖိုး ခန့်မှန်းခြင်း) ၏ အုတ်မြစ်ဖြစ်သည်။ ရိုးရှင်းသော်လည်း interpretable ဖြစ်ပြီး coefficient တစ်ခုချင်းက feature တစ်ခု၏ သက်ရောက်မှုကို တိုက်ရိုက်ဖော်ပြသည်။ ပိုမိုရှုပ်ထွေးသော model များ၏ baseline အဖြစ် အမြဲအသုံးပြုသည်။ ### How It Works 1. Feature များနှင့် weight များကို မြှောက်ပေါင်း၍ ခန့်မှန်းချက် ŷ ကို ထုတ်သည် 1. Loss (MSE) — ခန့်မှန်းချက်နှင့် အမှန်တန်ဖိုးကြား ကွာဟမှုကို တိုင်းတာသည် 1. Gradient Descent သို့မဟုတ် Normal Equation ဖြင့် weight များကို ချိန်ညှိသည် 1. Loss အနည်းဆုံးဖြစ်သည့် best-fit line ကို ရရှိသည် ### Formulas - **Linear prediction**: `ŷ = w₁x₁ + w₂x₂ + … + wₙxₙ + b` --- ## Entry № 005: Unsupervised Learning (UL) **Burmese Title**: ကြီးကြပ်မှုမဲ့ သင်ယူမှု **Level**: beginner | **Topics**: fundamentals, machine-learning **URL**: https://aipedia.burmesestack.com/terms/unsupervised-learning ### Definitions - **EL5 (Simple)**: Unsupervised Learning ဆိုတာ label (အဖြေမှန်) တွေ မပါပဲ data ကြီးတစ်ပုံတည်း ကြည့်ပြီး ကွန်ပျူတာက pattern တွေကို ကိုယ်တိုင်ရှာတဲ့ နည်းလမ်း။ ဥပမာ ဓာတ်ပုံတွေအများကြီးထဲကနေ ဆင်တူတဲ့အမျိုးအစားတွေကို သူဘာသာသူ စုဖို့ကြိုးစားတာမျိုး — ဘယ်သူမှ အမည်တပ်ပေးမထားဘူး။ - **Standard (Burmese)**: Unsupervised Learning သည် labeled data မလိုအပ်ဘဲ data ၏ အတွင်းပိုင်းဖွဲ့စည်းပုံ (hidden structure) ကို ရှာဖွေသော machine learning ပုံစံဖြစ်သည်။ Clustering, dimensionality reduction, anomaly detection စသည့် task များတွင် အသုံးပြုသည်။ - **Technical (Deep)**: Unsupervised learning တွင် training set {x_i} သာရှိပြီး label y မရှိပေ။ Density estimation p(x) ကို သင်ယူခြင်း သို့မဟုတ် low-dimensional representation z ရရှိရန် data ၏ structure ကို ရှာဖွေသည်။ အဓိက method များမှာ k-means, DBSCAN, PCA, autoencoders နှင့် t-SNE တို့ဖြစ်သည်။ ### Overview Unsupervised learning သည် data ၏ hidden pattern များကို ရှာဖွေ၍ grouping, compression နှင့် anomaly detection စသည့် အလုပ်များအတွက် အသုံးပြုသည်။ Recommendation systems နှင့် customer segmentation တွင် တွင်ကျယ်စွာ အသုံးပြုပြီး labeled data ရှားပါးသော နယ်ပယ်များတွင် အလွန်အသုံးဝင်သည်။ ### How It Works 1. Label မရှိသော data ကို စုဆောင်းပြီး preprocessing လုပ်ခြင်း 1. Algorithm ရွေးချယ်ခြင်း (k-means clustering, PCA, autoencoder...) 1. Similarity measure ဖြင့် data points များကို အုပ်စုဖွဲ့ခြင်း သို့မဟုတ် compress လုပ်ခြင်း 1. Result ကို အဓိပ္ပာယ်ဖွင့်ဆိုပြီး လက်တွေ့အသုံးချခြင်း ### Formulas - **K-means objective**: `argmin_S Σ_j Σ_{x∈S_j} ‖x − μ_j‖²` --- ## Entry № 006: Natural Language Processing (NLP) **Burmese Title**: သဘာဝဘာသာစကား စီမံဆောင်ရွက်မှု **Level**: beginner | **Topics**: fundamentals, nlp **URL**: https://aipedia.burmesestack.com/terms/nlp ### Definitions - **EL5 (Simple)**: NLP ဆိုတာ ကွန်ပျူတာကို လူတွေသုံးတဲ့ စကားလုံး/စာသား (language) တွေကို နားလည်အောင် သင်ပေးတဲ့ ပညာရပ်။ 'မနက်ဖြန် ရာသီဥတုဘယ်လိုလဲ?' လို့ မေးရင် အဖြေရှာပေးတာ၊ chatbot တွေ၊ Google Translate တွေ အားလုံး NLP ကို သုံးတယ်။ - **Standard (Burmese)**: Natural Language Processing (NLP) သည် စက်များက human language — text နှင့် speech — ကို နားလည်၊ ခွဲခြမ်းစိတ်ဖြာ၊ ထုတ်လုပ်နိုင်စေရန် လေ့လာသော AI ၏ နယ်ပယ်တစ်ခုဖြစ်သည်။ Text classification, translation, question answering, sentiment analysis စသည့် task များ ပါဝင်သည်။ - **Technical (Deep)**: NLP သည် raw text ကို numerical representation အဖြစ် ပြောင်း၍ model များဖြင့် process လုပ်သည်။ Pipeline — tokenization → normalization → embedding → sequence modeling (RNN/Transformer) → decoding။ Modern NLP သည် pre-trained transformers (BERT, GPT) ပေါ်တွင် အခြေခံပြီး fine-tuning ဖြင့် task-specific လုပ်ဆောင်သည်။ ### Overview NLP သည် စာသားများကို နားလည်ခြင်းနှင့် ထုတ်လုပ်ခြင်းတို့ကို ဆောင်ရွက်သည်။ Tokenization မှစ၍ embedding, attention mechanism, transformer များအထိ pipeline ဖြင့် အလုပ်လုပ်သည်။ Chatbots, search engines, translation နှင့် summarization တွင် အသုံးပြုသည်။ ### How It Works 1. Text ကို token များအဖြစ် ပိုင်းခြားခြင်း (tokenization) 1. Token တစ်ခုချင်းစီကို vector (embedding) အဖြစ် ပြောင်းခြင်း 1. Sequence model (RNN / Transformer) ဖြင့် context ကို process လုပ်ခြင်း 1. Task-specific output ထုတ်ခြင်း — classification, generation, translation 1. Trained model ကို fine-tune လုပ်ပြီး သီးသန့် task အတွက် အသုံးပြုခြင်း ### Formulas - **Text classification softmax**: `P(y=k|x) = exp(z_k) / Σ_j exp(z_j)` --- ## Entry № 006: Logistic Regression **Burmese Title**: လော့ဂျစ်စတစ် ရီဂရက်ရှင်း (Logistic Regression / အမျိုးအစားခွဲ ခန့်မှန်းနည်း) **Level**: beginner | **Topics**: machine-learning **URL**: https://aipedia.burmesestack.com/terms/logistic-regression ### Definitions - **EL5 (Simple)**: Logistic Regression ဆိုတာ 'ဟုတ်/မဟုတ်' လို အမျိုးအစား နှစ်ခုကြားက ဖြစ်နိုင်ခြေ (probability) ကို ခန့်မှန်းတဲ့ နည်းလမ်း။ ဥပမာ - အီးမေးလ်တစ်စောင်ဟာ spam ဖြစ်နိုင်ခြေ ဘယ်လောက်ရှိလဲ (0 ကနေ 1) ဆိုတာ တွက်ပြီး အဆုံးအဖြတ် ချတာမျိုးပါ။ - **Standard (Burmese)**: Logistic Regression သည် input များ၏ linear combination ရလဒ်ကို sigmoid function ဖြင့် 0–1 အတွင်း probability အဖြစ် ပြောင်းလဲပေးသော classification algorithm ဖြစ်သည်။ Threshold (ပုံမှန် 0.5) ကို ကျော်လွန်ပါက class တစ်ခု၊ မကျော်ပါက အခြား class ဟု ဆုံးဖြတ်သည်။ - **Technical (Deep)**: အမည်တွင် regression ပါသော်လည်း logistic regression သည် classification model ဖြစ်သည်။ z = wᵀx + b ကို တွက်၍ sigmoid σ(z) = 1/(1+e⁻ᶻ) မှတစ်ဆင့် P(y=1|x) ကို ရယူပြီး binary cross-entropy (log loss) ကို minimize လုပ်၍ လေ့ကျင့်သည်။ Softmax function က ၎င်းကို multi-class ပြဿနာများသို့ generalize လုပ်သည်။ ### Overview Logistic Regression သည် 'regression' ဟု အမည်ရသော်လည်း အမှန်တကယ်တွင် classification model ဖြစ်သည်။ ရိုးရှင်း၊ မြန်ဆန်ပြီး interpretable ဖြစ်၍ binary classification (spam/ham, ရောဂါရှိ/မရှိ) များတွင် baseline အဖြစ် ကျယ်ပြန့်စွာ အသုံးပြုသည်။ ### How It Works 1. Input များ၏ weighted sum z = wᵀx + b ကို တွက်သည် 1. Sigmoid function ဖြင့် z ကို 0–1 probability အဖြစ် ပြောင်းသည် 1. Binary cross-entropy loss ကို minimize ရန် weight များ ချိန်ညှိသည် 1. Threshold (ပုံမှန် 0.5) ဖြင့် class ကို ဆုံးဖြတ်သည် ### Formulas - **Sigmoid (logistic function)**: `σ(z) = 1 / (1 + e^(−z)), z = wᵀx + b` --- ## Entry № 006: K-Nearest Neighbors (KNN) **Burmese Title**: ကေ-အနီးဆုံး အိမ်နီးချင်းများ (K-Nearest Neighbors) **Level**: beginner | **Topics**: machine-learning **URL**: https://aipedia.burmesestack.com/terms/k-nearest-neighbors ### Definitions - **EL5 (Simple)**: K-Nearest Neighbors (KNN) ဆိုတာ အသစ်တစ်ခုကို ခွဲခြားဖို့ သူနဲ့ အနီးဆုံး အိမ်နီးချင်း K ခုကို ကြည့်ပြီး 'အများစုက ဘာအမျိုးအစားလဲ' ဆိုတာနဲ့ ဆုံးဖြတ်တဲ့ နည်း။ ဥပမာ - လူသစ်တစ်ယောက်ကို သူ့ပတ်ဝန်းကျင်က အနီးဆုံး သူငယ်ချင်း ၅ ယောက်ကြည့်ပြီး ဘယ်အုပ်စုနဲ့ တူလဲ ခန့်မှန်းသလိုပါ။ - **Standard (Burmese)**: KNN သည် training data ကို သိမ်းထားပြီး ခန့်မှန်းချိန်ကျမှသာ တွက်ချက်သော instance-based (lazy) algorithm ဖြစ်သည်။ Data point အသစ်တစ်ခုနှင့် အနီးဆုံး K ခုကို distance (ဥပမာ Euclidean) ဖြင့်ရှာ၍ ၎င်းတို့၏ label အများစု (classification) သို့မဟုတ် ပျမ်းမျှတန်ဖိုး (regression) ဖြင့် ဆုံးဖြတ်သည်။ - **Technical (Deep)**: KNN သည် non-parametric, instance-based method ဖြစ်သည်။ Query point တစ်ခုအတွက် distance metric (Euclidean, Manhattan, cosine) အောက် အနီးဆုံး training sample K ခုကို ရယူ၍ majority vote (classification) သို့မဟုတ် mean (regression) ဖြင့် ခန့်မှန်းသည်။ K ရွေးချယ်မှုသည် bias–variance ကို အလဲအလှယ်လုပ်; feature များ scale ရမည်ဖြစ်ပြီး naive search သည် query တစ်ခုလျှင် O(N) ကုန်ကျ — KD-tree သို့ approximate nearest-neighbor index ဖြင့် သက်သာစေသည်။ ### Overview KNN သည် သင်ယူစရာ parameter မရှိသော (non-parametric) ရိုးရှင်းသည့် algorithm ဖြစ်ပြီး 'ဆင်တူသည့်အရာများ အနီးအနားတွင် ရှိတတ်သည်' ဆိုသော အယူအဆပေါ် အခြေခံသည်။ K ရွေးချယ်မှုနှင့် feature scaling က performance ကို များစွာ သက်ရောက်သည်။ ### How It Works 1. Data point အသစ်နှင့် training point အားလုံးကြား distance တွက်သည် 1. အနီးဆုံး K ခုကို ရွေးသည် 1. Classification — label အများစု; Regression — ပျမ်းမျှတန်ဖိုး 1. K ကြီးလေ ချောမွေ့ (bias) လေ၊ K ငယ်လေ noise (variance) များလေ ### Formulas - **Euclidean distance**: `d(x, x′) = √( Σᵢ (xᵢ − x′ᵢ)² )` --- ## Entry № 007: Computer Vision (CV) **Burmese Title**: ကွန်ပျူတာအမြင် **Level**: beginner | **Topics**: fundamentals, computer-vision **URL**: https://aipedia.burmesestack.com/terms/computer-vision ### Definitions - **EL5 (Simple)**: Computer Vision ဆိုတာ ကွန်ပျူတာကို ဓာတ်ပုံ/ဗီဒီယိုတွေကနေ မြင်ပြီး နားလည်အောင် သင်ပေးတဲ့ ပညာရပ်။ ဓာတ်ပုံထဲမှာ ခွေးလား ကြောင်လား ခွဲခြားတာ၊ မျက်နှာကို မှတ်မိတာ၊ ကားတွေကို လမ်းပေါ်မှာ ရှာတွေ့တာတွေ အားလုံး computer vision ပဲ။ - **Standard (Burmese)**: Computer Vision သည် images နှင့် videos များမှ အဓိပ္ပာယ်ရှိသော information ကို ထုတ်ယူနိုင်စေရန် စက်များကို သင်ကြားပေးသော AI နယ်ပယ်တစ်ခုဖြစ်သည်။ Image classification, object detection, segmentation, face recognition စသည့် task များ ပါဝင်သည်။ - **Technical (Deep)**: Computer vision သည် raw pixels များကို convolutional neural networks (CNNs) ဖြင့် process လုပ်ကာ hierarchical features — edges, shapes, objects — များကို သင်ယူသည်။ ImageNet စသော large datasets များတွင် လေ့ကျင့်ထားသော models များကို transfer learning ဖြင့် အသုံးပြုလေ့ရှိသည်။ ### Overview Computer vision သည် images မှ visual information ကို နားလည်ရန် CNNs နှင့် vision transformers များကို အသုံးပြုသည်။ Pixels များကို tensors အဖြစ် ကိုယ်စားပြုပြီး model များက classification, detection, segmentation များ လုပ်ဆောင်သည်။ ### How It Works 1. Image ကို pixel matrix (height × width × channels) အဖြစ် load လုပ်ခြင်း 1. Convolutional layers များဖြင့် features (edges, textures) ထုတ်ယူခြင်း 1. Pooling ဖြင့် အတိုင်းအတာလျှော့ပြီး အရေးကြီးသော features များ ထိန်းသိမ်းခြင်း 1. Fully-connected layers များဖြင့် classification / detection output ထုတ်ခြင်း ### Formulas - **2D convolution**: `S(i,j) = (I * K)(i,j) = Σ_m Σ_n I(i+m, j+n) · K(m,n)` --- ## Entry № 007: K-Means Clustering **Burmese Title**: ကေ-မိန်း အစုဖွဲ့နည်း (K-Means Clustering) **Level**: beginner | **Topics**: machine-learning **URL**: https://aipedia.burmesestack.com/terms/k-means-clustering ### Definitions - **EL5 (Simple)**: K-Means ဆိုတာ label မပါတဲ့ data တွေကို ဆင်တူရာချင်း အုပ်စု K စု ခွဲပေးတဲ့ နည်း။ ဥပမာ - ဖောက်သည်တွေရဲ့ အမူအကျင့်ကို ကြည့်ပြီး 'ဒီအုပ်စု၊ ဟိုအုပ်စု' ဆိုပြီး ဘယ်သူမှ ကြိုမပြောဘဲ ကွန်ပျူတာက သူ့ဘာသာသူ စုပေးတာမျိုးပါ။ - **Standard (Burmese)**: K-Means သည် unsupervised clustering algorithm ဖြစ်ပြီး data များကို K အုပ်စု (cluster) ခွဲရာတွင် cluster center (centroid) နှင့် အနီးဆုံး point များကို တွဲစပ်ပေးသည်။ Point များကို centroid သို့ ပြန်တွဲ၊ centroid များကို ပြန်တွက် ဆိုသည့် အဆင့်များကို convergence ရသည်အထိ ထပ်ခါထပ်ခါ လုပ်ဆောင်သည်။ - **Technical (Deep)**: K-Means သည် within-cluster sum of squares ကို minimize လုပ်၍ point N ခုကို cluster K ခု ခွဲသည်။ Lloyd's algorithm က assignment step (point → အနီးဆုံး centroid) နှင့် update step (centroid = assigned point များ၏ mean) ကို convergence အထိ အလှည့်ကျ လုပ်သည်။ K ရွေးရ (elbow/silhouette)၊ initialization (k-means++ ဖြင့် သက်သာ) နှင့် feature scale အပေါ် အထိခိုက်လွယ်ပြီး roughly spherical, အရွယ်တူ cluster များဟု ယူဆသည်။ ### Overview K-Means သည် အသုံးအများဆုံး clustering algorithm ဖြစ်ပြီး customer segmentation, image compression, anomaly detection စသည်တို့တွင် သုံးသည်။ K ကို ကြိုတင်သတ်မှတ်ရခြင်းနှင့် initialization အပေါ်မူတည်၍ ရလဒ်ပြောင်းနိုင်ခြင်းက ၎င်း၏ အဓိက ကန့်သတ်ချက်များဖြစ်သည်။ ### How It Works 1. Cluster အရေအတွက် K ကို ရွေးသည် 1. Centroid K ခုကို ကျပန်း (သို့) k-means++ ဖြင့် စတင်ချသည် 1. Point တစ်ခုစီကို အနီးဆုံး centroid သို့ တွဲသည် (assignment) 1. Centroid များကို assigned point များ၏ ပျမ်းမျှဖြင့် ပြန်တွက်သည် (update) 1. ပြောင်းလဲမှု မရှိတော့သည်အထိ ထပ်ခါ လုပ်သည် --- ## Entry № 008: Generative AI (GenAI) **Burmese Title**: ဖန်တီးနိုင်သော AI **Level**: beginner | **Topics**: fundamentals, generative-ai **URL**: https://aipedia.burmesestack.com/terms/generative-ai ### Definitions - **EL5 (Simple)**: Generative AI ဆိုတာ မရှိသေးတဲ့ အကြောင်းအရာအသစ်တွေကို ကွန်ပျူတာက ကိုယ်တိုင်ဖန်တီးပေးတဲ့ AI။ ပန်းချီပုံအသစ်ဆွဲပေးတာ၊ စာစီစာကုံးရေးပေးတာ၊ သီချင်းဖန်တီးပေးတာ၊ ဓာတ်ပုံထဲက မျက်နှာကို ပြောင်းပေးတာတွေ လုပ်နိုင်တယ်။ သူလုပ်တာက သင်ယူထားတဲ့ ဥပမာတွေရဲ့ ပုံစံကို လိုက်ပြီး အသစ်တွေ ဖန်တီးတာပဲ။ - **Standard (Burmese)**: Generative AI သည် training data ၏ distribution ကို သင်ယူပြီး အသစ်သော၊ မမြင်ဖူးသော content — text, images, audio, video — များကို ဖန်တီးနိုင်သော AI နယ်ပယ်ဖြစ်သည်။ Large language models (LLMs), diffusion models နှင့် GANs များပေါ်တွင် အခြေခံသည်။ - **Technical (Deep)**: Generative models များသည် data distribution p(x) ကို model လုပ်ကာ samples အသစ်များ ထုတ်လုပ်သည်။ LLM များသည် sequence တစ်ခုစီအတွက် next token probability p(x_t | x_, — များကိုလည်း ထည့်လေ့ရှိသည်။ ### Overview Tokenization သည် text ကို discrete units များအဖြစ် ပိုင်းခြားကာ vocabulary ids များသို့ mapping လုပ်သည်။ Word-level, character-level, subword-level ဆိုပြီး အဆင့်သုံးမျိုးရှိပြီး subword (BPE/WordPiece) သည် modern models များတွင် အသုံးအများဆုံးဖြစ်သည်။ ### How It Works 1. Text ကို whitespace/punctuation ဖြင့် ကနဦး ခွဲခြားခြင်း (pre-tokenization) 1. Common subword units များကို dataset ပေါ်တွင် သင်ယူခြင်း (training a tokenizer) 1. Text ကို vocabulary ရှိ token ids များသို့ ပြောင်းခြင်း 1. Special tokens (start/end) နှင့် attention masks များ ပေါင်းထည့်ခြင်း ### Formulas - **BPE merge operation**: `freq(a,b) = count of adjacent pair (a,b); merge the most frequent pair into one token` --- ## Entry № 011: Loss Function **Burmese Title**: ဆုံးရှုံးမှုလုပ်ဆောင်ချက် **Level**: beginner | **Topics**: machine-learning, deep-learning **URL**: https://aipedia.burmesestack.com/terms/loss-function ### Definitions - **EL5 (Simple)**: Loss Function ဆိုတာ model ရဲ့ အဖြေက ဘယ်လောက်မှားနေလဲဆိုတာကို တိုင်းတဲ့ ကိရိယာ။ ဆရာ့အဖြေနဲ့ ကျောင်းသားရဲ့အဖြေ ကွာခြားချက်ကို နံပါတ်တစ်ခုနဲ့ ပြသလိုပဲ — နံပါတ်သေးလေ မှားမှုနည်းလေ၊ သုညဆိုရင် အဖြေမှန်ကန်လေဖြစ်တယ်။ - **Standard (Burmese)**: Loss Function သည် model ၏ prediction နှင့် true target အကြား ကွာခြားမှုကို scalar value တစ်ခုဖြင့် တိုင်းတာသော function ဖြစ်သည်။ Training ၏ ပန်းတိုင်မှာ ဤ loss ကို လျှော့ချရန်ဖြစ်ပြီး gradient descent ကဲ့သို့ optimizers များက loss အပေါ်အခြေခံ၍ parameters များကို ပြုပြင်သည်။ - **Technical (Deep)**: Loss L(θ) = (1/N) Σ_i ℓ(f(x_i; θ), y_i) ကို model parameters θ ၏ function အဖြစ် minimize လုပ်သည်။ Task အလိုက် loss ရွေးသည် — regression: MSE/MAE; classification: cross-entropy; robust regression: Huber loss။ Soft labels သို့မဟုတ် auxiliary losses များဖြင့် optimization landscape ကို ချောမွေ့စေနိုင်သည်။ ### Overview Loss function သည် optimization ၏ compass ဖြစ်သည် — model ဘယ်လောက်မှားနေလဲကို တိုင်းတာပြီး gradient descent ကို လမ်းညွှန်သည်။ Task အလိုက် loss ရွေးချယ်မှုက training ၏ အရည်အသွေးကို တိုက်ရိုက်သက်ရောက်သည်။ ### How It Works 1. Forward pass ဖြင့် model prediction ŷ ကို တွက်ခြင်း 1. Prediction ŷ နဲ့ true label y ကို loss function ထဲထည့်တွက်ခြင်း 1. Scalar loss value တစ်ခု ရရှိခြင်း 1. Backward pass ဖြင့် loss ၏ gradient များကို weights တွေဆီ ဖြန့်ခြင်း 1. Optimizer က gradient ကိုသုံးပြီး weights တွေကို update လုပ်ခြင်း ### Formulas - **Cross-entropy loss**: `L = −Σ_c y_c · log(ŷ_c)` - **Mean squared error**: `L = (1/N) Σ_i (ŷ_i − y_i)²` --- ## Entry № 012: Gradient Descent (GD) **Burmese Title**: ဂရေဒီယင့်ဆင်းသက်နည်း **Level**: beginner | **Topics**: machine-learning, deep-learning **URL**: https://aipedia.burmesestack.com/terms/gradient-descent ### Definitions - **EL5 (Simple)**: Gradient Descent ဆိုတာ နှင်းတွေထူတဲ့ညမှာ တောင်ပေါ်ကနေ အောက်ဆုံးကို မမြင်ရဘဲ ဆင်းဖို့ကြိုးစားသလိုပဲ — ခြေလှမ်းတိုင်း မြေကြီးက ဘယ်ဘက်စောင်းနေလဲ စမ်းသပ်ပြီး အောက်ဘက်ကို နည်းနည်းချင်း လျှောက်သွားတယ်။ Model ရဲ့ loss ဆိုတဲ့တောင်ကို အောက်ဆုံးရောက်အောင် (loss အနည်းဆုံး) ဆင်းတာပဲ။ - **Standard (Burmese)**: Gradient Descent သည် loss function ကို minimize လုပ်ရန် model parameters များကို ထပ်ခါတလဲလဲ update လုပ်သော optimization algorithm ဖြစ်သည်။ ခြေလှမ်းတိုင်း gradient (loss ၏ slope) ၏ ဦးတည်ချက်ဆန့်ကျင်ဘက်သို့ learning rate အတိုင်း ရွေ့သည်။ - **Technical (Deep)**: Parameters θ ကို rule θ := θ − η·∇θL(θ) ဖြင့် update လုပ်သည် — η သည် learning rate၊ ∇θL သည် loss ၏ gradient ဖြစ်သည်။ Variants — batch GD (full dataset), stochastic GD (single sample), mini-batch GD (subset)။ Optimizers များ (Momentum, RMSprop, Adam) က convergence speed နှင့် stability ကို မြှင့်တင်သည်။ ### Overview Gradient descent သည် neural networks များကို လေ့ကျင့်ရာတွင် အဓိက engine ဖြစ်သည်။ Loss landscape ၏ gradient ကို backpropagation ဖြင့် တွက်ပြီး parameters များကို steepest descent ဦးတည်ချက်သို့ ရွေ့သည်။ Learning rate ရွေးချယ်မှုသည် convergence အတွက် အရေးကြီးသည်။ ### How It Works 1. Forward pass ဖြင့် loss L(θ) ကို တွက်ခြင်း 1. Backpropagation ဖြင့် gradient ∇L ကို parameters တစ်ခုစီအတွက် တွက်ခြင်း 1. Update rule θ := θ − η·∇L ဖြင့် parameters များကို ရွေ့ခြင်း 1. Loss converge ရသည်အထိ ထပ်တလဲလဲ ပြုလုပ်ခြင်း ### Formulas - **Gradient descent update rule**: `w := w − η · ∇w L(w)` - **Mini-batch SGD update**: `θ := θ − η · (1/B) Σ_i ∇θ ℓ(x_i, y_i; θ)` --- ## Entry № 013: Overfitting **Burmese Title**: အလွန်အကျွံ လိုက်ဖက်ခြင်း **Level**: intermediate | **Topics**: machine-learning, deep-learning **URL**: https://aipedia.burmesestack.com/terms/overfitting ### Definitions - **EL5 (Simple)**: Overfitting ဆိုတာ ကျောင်းသားတစ်ယောက်က စာမေးပွဲမေးခွန်းတွေကို အာဂုံမှတ်ပြီး မေးခွန်းအသစ်တွေကျတော့ မဖြေနိုင်တာနဲ့တူတယ်။ Model က training data ကို အလွန်အကျွံ မှတ်မိပြီး data အသစ် (test) မှာ စွမ်းဆောင်ရည်ကျသွားတာ။ Training မှာ score မြင့်ပေမယ့် test မှာနိမ့်ရင် overfitting ဖြစ်နေပြီ။ - **Standard (Burmese)**: Overfitting သည် model က training data ၏ noise နှင့် details များကို အလွန်အကျွံ သင်ယူမိ၍ generalization ဆိုးသွားသည့် ဖြစ်စဉ်ဖြစ်သည်။ ၎င်းကို training performance မြင့်မားသော်လည်း validation/test performance နိမ့်ကျခြင်းဖြင့် ဖော်ပြသည် — train-test gap ဟုခေါ်သည်။ - **Technical (Deep)**: Overfitting သည် model capacity က training data ၏ signal ထက် ကျော်လွန်သောအခါ ဖြစ်ပွားသည် — loss curve တွင် training loss ဆက်ကျသော်လည်း validation loss ပြန်တက်လာသည်။ Remedies — more data, regularization (L1/L2, dropout), early stopping, data augmentation, cross-validation, simpler models။ Bias-variance tradeoff ၏ variance ဘက်သို့ ယိမ်းခြင်းဖြစ်သည်။ ### Overview Overfitting သည် ML ၏ အဖြစ်အများဆုံး ပြဿနာတစ်ခုဖြစ်သည် — model က training data ကို memorize လုပ်လွန်းလို့ မမြင်ဖူးသော data တွင် စွမ်းဆောင်ရည်ကျသည်။ Train-test gap ကို စောင့်ကြည့်ခြင်းဖြင့် သိရှိနိုင်ပြီး regularization နှင့် data ပိုခြင်းဖြင့် ကာကွယ်နိုင်သည်။ ### How It Works 1. Model က training data ရှိ pattern နှင့် noise နှစ်မျိုးလုံးကို သင်ယူခြင်း 1. Training loss ကျဆင်းသော်လည်း validation loss ပြန်တက်လာခြင်း 1. Train-test gap ကြီးထွားလာခြင်း 1. Regularization (dropout, weight decay) သို့မဟုတ် early stopping ဖြင့် ထိန်းချုပ်ခြင်း ### Formulas - **Train–test gap**: `Gap = L_train(θ) − L_val(θ); large positive gap → overfitting` - **L2 (ridge) regularization**: `L'(θ) = L(θ) + λ · Σ θ_i²` --- ## Entry № 013: Hallucination **Burmese Title**: ဟေလူစီနေးရှင်း (Hallucination / AI လုပ်ကြံ အဖြေ) **Level**: intermediate | **Topics**: generative-ai, nlp **URL**: https://aipedia.burmesestack.com/terms/hallucination ### Definitions - **EL5 (Simple)**: Hallucination ဆိုတာ AI (အထူးသဖြင့် LLM) က တကယ်မဟုတ်တဲ့၊ မှားနေတဲ့ အချက်အလက်တွေကို 'သေချာသလိုပုံစံနဲ့' ယုံကြည်လောက်အောင် လုပ်ပြီး ပြောဆိုတာ။ ဥပမာ - မရှိတဲ့ စာအုပ်နာမည်၊ မှားနေတဲ့ ရက်စွဲ၊ လုပ်ကြံ ကိုးကားချက်တွေ ဖန်တီးပြောတာမျိုးပါ။ AI က 'မသိဘူး' လို့ မပြောဘဲ ဟန်ဆောင် ဖြေဆိုတတ်လို့ သတိထားရတယ်။ - **Standard (Burmese)**: Hallucination သည် LLM က training data တွင် အခြေအမြစ်မရှိသော သို့မဟုတ် အမှန်တရားနှင့် ကွဲလွဲသော အချက်အလက်များကို ယုံကြည်လောက်ဖွယ် ပုံစံဖြင့် ထုတ်ပေးသည့် ဖြစ်စဉ်ဖြစ်သည်။ Model သည် အမှန်တရားကို 'သိ'သည်မဟုတ်ဘဲ ဖြစ်နိုင်ခြေအရှိဆုံး နောက်စကားလုံးကို ခန့်မှန်းသောကြောင့် ဖြစ်ပေါ်ခြင်းဖြစ်သည်။ RAG, grounding, guardrails တို့ဖြင့် လျှော့ချနိုင်သည်။ - **Technical (Deep)**: Hallucination ဆိုသည်မှာ model က ချောမွေ့သော်လည်း အခြေအမြစ်မရှိ/လုပ်ကြံ content ကို ထုတ်ခြင်းဖြစ်သည်။ LLM များကို ဖြစ်နိုင်ခြေရှိသော token ခန့်မှန်းရန် လေ့ကျင့်ထားခြင်း (truth စစ်ဆေးရန် မဟုတ်) နှင့် အချက်အလက် ရင်းမြစ်သို့ built-in grounding မရှိခြင်းကြောင့် ဖြစ်ပေါ်သည်။ လျှော့ချနည်း — retrieval-augmented generation (document သို့ grounding), citation/attribution, temperature နိမ့်, tool use, guardrail/verification layer; အပြီးအပိုင် ဖျောက်၍ မရ။ ### Overview Hallucination သည် LLM များ၏ အဓိက ယုံကြည်စိတ်ချရမှု (reliability) ပြဿနာဖြစ်ပြီး production application များတွင် အရေးအကြီးဆုံး စိန်ခေါ်မှုတစ်ခုဖြစ်သည်။ Model က 'မသိပါ' ဟု ဝန်ခံမည့်အစား ယုံကြည်လောက်ဖွယ် အမှားကို ဖန်တီးတတ်သောကြောင့် factual task များတွင် grounding နှင့် verification မဖြစ်မနေ လိုအပ်သည်။ ### How It Works 1. LLM သည် အမှန်တရားကို စစ်ဆေးခြင်းမဟုတ်ဘဲ ဖြစ်နိုင်ခြေအရှိဆုံး token ကို ခန့်မှန်းသည် 1. Training data တွင် မရှိ/မှားသော အချက်ကိုပင် ချောမွေ့စွာ ဖန်တီးထုတ်သည် 1. Confidence မြင့်ပုံစံဖြင့် ပြောသောကြောင့် ခွဲခြားရ ခက်သည် 1. RAG, grounding, temperature နိမ့်, verification ဖြင့် လျှော့ချသည် --- ## Entry № 014: Deep Learning **Burmese Title**: နက်ရှိုင်းသောသင်ယူမှု **Level**: intermediate | **Topics**: machine-learning, deep-learning **URL**: https://aipedia.burmesestack.com/terms/deep-learning ### Definitions - **EL5 (Simple)**: Deep Learning ဆိုတာ neural network ကို အလွှာတွေ အများကြီးထပ်ပြီး ဆောက်ထားတဲ့ Machine Learning နည်းလမ်း။ အလွှာတစ်ခုချင်းစီက ရုပ်ပုံရဲ့ အစိတ်အပိုင်းလေးတွေ (မျဉ်း၊ ထောင့်၊ မျက်နှာလေးတွေ) ကို တစ်ဆင့်ချင်း သင်ယူသွားတယ်။ ကလေးတစ်ယောက် ရုပ်ပုံကြည့်ပြီး 'ဒါ မျက်လုံး၊ ဒါ မျက်နှာ' လို့ အဆင့်ဆင့် မှတ်မိသလိုပဲ။ - **Standard (Burmese)**: Deep Learning သည် hidden layers များစွာပါဝင်သော neural networks များကို အသုံးပြု၍ data မှ hierarchical representations များကို သင်ယူသော Machine Learning ၏ အခွဲတစ်ရပ်ဖြစ်သည်။ Layer အရေအတွက် များလေလေ abstraction အဆင့်များ ပိုနက်ရှိုင်းလေလေဖြစ်ပြီး images, audio, text ကဲ့သို့ high-dimensional data များတွင် ထူးချွန်စွာ လုပ်ဆောင်နိုင်သည်။ - **Technical (Deep)**: Deep learning ဆိုသည်မှာ hidden layer များစွာ (depth) ပါဝင်သော neural network များကို backpropagation နှင့် gradient descent ဖြင့် end-to-end လေ့ကျင့်ခြင်းဖြစ်သည်။ Deep architecture များက feature hierarchy ကို သင်ယူသည် — အောက် layer များက low-level feature (edge, phoneme)၊ အထက် layer များက abstract concept (object, semantics) ကို ဖမ်းယူသည်။ Data, compute (GPU) နှင့် CNN, RNN, Transformer ကဲ့သို့ architecture များ၏ scale က ၎င်း၏ အောင်မြင်မှုကို တွန်းအားပေးသည်။ ### Overview Deep Learning ၏ 'deep' ဆိုသည်မှာ input နှင့် output ကြားရှိ hidden layers အရေအတွက် များပြားခြင်းကို ဆိုလိုသည်။ Shallow model များက feature engineering ကို လူကိုယ်တိုင် လုပ်ရသော်လည်း deep networks များက features များကို ကိုယ်တိုင်အလိုအလျောက် သင်ယူသည် — ၎င်းကို representation learning ဟုခေါ်သည်။ ### How It Works 1. Forward pass — input ကို layers များဖြတ်၍ nonlinear activations (ReLU စသည်) ဖြင့် အသွင်ပြောင်းကာ output ထုတ်ခြင်း 1. Loss computation — output နှင့် true label ကွာခြားမှုကို loss function ဖြင့် တိုင်းတာခြင်း 1. Backpropagation — chain rule ဖြင့် loss ၏ gradients များကို နောက်ပြန်ဖြန့်ခြင်း 1. Gradient descent — optimizer (Adam, SGD) က weights များကို update လုပ်ခြင်း 1. Repeat — epochs များစွာ iterate လုပ်၍ loss ကို လျှော့ချခြင်း ### Formulas - **Depth of a network**: `depth = number of hidden layers L + output layer` --- ## Entry № 014: Few-Shot & Zero-Shot Learning **Burmese Title**: ဖယူး-ရှော့ / ဇီးရိုး-ရှော့ သင်ယူမှု (Few-Shot & Zero-Shot Learning) **Level**: intermediate | **Topics**: generative-ai, nlp **URL**: https://aipedia.burmesestack.com/terms/few-shot-learning ### Definitions - **EL5 (Simple)**: Few-shot နဲ့ Zero-shot Learning ဆိုတာ AI ကို ဥပမာ အနည်းငယ် (few-shot) ဒါမှမဟုတ် လုံးဝ မပြဘဲ (zero-shot) အလုပ်အသစ်တစ်ခု လုပ်ခိုင်းတာ။ ဥပမာ - 'ဒီစာက positive လား negative လား' ဆိုတာ ဥပမာ ၂-၃ ခုပဲ ပြပြီး ကျန်တာ ဆက်ခွဲခိုင်းတာမျိုး — model အသစ် retrain မလုပ်ဘဲ prompt ထဲမှာပဲ သင်ပေးလိုက်တာပါ။ - **Standard (Burmese)**: Zero-shot learning သည် ဥပမာ မပြဘဲ instruction သက်သက်ဖြင့်၊ few-shot learning သည် prompt အတွင်း ဥပမာ အနည်းငယ် (demonstrations) ပြ၍ LLM အား task အသစ်တစ်ခုကို လုပ်ဆောင်စေခြင်းဖြစ်သည်။ Model ၏ weight ကို ပြောင်းလဲခြင်းမရှိဘဲ prompt အတွင်း context မှတစ်ဆင့် သင်ကြားခြင်းဖြစ်သောကြောင့် in-context learning ၏ ပုံစံများဖြစ်သည်။ - **Technical (Deep)**: Zero-shot သည် instruction သက်သက်ဖြင့် task လုပ်ဆောင်; few-shot သည် prompt ထဲ input–output ဥပမာ အနည်းငယ် ထည့်သည်။ နှစ်ခုစလုံး model weight ကို update မလုပ် — model က inference အချိန်တွင် demonstration များအပေါ် condition လုပ်သည် (in-context learning)။ Zero- မှ few-shot သို့နှင့် ဥပမာ ရွေးချယ်/စီစဉ်မှု ကောင်းလာသည်နှင့်အမျှ စွမ်းဆောင်ရည် တိုးသော်လည်း context length ဖြင့် ကန့်သတ်ခံရ၍ format အပေါ် အထိခိုက်လွယ်သည်။ ### Overview Few-shot နှင့် zero-shot prompting သည် LLM များကို fine-tuning မလိုဘဲ task အမျိုးမျိုးတွင် ချက်ချင်း အသုံးချနိုင်စေသည့် အဓိကနည်းလမ်းဖြစ်သည်။ GPT-3 paper ('Language Models are Few-Shot Learners') က ဤစွမ်းရည်ကို ကျယ်ပြန့်စွာ မိတ်ဆက်ပေးခဲ့သည်။ ### How It Works 1. Zero-shot — instruction သက်သက်ပေး၍ task လုပ်ခိုင်းသည် 1. Few-shot — prompt ထဲ input→output ဥပမာ အနည်းငယ် ထည့်သည် 1. Model က ဥပမာများ၏ pattern ကို context မှ 'သင်ယူ' သည် (weight မပြောင်း) 1. ဥပမာ ရွေးချယ်မှုနှင့် အစီအစဉ်က ရလဒ်ကို သက်ရောက်သည် --- ## Entry № 015: Activation Function **Burmese Title**: အသက်သွင်းလုပ်ဆောင်ချက် **Level**: intermediate | **Topics**: deep-learning, machine-learning **URL**: https://aipedia.burmesestack.com/terms/activation-function ### Definitions - **EL5 (Simple)**: Activation function ဆိုတာ neuron တစ်ခုရဲ့ 'ဖွင့်ဖို့ / ပိတ်ဖို့' ဆုံးဖြတ်ပေးတဲ့ ခလုတ်တစ်ခုလိုပဲ။ Input တွေရဲ့ ပေါင်းလဒ်က သတ်မှတ်ချက်ထက်ကျော်ရင် ဖွင့်ပြီး နောက်တစ်ဆင့်ကို အချက်ပြပို့တယ်။ ဒီခလုတ်က linear မဟုတ်တဲ့အတွက် network က ရှုပ်ထွေးတဲ့ pattern တွေကို သင်ယူနိုင်တယ်။ - **Standard (Burmese)**: Activation function သည် neural network သို့ non-linearity ကို ထည့်သွင်းပေးသော function ဖြစ်ပြီး network များကို linear မဟုတ်သော ဆက်စပ်မှုများ သင်ယူနိုင်စေသည်။ - **Technical (Deep)**: Non-linear activation များ မရှိပါက linear layer များ ဘယ်လောက်ပင် ထပ်ထားစေ single linear transformation တစ်ခုအဖြစ်သာ ကျရောက်သွားသည်။ အသုံးများသော choice များ — ReLU max(0, x) (sparse, မြန်)၊ Sigmoid 1/(1+e^{−x}) (output (0,1) အတွင်း)၊ Tanh, GELU, LeakyReLU — တစ်ခုစီတွင် gradient flow နှင့် saturation အတွက် အားသာ/အားနည်းချက် ရှိသည်။ ### Overview Activation functions များက neural network ၏ 'expressiveness' ကို ပေးသည်။ ၎င်းတို့မရှိပါက deep network သည် single linear layer နှင့် ညီမျှပြီး XOR ကဲ့သို့ simple patterns များကိုပင် မသင်ယူနိုင်ပါ။ Simulator တွင် activation ပြောင်းလိုက်တိုင်း output ပြောင်းပုံကို real-time ကြည့်နိုင်သည်။ ### How It Works 1. Neuron က weighted sum z = Σw·x + b ကို တွက်ချက်သည် 1. Activation function f ကို z ပေါ်တွင် အသုံးပြုသည်: a = f(z) 1. ReLU က positive input များကို ပို့ပြီး negative များကို 0 ပြုလုပ်သည် 1. Sigmoid က output ကို 0–1 ကြား ညှစ်ပေးပြီး probability ကဲ့သို့ အသုံးပြုနိုင်သည် 1. Backpropagation သည် gradient များကို activation ၏ derivative ဖြင့် flow လုပ်သည် ### Formulas - **ReLU**: `ReLU(x) = max(0, x); ReLU′(x) = 1 if x>0 else 0` - **Sigmoid**: `σ(x) = 1 / (1 + e^(−x)); σ′(x) = σ(x)·(1 − σ(x))` - **Tanh**: `tanh(x) = (e^(2x) − 1) / (e^(2x) + 1) ∈ (−1, 1)` --- ## Entry № 015: Support Vector Machine (SVM) **Burmese Title**: ပံ့ပိုးကွက် ဗက်တာ စက် (Support Vector Machine) **Level**: intermediate | **Topics**: machine-learning **URL**: https://aipedia.burmesestack.com/terms/support-vector-machine ### Definitions - **EL5 (Simple)**: Support Vector Machine (SVM) ဆိုတာ အုပ်စုနှစ်ခုကြားမှာ နှစ်ဖက်လုံးနဲ့ အကွာအဝေး အကျယ်ဆုံးဖြစ်တဲ့ နယ်ခြားမျဉ်း ဆွဲပြီး ခွဲခြားတဲ့ နည်း။ ဥပမာ - စားပွဲပေါ်က အနီနဲ့ အပြာ ဂေါ်လီတွေကြားမှာ နှစ်ဖက်စလုံးနဲ့ အဝေးဆုံးဖြစ်အောင် မျဉ်းတစ်ကြောင်း ဆွဲပေးတာမျိုးပါ။ - **Standard (Burmese)**: SVM သည် class နှစ်ခုကြား margin (အကွာအဝေး) အကျယ်ဆုံးဖြစ်စေသည့် decision boundary (hyperplane) ကို ရှာသော supervised algorithm ဖြစ်သည်။ Boundary အနီးဆုံး point များ (support vectors) သာ boundary ကို သတ်မှတ်သည်။ Kernel trick ဖြင့် linear မဟုတ်သော data များကိုပါ ခွဲခြားနိုင်သည်။ - **Technical (Deep)**: SVM သည် class များ ခွဲသည့် maximum-margin hyperplane ကို ရှာ၍ support vector များဖြင့်သာ သတ်မှတ်သည်။ Soft margin (regularization C) က generalization ပိုကောင်းစေရန် misclassification အနည်းငယ် ခွင့်ပြုသည်။ Kernel trick (RBF, polynomial) က input များကို higher-dimensional space သို့ implicit map လုပ်၍ mapping ကို တိုက်ရိုက် မတွက်ဘဲ non-linear boundary ကို ကိုင်တွယ်သည်။ SVM များသည် high dimension တွင် ထိရောက်သော်လည်း dataset အလွန်ကြီးများတွင် scale ညံ့သည်။ ### Overview SVM သည် ML ခေတ်ဦးက အလွန်ဩဇာကြီးခဲ့သော algorithm ဖြစ်ပြီး margin maximization ဆိုသည့် သီအိုရီ အခိုင်အမာရှိသည်။ High-dimensional data (ဥပမာ text) များနှင့် dataset သေးငယ်သည့်အခါ ယနေ့တိုင် အသုံးဝင်ဆဲဖြစ်သည်။ ### How It Works 1. Class များကို ခွဲသည့် hyperplane များစွာထဲမှ margin အကျယ်ဆုံးကို ရွေးသည် 1. Boundary အနီးဆုံး point (support vectors) များသာ အရေးပါသည် 1. Soft margin (C) ဖြင့် အမှားအနည်းငယ်ကို ခွင့်ပြု၍ generalization တိုးသည် 1. Kernel trick ဖြင့် non-linear data ကို ခွဲခြားသည် --- ## Entry № 015: Context Window **Burmese Title**: ကွန်တက်စ် ဝင်းဒိုး (Context Window / အကြောင်းအရာ နယ်ပယ်) **Level**: intermediate | **Topics**: generative-ai, nlp **URL**: https://aipedia.burmesestack.com/terms/context-window ### Definitions - **EL5 (Simple)**: Context Window ဆိုတာ LLM တစ်ခုက တစ်ကြိမ်တည်းမှာ 'ဖတ်/မှတ်' နိုင်တဲ့ စာ ပမာဏ အများဆုံး။ token (စကားလုံးအပိုင်းအစ) အရေအတွက်နဲ့ တိုင်းတာတယ်။ Window သေးရင် စာရှည်ကြီးရဲ့ အစပိုင်းကို 'မေ့' သွားတတ်ပြီး၊ ကြီးရင် စာအုပ်တစ်အုပ်လုံးကိုတောင် တစ်ခါတည်း ဖတ်ပြီး ဖြေနိုင်တယ်။ - **Standard (Burmese)**: Context Window သည် LLM က တစ်ကြိမ်လျှင် process လုပ်နိုင်သော token (input + output) ၏ အများဆုံး အရေအတွက်ဖြစ်သည်။ Prompt, conversation history, retrieved documents အားလုံး ဤ window အတွင်း အံဝင်ရမည်ဖြစ်ပြီး ကျော်လွန်ပါက အဟောင်းဆုံး အချက်များ ကျန်ခဲ့/ဖြတ်တောက်ခံရသည်။ Window ကြီးလေ ဆက်စပ်အချက်အလက် ပိုသိုလှောင်နိုင်လေဖြစ်သည်။ - **Technical (Deep)**: Context window သည် model က forward pass တစ်ကြိမ်တွင် attend လုပ်နိုင်သော token (prompt + generated output) အများဆုံး အရေအတွက်ဖြစ်သည်။ ၎င်းက instruction, history, retrieved context မည်မျှ တစ်ပြိုင်နက် အံဝင်နိုင်မည်ကို ကန့်သတ်သည်; ကျော်လွန်ပါက truncation သို့မဟုတ် summarization လုပ်ရသည်။ Window ကြီးလျှင် long-document/long-conversation ကို ဖြစ်နိုင်စေသော်လည်း compute/memory ကုန်ကျ (attention သည် sequence length အလိုက် တိုးသည်) ပြီး 'lost in the middle' ယိုယွင်းမှု ဖြစ်နိုင်သည်။ ### Overview Context Window သည် LLM application ဒီဇိုင်း၏ အခြေခံ ကန့်သတ်ချက်တစ်ခုဖြစ်သည် — RAG, chunking, memory strategy များ လိုအပ်ရခြင်း၏ အဓိက အကြောင်းရင်းဖြစ်သည်။ Model အသစ်များတွင် window သည် များစွာ ကြီးလာသော်လည်း token ကုန်ကျစရိတ်နှင့် 'lost in the middle' ပြဿနာက ရှိနေဆဲဖြစ်သည်။ ### How It Works 1. Token = စကားလုံး/စကားလုံးအပိုင်းအစ; window ကို token အရေအတွက်ဖြင့် တိုင်းသည် 1. Prompt + history + retrieved context + output အားလုံး window အတွင်း အံဝင်ရသည် 1. ကျော်လွန်ပါက အဟောင်းဆုံးအပိုင်းများ ဖြတ်တောက်/အနှစ်ချုပ်ခံရသည် 1. Window ကြီးလေ ဆက်စပ်အချက် ပိုသိုလှောင်နိုင်လေ (သို့သော် compute ပိုကုန်) --- ## Entry № 016: Backpropagation **Burmese Title**: နောက်ပြန်ဖြန့်ခြင်း **Level**: intermediate | **Topics**: machine-learning, deep-learning **URL**: https://aipedia.burmesestack.com/terms/backpropagation ### Definitions - **EL5 (Simple)**: Backpropagation ဆိုတာ neural network ရဲ့ error ကို နောက်ဆုံးအလွှာကနေ ရှေ့ဆုံးအလွှာအထိ ပြန်ပို့ပြီး 'ဒီ weight က ဘယ်လောက် မှားခဲ့လဲ' ဆိုတာ တွက်တဲ့ နည်းလမ်း။ ရလဒ်ကို စာမေးပွဲဖြေပြီးမှ အမှတ်ကြည့်ပြီး ဘယ်မေးခွန်းက ဘယ်မှားလဲ ပြန်စစ်တာနဲ့ တူတယ်။ - **Standard (Burmese)**: Backpropagation သည် neural network ရှိ weights များအတွက် loss ၏ gradients များကို chain rule ဖြင့် တွက်ချက်သော algorithm ဖြစ်သည်။ Error ကို output layer မှ input layer သို့ နောက်ပြန်ဖြန့်ကာ layer တစ်ခုချင်းစီ၏ parameter updates များကို gradient descent ဖြင့် လုပ်ဆောင်နိုင်စေသည်။ - **Technical (Deep)**: Backpropagation သည် chain rule ဖြင့် layer l တိုင်းအတွက် ∂L/∂W(l) ကို တွက်သည်။ Error signal δ(l) သည် နောက်ပြန် ဖြန့်သွားသည် — δ(l) = (W(l+1))ᵀ δ(l+1) ⊙ σ'(z(l))၊ gradient မှာ ∂L/∂W(l) = δ(l) (a(l−1))ᵀ ဖြစ်သည်။ ၎င်းသည် network ၏ computational graph ပေါ် အသုံးပြုသော reverse-mode automatic differentiation ၏ special case တစ်ခုဖြစ်သည်။ ### Overview Backpropagation သည် forward pass တွင် တွက်ချက်ထားသော activations များကို သုံး၍ loss ၏ gradients များကို output layer မှ input layer သို့ ဆင့်ကဲတွက်ချက်သည်။ ၎င်းသည် gradient descent ၏ အုတ်မြစ်ဖြစ်ပြီး deep networks များ လေ့ကျင့်နိုင်စေသော အဓိက algorithm ဖြစ်သည်။ ### How It Works 1. Forward pass — activations a(l) များကို layer များဖြတ်၍ တွက်ချက်ပြီး loss L ရယူခြင်း 1. Output error — δ(L) = ∇a L ⊙ σ'(z(L)) ကို output layer တွင် တွက်ချက်ခြင်း 1. Backward recursion — δ(l) = (W(l+1))T δ(l+1) ⊙ σ'(z(l)) ဖြင့် error signal ကို layer များတစ်လျှောက် နောက်ပြန်ဖြန့်ခြင်း 1. Gradient — ∂L/∂W(l) = δ(l) (a(l−1))T ကို layer တစ်ခုစီအတွက် တွက်ခြင်း 1. Update — optimizer (SGD/Adam) က weights များကို gradient ၏ ဆန့်ကျင်ဘက်ဦးတည်ချက်ဖြင့် update လုပ်ခြင်း ### Formulas - **Backpropagation recursion**: `δ(l) = (W(l+1))T δ(l+1) ⊙ σ'(z(l)); ∂L/∂W(l) = δ(l) (a(l−1))T` - **Chain rule**: `∂L/∂W(l) = (∂L/∂a(l)) · (∂a(l)/∂z(l)) · (∂z(l)/∂W(l))` --- ## Entry № 016: Ensemble Learning **Burmese Title**: အစုအဖွဲ့ သင်ယူမှု (Ensemble Learning) **Level**: intermediate | **Topics**: machine-learning **URL**: https://aipedia.burmesestack.com/terms/ensemble-learning ### Definitions - **EL5 (Simple)**: Ensemble Learning ဆိုတာ model တစ်ခုတည်းကို အားကိုးမနေဘဲ model များစွာ၏ အဖြေတွေ ပေါင်းစပ်ပြီး ပိုကောင်းတဲ့ အဖြေ ထုတ်တဲ့ နည်း။ ကျွမ်းကျင်သူ တစ်ဦးတည်းထက် အဖွဲ့လိုက် တိုင်ပင်ဆုံးဖြတ်ရင် အမှားနည်းသလိုပါ။ - **Standard (Burmese)**: Ensemble Learning သည် model (learner) များစွာ၏ ခန့်မှန်းချက်များကို ပေါင်းစပ်ခြင်းဖြင့် တစ်ခုတည်းထက် ပိုတိကျ၊ ပိုတည်ငြိမ်သော ရလဒ်ရရှိစေသည့် နည်းလမ်းဖြစ်သည်။ အဓိကနည်းသုံးမျိုးမှာ Bagging (parallel — variance လျော့), Boosting (sequential — bias လျော့) နှင့် Stacking (model များ၏ output ကို meta-model ဖြင့် ပေါင်း) တို့ဖြစ်သည်။ - **Technical (Deep)**: Ensemble method များသည် base learner များစွာကို ပေါင်းစပ်၍ တစ်ခုတည်းထက် error လျှော့ချသည်။ Bagging (ဥပမာ Random Forest) က learner များကို bootstrap sample ပေါ် parallel လေ့ကျင့်၍ ပျမ်းမျှကာ variance လျှော့; boosting (ဥပမာ gradient boosting) က sequential လေ့ကျင့်၍ residual error ပြင်ကာ bias လျှော့; stacking က base-model output များပေါ် meta-model သင်ယူသည်။ အကျိုးရရှိရန် base learner များ တိကျ၍ ကွဲပြား (decorrelated error) ရမည်။ ### Overview Ensemble Learning သည် modern applied ML ၏ အောင်မြင်မှု၏ နောက်ကွယ်တွင် ရှိသည် — Random Forest, XGBoost အစရှိသည်တို့သည် ensemble များဖြစ်သည်။ Base learner များ တိကျ၍ ကွဲပြား (diverse) လေ ensemble ၏ အကျိုးကျေးဇူး များလေဖြစ်သည်။ ### How It Works 1. Base learner (model) များစွာကို လေ့ကျင့်သည် 1. Bagging — parallel လေ့ကျင့်၍ ပျမ်းမျှ/vote (variance လျော့) 1. Boosting — sequential လေ့ကျင့်၍ error ပြင် (bias လျော့) 1. Stacking — base output များကို meta-model ဖြင့် ပေါင်းစပ်သည် --- ## Entry № 016: Pooling **Burmese Title**: ပူးလင်း (Pooling / Feature ချုံ့ခြင်း) **Level**: intermediate | **Topics**: computer-vision, deep-learning **URL**: https://aipedia.burmesestack.com/terms/pooling ### Definitions - **EL5 (Simple)**: Pooling ဆိုတာ CNN ထဲမှာ ဓာတ်ပုံ feature တွေကို 'ချုံ့' ပေးတဲ့ အဆင့်။ အနီးအနားက pixel အုပ်စုလေးတစ်ခုကနေ အရေးအကြီးဆုံး တန်ဖိုး (ဥပမာ အမြင့်ဆုံး) ကိုသာ ယူလိုက်တာ။ ဒါကြောင့် ပုံအရွယ် သေးသွားပြီး တွက်ရ ပိုမြန်၊ အရေးကြီးတဲ့ အချက်တွေ ကျန်နေတယ်။ - **Standard (Burmese)**: Pooling သည် CNN တွင် feature map ၏ spatial dimension (အကျယ် × အမြင့်) ကို လျှော့ချပေးသည့် downsampling operation ဖြစ်သည်။ Max pooling က window တစ်ခုအတွင်း အမြင့်ဆုံးတန်ဖိုးကို၊ average pooling က ပျမ်းမျှကို ယူသည်။ ၎င်းက computation ကို လျှော့ချ၊ receptive field ကို ကျယ်စေ၊ translation invariance ကို တိုးစေသည်။ - **Technical (Deep)**: Pooling သည် local window (ဥပမာ 2×2, stride 2) များပေါ် aggregate လုပ်၍ feature map ကို downsample လုပ်သည်။ Max pooling က အားအကောင်းဆုံး activation ကို ထိန်း; average pooling က mean ကို ယူသည်။ ၎င်းက spatial resolution နှင့် computation လျှော့၊ receptive field ကျယ်စေ၊ translation invariance အနည်းငယ် ပေးသည်။ Global average pooling က channel တစ်ခုစီကို တန်ဖိုးတစ်ခုအဖြစ် ချုံ့၍ classifier ရှေ့ dense layer များကို အစားထိုးလေ့ရှိသည်။ ### Overview Pooling သည် CNN ၏ convolution–activation–pooling pipeline ၏ အစိတ်အပိုင်းတစ်ခုဖြစ်ပြီး feature map ကို ဆင့်ကဲ ချုံ့ရင်း အရေးပါသော pattern များကို ထိန်းသိမ်းသည်။ Modern architecture အချို့တွင် pooling အစား strided convolution ကို သုံးလာသည်။ ### How It Works 1. Feature map ကို window (ဥပမာ 2×2) များအဖြစ် ပိုင်းသည် 1. Max pooling — window တစ်ခုစီမှ အမြင့်ဆုံးတန်ဖိုးကို ယူသည် 1. Average pooling — ပျမ်းမျှတန်ဖိုးကို ယူသည် 1. Spatial size သေးသွား၍ computation နှင့် parameter လျော့သည် --- ## Entry № 016: Temperature & Sampling **Burmese Title**: တမ်ပရေချာ နှင့် Sampling (Temperature & Sampling) **Level**: intermediate | **Topics**: generative-ai **URL**: https://aipedia.burmesestack.com/terms/temperature-sampling ### Definitions - **EL5 (Simple)**: Temperature နဲ့ Sampling ဆိုတာ LLM ရဲ့ အဖြေ ဘယ်လောက် 'တီထွင်ဆန်း'/'တည်ငြိမ်' ဖြစ်မလဲ ချိန်ပေးတဲ့ ခလုတ်တွေ။ Temperature နိမ့်ရင် အမြဲတမ်း အလုံခြုံဆုံး၊ တူညီတဲ့ အဖြေ ထုတ်တယ်။ မြင့်ရင် ပိုဆန်းသစ်၊ မထင်မှတ်တဲ့ အဖြေ ထွက်လာတယ် (ဒါပေမယ့် အမှားလည်း များနိုင်)။ ကဗျာဆိုရင် မြင့်၊ သင်္ချာဆိုရင် နိမ့် ထားသင့်တယ်။ - **Standard (Burmese)**: Temperature သည် LLM ၏ နောက်စကားလုံး probability distribution ကို ချောမွေ့/ချွန်ထက်စေသည့် parameter ဖြစ်သည် — နိမ့်လျှင် deterministic/တည်ငြိမ်၊ မြင့်လျှင် ကွဲပြား/တီထွင်ဆန်းသည်။ Top-k နှင့် Top-p (nucleus) sampling တို့သည် ရွေးချယ်ရာတွင် ဖြစ်နိုင်ခြေ အမြင့်ဆုံး token များကိုသာ ကန့်သတ်၍ အရည်အသွေးနှင့် ကွဲပြားမှုကို ချိန်ညှိပေးသည်။ - **Technical (Deep)**: Sampling သည် model ၏ probability distribution မှ နောက် token ကို ဘယ်လို ဆွဲထုတ်မည်ကို ထိန်းချုပ်သည်။ Temperature T က logit များကို rescale လုပ်သည် (softmax(logits/T)) — T→0 သည် greedy/deterministic သို့ ချဉ်းကပ်၊ T မြင့်လျှင် distribution ကို ချောမွေ့စေ၍ diversity တိုးသည်။ Top-k က ဖြစ်နိုင်ခြေအမြင့်ဆုံး token k ခုသို့ ကန့်သတ်; top-p (nucleus) က cumulative probability ≥ p ဖြစ်သည့် အနည်းဆုံး set သို့။ ၎င်းတို့သည် coherence/factuality နှင့် creativity/variety ကို အလဲအလှယ်လုပ်သည်။ ### Overview Temperature နှင့် sampling parameter များသည် LLM output ၏ ကွဲပြားမှုနှင့် တည်ငြိမ်မှုကို ထိန်းချုပ်ရာတွင် အခြေခံကျသည်။ Factual/code task များတွင် temperature နိမ့် (တိကျ)၊ creative writing တွင် မြင့် (ဆန်းသစ်) ထားခြင်းက အကောင်းဆုံး ရလဒ်ပေးသည်။ ### How It Works 1. Model က token တစ်ခုစီအတွက် probability ထုတ်သည် 1. Temperature — နိမ့်=deterministic/တည်ငြိမ်, မြင့်=ကွဲပြား/တီထွင်ဆန်း 1. Top-k — ဖြစ်နိုင်ခြေ အမြင့်ဆုံး k ခုထဲမှသာ ရွေးသည် 1. Top-p (nucleus) — cumulative probability ≥ p ဖြစ်သည့် အနည်းဆုံး set ထဲမှ ရွေးသည် ### Formulas - **Temperature-scaled softmax**: `P(token) = softmax(logits / T)` --- ## Entry № 017: Attention Mechanism **Burmese Title**: အာရုံစိုက်မှု ယန္တရား **Level**: intermediate | **Topics**: deep-learning, nlp **URL**: https://aipedia.burmesestack.com/terms/attention-mechanism ### Definitions - **EL5 (Simple)**: Attention ဆိုတာ model က စာကြောင်းထဲက စကားလုံးတစ်လုံးကို ကြည့်တဲ့အခါ ဘယ်စကားလုံးတွေကို ပိုအရေးကြီးတယ်ဆိုတာ ဆုံးဖြတ်တဲ့ နည်း။ 'ဘဏ်ကို သွားတယ်' ဆိုရင် 'ဘဏ်' က ငွေဘဏ်လား မြစ်ကမ်းလား ဆိုတာကို ဘေးပတ်ဝန်းကျင် စကားလုံးတွေကနေ ချိန်ဆကြည့်သလိုပဲ။ - **Standard (Burmese)**: Attention mechanism သည် model များကို input sequence ၏ သက်ဆိုင်ရာ အစိတ်အပိုင်းများပေါ်တွင် အလေးပေးအာရုံစိုက်နိုင်စေပြီး position မခွဲခြားဘဲ dependencies များကို ဖမ်းယူနိုင်စေသည်။ - **Technical (Deep)**: Scaled dot-product attention သည် Attention(Q,K,V) = softmax(QKᵀ/√dₖ)V ကို တွက်ချက်သည် — Q/K/V များသည် token representation များ၏ linear projection များဖြစ်သည်။ Multi-head attention က ၎င်းကို subspace များတွင် တစ်ပြိုင်နက် လုပ်ဆောင်စေ၍ head တစ်ခုစီက မတူညီသော linguistic relationship များအပေါ် အထူးပြုသည်။ ### Overview Attention က sequence-to-sequence နှင့် transformer architectures များ၏ အခြေခံဖြစ်သည်။ RNN များနှင့် မတူဘဲ sequence ရှည်ရှည်များတွင် မှေးမှိန်ခြင်း (vanishing) မရှိဘဲ any two positions ကြား direct dependency ကို capture လုပ်နိုင်သည်။ Attention heatmap simulator တွင် weight များ ဖြန့်ဝေပုံကို ကြည့်နိုင်သည်။ ### How It Works 1. Token တစ်ခုစီမှ Query, Key, Value ဟူ၍ vector သုံးမျိုး ထုတ်ယူသည် 1. Similarity score = Query · Key (scaled) — token နှစ်ခုမည်မျှ ဆက်စပ်သည်ကို တိုင်းသည် 1. Softmax ဖြင့် scores များကို probability (weights) အဖြစ်ပြောင်းသည် 1. Output = weighted sum of Values — အရေးကြီးသော tokens များက ပိုလွှမ်းမိုးသည် 1. Multi-head ဖြင့် မတူညီသော relation အမျိုးမျိုးကို ရှာဖွေသည် ### Formulas - **Scaled dot-product attention**: `Attention(Q, K, V) = softmax(Q·Kᵀ / √dₖ)·V` --- ## Entry № 017: Random Forest **Burmese Title**: ကျပန်း သစ်တော (Random Forest) **Level**: intermediate | **Topics**: machine-learning **URL**: https://aipedia.burmesestack.com/terms/random-forest ### Definitions - **EL5 (Simple)**: Random Forest ဆိုတာ Decision Tree အများကြီးကို တွဲသုံးပြီး 'အများဆန္ဒ' နဲ့ ဆုံးဖြတ်တဲ့ နည်း။ သစ်ပင်တစ်ပင်တည်း မေးရင် မှားနိုင်ပေမယ့် သစ်ပင် ရာနဲ့ချီ မေးပြီး အများစုအဖြေကို ယူတာမို့ ပိုတိကျပါတယ်။ ကျွမ်းကျင်သူ များများကို မေးပြီး အများသဘောတူတဲ့ အဖြေ ယူသလိုပါ။ - **Standard (Burmese)**: Random Forest သည် decision tree များစွာကို data နှင့် feature ၏ ကျပန်း subset များပေါ်တွင် သီးခြားစီ လေ့ကျင့်ပြီး ၎င်းတို့၏ ခန့်မှန်းချက်များကို ပေါင်းစပ် (classification — အများစု vote; regression — ပျမ်းမျှ) သည့် ensemble algorithm ဖြစ်သည်။ Tree များ၏ ကွဲပြားမှုက overfitting ကို လျှော့ချ၍ တစ်ပင်တည်းထက် ပိုတည်ငြိမ်စေသည်။ - **Technical (Deep)**: Random Forest သည် decision tree များ၏ bagging ensemble ဖြစ်သည်။ Tree တစ်ပင်စီကို bootstrap sample ပေါ် လေ့ကျင့်ပြီး split တိုင်းတွင် feature subset ကျပန်း (feature bagging) စဉ်းစားခြင်းက tree များကို decorrelate လုပ်သည်။ Prediction များကို majority vote သို့ averaging ဖြင့် ပေါင်းသည်။ ၎င်းက single deep tree နှင့်ယှဉ်လျှင် variance လျှော့၍ feature-importance ခန့်မှန်းချက်နှင့် out-of-bag error ပေးသည် — interpretability နှင့် model size ကြီးခြင်းကို အလဲအလှယ်လုပ်၍။ ### Overview Random Forest သည် decision tree များ၏ variance ကို bagging ဖြင့် လျှော့ချပြီး tabular data တွင် အားကောင်း၊ ချိန်ညှိရ လွယ်ကူသော algorithm ဖြစ်သည်။ Feature importance ကို ဖော်ပြနိုင်ခြင်းက ၎င်း၏ အားသာချက်တစ်ခုဖြစ်သည်။ ### How It Works 1. Training data မှ bootstrap sample များ ကျပန်းယူသည် 1. Tree တစ်ပင်စီကို split တိုင်းတွင် feature subset ကျပန်းဖြင့် ကြီးထွားစေသည် 1. Tree များ၏ ခန့်မှန်းချက်ကို vote / average ဖြင့် ပေါင်းစပ်သည် 1. Tree များ ကွဲပြားခြင်းက variance ကို လျှော့ချသည် --- ## Entry № 017: Optical Character Recognition (OCR) **Burmese Title**: အက္ခရာ အလိုအလျောက် ဖတ်ရှုခြင်း (Optical Character Recognition) **Level**: intermediate | **Topics**: computer-vision, nlp **URL**: https://aipedia.burmesestack.com/terms/ocr ### Definitions - **EL5 (Simple)**: OCR ဆိုတာ ဓာတ်ပုံ ဒါမှမဟုတ် စကင်န်ဖတ်ထားတဲ့ စာရွက်ထဲက စာသားကို ကွန်ပျူတာ ဖတ်လို့ရတဲ့ digital စာသားအဖြစ် ပြောင်းပေးတဲ့ နည်း။ ဥပမာ - စာအုပ်စာမျက်နှာ ဓာတ်ပုံရိုက်ပြီး အဲဒီထဲက စာတွေကို copy-paste လုပ်လို့ရအောင် ပြောင်းပေးတာမျိုးပါ။ - **Standard (Burmese)**: OCR (Optical Character Recognition) သည် ဓာတ်ပုံ သို့မဟုတ် scan လုပ်ထားသော document ထဲရှိ ရေးသား/ရိုက်နှိပ်ထားသော စာသားများကို machine-readable digital text အဖြစ် ပြောင်းလဲပေးသည့် computer vision နည်းပညာဖြစ်သည်။ Text detection (စာ ဘယ်နေရာ) နှင့် text recognition (ဘာစာ) ဟူသော အဆင့်နှစ်ဆင့် ပါဝင်သည်။ - **Technical (Deep)**: OCR သည် ဓာတ်ပုံရှိ စာသားကို machine-readable character အဖြစ် ပြောင်းသည် — များသောအားဖြင့် detection (text region/line ရှာ) ပြီးနောက် recognition (character decode, CNN+RNN+CTC သို့ Transformer decoder ဖြင့်) ဟူ၍ ဖြစ်သည်။ Modern system များက scene text, လက်ရေးနှင့် multilingual script များကို ကိုင်တွယ်နိုင်သည်။ ၎င်းသည် computer vision နှင့် NLP ကို ချိတ်ဆက်၍ document digitization, translation app, automated data entry တို့ကို ဖြစ်စေသည်။ ### Overview OCR သည် computer vision နှင့် NLP ကို ချိတ်ဆက်ပေးသည့် နည်းပညာဖြစ်ပြီး document digitization, license plate reading, receipt scanning, ဘာသာပြန် app များတွင် အသုံးဝင်သည်။ ရိုးရိုး ရိုက်နှိပ်စာမှ လက်ရေးနှင့် မြန်မာစာ ကဲ့သို့ ရှုပ်ထွေးသော script များအထိ တိုးတက်လာသည်။ ### How It Works 1. Text detection — ဓာတ်ပုံထဲ စာသား ရှိရာ နေရာ (line / word) ကို ရှာသည် 1. Preprocessing — deskew, denoise, binarize လုပ်၍ ရှင်းလင်းစေသည် 1. Text recognition — စာလုံးများကို decode လုပ်သည် (CNN+RNN+CTC / Transformer) 1. Digital text အဖြစ် ထုတ်ပေးသည် --- ## Entry № 018: Convolutional Neural Network (CNN) **Burmese Title**: လိမ်ယှက်အာရုံကြောကွန်ယက် **Level**: intermediate | **Topics**: deep-learning, computer-vision **URL**: https://aipedia.burmesestack.com/terms/cnn ### Definitions - **EL5 (Simple)**: Convolutional Neural Network ဆိုတာ ရုပ်ပုံတွေကို ကြည့်ဖို့ အထူးဆောက်ထားတဲ့ neural network။ သူက ရုပ်ပုံကို ကြည့်တဲ့အခါ နေရာတိုင်းကို filter လေးတွေနဲ့ လျှောက်စကင်န်ဖတ်တယ်။ ဒါကြောင့် ခွေးတစ်ကောင်က ဘယ်ဘက်မှာရှိ၊ ညာဘက်မှာရှိ ဘာနေရာပဲဖြစ်ဖြစ် မှတ်မိနိုင်တယ် — ပြောင်းလိုက်တဲ့နေရာမှာ အထိခိုက်မရှိဘူး။ - **Standard (Burmese)**: Convolutional Neural Network (CNN) သည် images များကို လုပ်ဆောင်ရန် ဒီဇိုင်းထုတ်ထားသော deep learning architecture ဖြစ်ပြီး convolution layers, activation (ReLU), pooling layers နှင့် fully-connected layers များဖြင့် ဖွဲ့စည်းထားသည်။ Convolution များက local spatial features (edges, textures) များကို သင်ယူကာ feature map ၏ dimension များကို လျှော့ချပေးသည်။ - **Technical (Deep)**: CNN သည် input feature map များပေါ် လျှောက်သွားသော convolutional filter များကို ထပ်ဆင့်တည်ဆောက်သည် — output(l) = σ(W(l) ∗ input(l) + b(l))၊ ∗ သည် convolution operator ဖြစ်သည်။ Pooling (max/avg) downsampling က local translation invariance ကို ပေး၍ computation လျှော့သည်; နောက်ပိုင်း layer များက higher-level feature များကို ဖမ်းယူ၊ fully-connected layer နောက်ဆုံးတွင် class score ထုတ်သည်။ Conv2d filter များသည် spatial position များတစ်လျှောက် weight မျှဝေသုံးသောကြောင့် dense layer နှင့်ယှဉ်လျှင် parameter များစွာ လျော့သည်။ ### Overview CNN များသည် images ကဲ့သို့ grid-structured data တွင် ထူးချွန်သည်။ Convolution က သေးငယ်သော kernel ဖြင့် input ကို လျှောက်ဖြတ်၍ feature maps များထုတ်ကာ၊ pooling က resolution ကို လျှော့ချပြီး နောက်ဆုံးတွင် fully-connected layers များက classification လုပ်သည်။ ### How It Works 1. Convolution — kernel (filter) က input image ပေါ်တွင် လျှောက်၍ element-wise product များကို ပေါင်းပြီး feature map ထုတ်ခြင်း 1. ReLU activation — negative values များကို ဖယ်ကာ non-linearity ထည့်ခြင်း 1. Pooling — max/avg pooling ဖြင့် spatial size လျှော့ချပြီး နေရာပြောင်းလဲမှုကို ခံနိုင်ရည်ရှိစေခြင်း 1. Flatten + fully-connected — feature map များကို vector အဖြစ်ပြောင်းကာ class scores ထုတ်ခြင်း 1. Training — backpropagation + gradient descent ဖြင့် kernels များကို သင်ယူခြင်း ### Formulas - **2D convolution**: `y[i,j] = Σ_m Σ_n x[i+m, j+n] · k[m,n] + b` - **Small kernel example (3×3)**: `Edge filter k = [[-1,-1,-1],[0,0,0],[1,1,1]] applied at pixel (i,j) highlights horizontal edges` - **Max pooling**: `y[i,j] = max{ x[2i:2i+2, 2j:2j+2] }` --- ## Entry № 018: Gradient Boosting (XGBoost) **Burmese Title**: ဂရေဒီယင့် ဘွတ်စတင်း (Gradient Boosting) **Level**: intermediate | **Topics**: machine-learning **URL**: https://aipedia.burmesestack.com/terms/gradient-boosting ### Definitions - **EL5 (Simple)**: Gradient Boosting ဆိုတာ Decision Tree သေးသေးလေးတွေကို တစ်ပင်ပြီးတစ်ပင် ထပ်ဆောက်ပြီး၊ တစ်ပင်စီက အရင်တစ်ပင်ရဲ့ အမှားကို ဖာထေးပြင်ဆင်ပေးသွားတဲ့ နည်း။ အမှားကို တစ်ဆင့်ချင်း ပြင်သွားတာမို့ နောက်ဆုံးမှာ အလွန်တိကျတဲ့ model ဖြစ်လာတယ်။ - **Standard (Burmese)**: Gradient Boosting သည် အားနည်းသေးငယ်သော tree များကို အဆင့်ဆင့် (sequential) ထပ်ဆောက်ပြီး tree အသစ်တစ်ပင်စီက ယခင် model ၏ error (residual / gradient) ကို ပြင်ဆင်ပေးသည့် ensemble နည်းလမ်းဖြစ်သည်။ XGBoost, LightGBM, CatBoost တို့သည် ၎င်း၏ အမြန်ဆုံးနှင့် အသုံးအများဆုံး implementation များဖြစ်သည်။ - **Technical (Deep)**: Gradient boosting သည် additive model ကို အဆင့်ဆင့် (stage-wise) တည်ဆောက်သည် — weak learner အသစ်တစ်ခုစီ (များသောအားဖြင့် shallow tree) ကို လက်ရှိ prediction အပေါ် loss ၏ negative gradient နှင့် fit လုပ်ပြီး learning-rate shrinkage ဖြင့် ပေါင်းထည့်သည်။ Regularization, subsampling နှင့် second-order optimization (XGBoost) က overfitting ကို ထိန်းသည်။ Tabular data benchmark များတွင် ဦးဆောင်လေ့ရှိသော်လည်း Random Forest ထက် tuning ပိုအထိခိုက်လွယ်သည်။ ### Overview Gradient Boosting သည် tabular data ပြိုင်ပွဲ (Kaggle) များ၏ အနိုင်ရ algorithm ဖြစ်လေ့ရှိသည်။ Random Forest က tree များကို တစ်ပြိုင်နက် (parallel, bagging) ဆောက်သော်လည်း boosting က အဆင့်ဆင့် (sequential) ဆောက်၍ error ကို ဖာထေးသွားခြင်း ကွာခြားသည်။ ### How It Works 1. ပထမ model ၏ ခန့်မှန်းချက်နှင့် အမှား (residual) ကို တွက်သည် 1. Tree အသစ်တစ်ပင်ကို အဲဒီ error (negative gradient) ကို ခန့်မှန်းရန် လေ့ကျင့်သည် 1. Learning rate ဖြင့် တစ်ပင်စီ၏ contribution ကို ချုံ့၍ ပေါင်းထည့်သည် 1. Error ငယ်သွားသည်အထိ ထပ်ခါ ထပ်ဆောက်သည် --- ## Entry № 019: Recurrent Neural Network (RNN) **Burmese Title**: ပြန်လည်လည်ပတ်အာရုံကြောကွန်ယက် **Level**: intermediate | **Topics**: deep-learning, nlp **URL**: https://aipedia.burmesestack.com/terms/rnn ### Definitions - **EL5 (Simple)**: Recurrent Neural Network ဆိုတာ စာကြောင်း၊ သီချင်း၊ အချိန်နဲ့အလိုက် အချက်အလက်တွေလို တစ်ခုပြီးတစ်ခု အစဉ်လိုက်ပါတဲ့ data တွေကို ကိုင်တွယ်ဖို့ ဆောက်ထားတဲ့ neural network။ သူမှာ memory လေးရှိပြီး အရင်ဖတ်ခဲ့တာတွေကို မှတ်မိပြီး နောက်စာလုံးကို ခန့်မှန်းတယ်။ ဥပမာ 'နေ့လည်' ဆိုတဲ့စကားလုံးပြီးရင် 'ခင်းကျင်းမှု' လို့ ခန့်မှန်းနိုင်တယ်။ - **Standard (Burmese)**: Recurrent Neural Network (RNN) သည် sequential data (text, time series, speech) များကို လုပ်ဆောင်ရန် ဒီဇိုင်းထုတ်ထားသော neural network အမျိုးအစားဖြစ်ပြီး hidden state မှတစ်ဆင့် အချိန်တစ်ဆင့်ချင်း၏ information များကို သယ်ဆောင်သည်။ ၎င်းသည် sequence တစ်ခုလုံး၏ context ကို ထည့်သွင်းစဉ်းစားနိုင်သည်။ - **Technical (Deep)**: RNN သည် hidden state ကို ထိန်းသိမ်းရင်း sequence ကို အဆင့်လိုက် process လုပ်သည်။ Recurrence မှာ h_t = f(W_h h_{t−1} + W_x x_t + b)၊ output o_t = g(W_o h_t) ဖြစ်သည်။ Weight များကို time step များတစ်လျှောက် မျှဝေသုံးသည်။ Backpropagation through time (BPTT) ဖြင့် လေ့ကျင့်ရာတွင် vanishing gradient ဖြစ်နိုင်၍ LSTM, GRU ကဲ့သို့ gated variant များကို လှုံ့ဆော်ပေးသည်။ ### Overview RNN များသည် fixed-size input သာလက်ခံနိုင်သော feedforward networks များနှင့် မတူဘဲ variable-length sequences များကို လက်ခံကာ hidden state ဖြင့် အချိန်တစ်လျှောက် memory ကို ထိန်းသိမ်းသည်။ ၎င်းတို့သည် language modeling, machine translation, speech recognition ကဲ့သို့ tasks များအတွက် အခြေခံဖြစ်သည် — သို့သော် long sequences တွင် vanishing gradients ပြဿနာ ရှိသည်။ ### How It Works 1. Sequence ကို time step အလိုက် ဖြတ်သန်းခြင်း — token/timestep တစ်ခုစီကို input x_t အဖြစ် ယူခြင်း 1. Hidden state update — h_t = f(W_h h_{t−1} + W_x x_t) ဖြင့် ယခင် state နှင့် input ကို ပေါင်းစပ်ခြင်း 1. Output generation — လိုအပ်လျှင် h_t မှ output o_t ထုတ်ခြင်း 1. Backpropagation through time (BPTT) — loss ကို time steps အားလုံးတစ်လျှောက် နောက်ပြန်ဖြန့်ခြင်း 1. Repeat — sequence အားလုံးကို သင်ယူသည့်အထိ epochs များစွာ လေ့ကျင့်ခြင်း ### Formulas - **RNN hidden state recurrence**: `h_t = f(W_h h_{t−1} + W_x x_t + b)` - **Output computation**: `o_t = g(W_o h_t)` --- ## Entry № 019: Principal Component Analysis (PCA) **Burmese Title**: အဓိက အစိတ်အပိုင်း ခွဲခြမ်းစိတ်ဖြာမှု (Principal Component Analysis) **Level**: intermediate | **Topics**: machine-learning **URL**: https://aipedia.burmesestack.com/terms/pca ### Definitions - **EL5 (Simple)**: PCA ဆိုတာ အချက်အလက် (feature) အများကြီးကို အဓိကအရေးပါဆုံး အနည်းငယ်အဖြစ် ချုံ့ပေးတဲ့ နည်း။ ဥပမာ - မှတ်စု စာမျက်နှာ ၅၀ ကို အဓိကအချက် ၂-၃ ချက်အဖြစ် အနှစ်ချုပ်လိုက်တာမျိုး — အရေးကြီးတဲ့ သတင်းအချက်အလက် အများစုကို ထိန်းသိမ်းရင်း ပိုရိုးရှင်းအောင် လုပ်တာပါ။ - **Standard (Burmese)**: PCA သည် feature များစွာ (high-dimensional data) ကို ၎င်းတို့၏ variance အများဆုံး ဦးတည်ချက် (principal components) များပေါ်သို့ project လုပ်၍ dimension လျှော့ချသော unsupervised နည်းလမ်းဖြစ်သည်။ သတင်းအချက်အလက် အများစုကို ထိန်းသိမ်းရင်း data ကို ပိုရိုးရှင်း၊ visualize လုပ်ရ လွယ်ကူစေသည်။ - **Technical (Deep)**: PCA သည် principal component များ၏ orthogonal basis — data covariance matrix ၏ eigenvector များ (SVD ဖြင့်လည်း) — ကို explained variance အလိုက် စီ၍ ရှာသည်။ Top-k component များပေါ် project လုပ်ခြင်းက ထိန်းသိမ်းသော variance ကို အမြင့်ဆုံးဖြစ်စေသည့် lower-dimensional representation ပေးသည်။ Data ကို center (နှင့် များသောအားဖြင့် standardize) လုပ်ရသည်။ ၎င်းသည် compression, denoising, visualization, decorrelation တို့တွင် သုံးသော linear, unsupervised transform ဖြစ်သည်။ ### Overview PCA သည် dimensionality reduction ၏ အခြေခံအကျဆုံး နည်းလမ်းဖြစ်ပြီး feature များ များပြားလွန်းသည့် (curse of dimensionality) ပြဿနာကို ဖြေရှင်းရာတွင်၊ data ကို 2D/3D ဖြင့် visualize လုပ်ရာတွင် သုံးသည်။ Linear projection ဖြစ်သောကြောင့် non-linear structure များအတွက် t-SNE / UMAP ကို သုံးလေ့ရှိသည်။ ### How It Works 1. Data ကို center (mean = 0) ပြု၊ များသောအားဖြင့် standardize လုပ်သည် 1. Covariance matrix ၏ eigenvector / eigenvalue (သို့) SVD ကို တွက်သည် 1. Variance အများဆုံး principal component များကို ရွေးသည် 1. Top-k component များပေါ်သို့ data ကို project လုပ်၍ dimension လျှော့သည် --- ## Entry № 020: Transformer **Burmese Title**: Transformer (နည်းပညာအမည်) **Level**: intermediate | **Topics**: deep-learning, nlp **URL**: https://aipedia.burmesestack.com/terms/transformer ### Definitions - **EL5 (Simple)**: Transformer ဆိုတာ စာသားတွေကို နားလည်ပြီး ဆက်ရေးနိုင်အောင် လုပ်ပေးတဲ့ အဆောက်အအုံပုံစံ။ စကားလုံးတိုင်းကို မျဉ်းဖြောင့်အတိုင်း ဆက်ဖတ်တဲ့ စာအုပ်လိုမဟုတ်ဘဲ စကားလုံးအားလုံးကို တစ်ပြိုင်နက် ကြည့်ပြီး ဘယ်ဟာတွေ ဆက်စပ်နေလဲဆိုတာကို attention နဲ့ ချိန်ဆတယ်။ ChatGPT လိုဟာမျိုးတွေရဲ့ နောက်ကွယ်မှာ ဒီပုံစံပဲ။ - **Standard (Burmese)**: Transformer သည် attention mechanisms များကို အခြေခံထားသော neural network architecture ဖြစ်ပြီး recurrence ကို အသုံးမပြုဘဲ sequence data ကို parallel လုပ်ဆောင်နိုင်သောကြောင့် modern LLMs များ၏ ကျောရိုးဖြစ်လာသည်။ - **Technical (Deep)**: Transformer သည် တူညီသော block N ခုကို ထပ်ဆင့်တည်ဆောက်သည်။ Block တစ်ခုစီတွင် multi-head self-attention + position-wise feed-forward network ပါ၍ residual connection နှင့် layer norm ဖြင့် တွဲထားသည်။ Encoder က contextual representation များ ထုတ်; decoder (generative model အတွက်) က masked self-attention နှင့် cross-attention ဖြင့် token များကို autoregressive ခန့်မှန်းသည်။ ### Overview Transformer ၏ အဓိကအောင်မြင်မှုမှာ parallelism — RNN များနှင့်မတူဘဲ sequence အားလုံးကို တစ်ပြိုင်နက် process လုပ်နိုင်ခြင်းဖြစ်သည်။ ထို့ကြောင့် training data အမြောက်အမြားနှင့် scaling လုပ်နိုင်ပြီး GPT, Gemini, BERT ကဲ့သို့ ကြီးမားသော models များ ဖြစ်ပေါ်လာသည်။ ### How It Works 1. Tokenize + embedding — စာသားကို vectors အဖြစ်ပြောင်း၊ position encoding ထည့်သည် 1. Self-attention — token တစ်ခုစီက sequence တစ်ခုလုံးပေါ် attend လုပ်သည် 1. Feed-forward network — မျဉ်းကြောင်းမဟုတ်သော transformation ဖြင့် features ချဲ့သည် 1. Residual + LayerNorm — နက်ရှိုင်းစွာ stacking လုပ်နိုင်ရန် stabilization 1. Repeat N layers — contextual representation များကို အဆင့်ဆင့် သန်းစေသည် ### Formulas - **Multi-head attention**: `MultiHead(Q,K,V) = Concat(head₁,…,headₕ)Wᴼ, headᵢ = Attention(QWᵢ^Q, KWᵢ^K, VWᵢ^V)` --- ## Entry № 020: LSTM & GRU (LSTM) **Burmese Title**: အယ်လ်အက်စ်တီအမ် နှင့် ဂျီအာယူ (LSTM & GRU) **Level**: intermediate | **Topics**: deep-learning, nlp **URL**: https://aipedia.burmesestack.com/terms/lstm ### Definitions - **EL5 (Simple)**: LSTM နဲ့ GRU ဆိုတာ RNN ရဲ့ 'မှတ်ဉာဏ်' ကို ပိုကောင်းအောင် လုပ်ပေးတဲ့ အဆင့်မြှင့် ပုံစံတွေ။ ရိုးရိုး RNN က စာကြောင်း ရှည်လာရင် အစပိုင်းက အကြောင်းအရာတွေ မေ့သွားတတ်တယ်။ LSTM/GRU မှာတော့ 'ဂိတ်' (gate) လေးတွေ ပါလာပြီး 'ဘာကို မှတ်ထား၊ ဘာကို မေ့လိုက်' ဆိုတာ ထိန်းချုပ်နိုင်တာမို့ ကြာရှည် မှတ်ဉာဏ် ပိုကောင်းတယ်။ - **Standard (Burmese)**: LSTM (Long Short-Term Memory) နှင့် GRU (Gated Recurrent Unit) သည် ရိုးရိုး RNN ၏ vanishing gradient နှင့် long-term memory ပြဿနာကို gate mechanism များဖြင့် ဖြေရှင်းသော RNN အမျိုးအစားများဖြစ်သည်။ Gate များက information ကို ဘယ်လောက် သိမ်းမည်၊ မေ့မည်၊ ထုတ်မည် ဆိုသည်ကို ထိန်းချုပ်၍ sequence ရှည်များတွင်ပါ dependency များကို ထိန်းသိမ်းနိုင်သည်။ - **Technical (Deep)**: LSTM သည် cell state နှင့် gate သုံးခု (input, forget, output) ကို မိတ်ဆက်၍ information flow ကို ထိန်းညှိကာ gradient ကို sequence ရှည်များတွင် propagate စေ၍ vanishing gradient ကို သက်သာစေသည်။ GRU က ၎င်းကို gate နှစ်ခု (reset, update) အဖြစ် ရိုးရှင်းစေ၍ parameter နည်းသော်လည်း တူညီသော performance ရသည်။ နှစ်ခုစလုံး Transformer က long-range task များတွင် အစားထိုးမီ sequence model ၏ ဩဇာအရှိဆုံး ဖြစ်ခဲ့သည်။ ### Overview LSTM နှင့် GRU သည် Transformer မပေါ်မီ NLP နှင့် sequence modeling ၏ အဓိက architecture များဖြစ်ခဲ့သည်။ Gate mechanism က ရိုးရိုး RNN ၏ 'ကြာရှည် မမှတ်နိုင်' ပြဿနာကို ဖြေရှင်းပေးသည်။ GRU သည် LSTM ထက် ရိုးရှင်း၍ parameter နည်းသော်လည်း များသောအခါ တူညီသော performance ရရှိသည်။ ### How It Works 1. Cell state — ကြာရှည် information သယ်ဆောင်သည့် လမ်းကြောင်း (LSTM) 1. Forget gate — မလိုတော့သော အချက်အလက်ကို ဖယ်ရှားသည် 1. Input gate — အသစ်ဝင်လာသော information ကို ဘယ်လောက် သိမ်းမည် ဆုံးဖြတ်သည် 1. Output gate — hidden state အဖြစ် ဘာထုတ်မည် ဆုံးဖြတ်သည် 1. GRU — reset နှင့် update gate နှစ်ခုသာဖြင့် ပိုရိုးရှင်းစွာ လုပ်ဆောင်သည် --- ## Entry № 021: Large Language Model (LLM) **Burmese Title**: ကြီးမားသောဘာသာစကားမော်ဒယ် **Level**: intermediate | **Topics**: deep-learning, nlp, generative-ai **URL**: https://aipedia.burmesestack.com/terms/llm ### Definitions - **EL5 (Simple)**: Large Language Model ဆိုတာ စာတွေ အများကြီးဖတ်ပြီး နောက်စာလုံးကို ခန့်မှန်းဖို့ လေ့ကျင့်ထားတဲ့ ကြီးမားတဲ့ neural network။ စာတစ်ကြောင်းဖတ်ပြီး 'ဒီပြီးရင် ဘယ်စကားလုံး လာနိုင်မလဲ' လို့ အဆင့်ဆင့် ခန့်မှန်းရင်း စာတွေ၊ အဖြေတွေ ရေးတယ်။ ဥပမာ ChatGPT လိုမျိုး။ - **Standard (Burmese)**: Large Language Model (LLM) သည် internet-scale text data များဖြင့် လေ့ကျင့်ထားသော ကြီးမားသော transformer-based language models များဖြစ်သည်။ Autoregressive နည်းဖြင့် နောက်လာမည့် token ကို ခန့်မှန်းခြင်းဖြင့် text generation, summarization, translation, question answering စသည့် tasks များကို လုပ်ဆောင်နိုင်သည်။ - **Technical (Deep)**: LLM များသည် massive corpus ပေါ် next-token prediction ဖြင့် လေ့ကျင့်ထားသော transformer decoder များဖြစ်ပြီး parameter သန်း/ဘီလီယံ ဂဏန်းများစွာ ပါဝင်သည်။ Scaling law သက်သေများက model size, data, compute တိုးလာသည်နှင့်အမျှ စွမ်းဆောင်ရည် ခန့်မှန်းနိုင်စွာ တိုးကြောင်း ပြသည်။ များသောအားဖြင့် instruction tuning (SFT) နှင့် preference optimization (RLHF) ဖြင့် align လုပ်သည်။ Self-attention က token များကြား long-range dependency ကို ဖြစ်နိုင်စေသည်။ ### Overview LLM များ၏ အဓိကစွမ်းအားမှာ scale နှင့် next-token prediction တို့မှ လာသည်။ GPT မျိုးဆက်များသည် autoregressive decoder များဖြစ်ပြီး token တစ်ခုချင်းကို အစဉ်လိုက်ထုတ်လုပ်သည်။ Training သည် pretrain → instruction tuning (SFT) → preference alignment (RLHF) ဟူသော အဆင့်သုံးဆင့်ဖြင့် လုပ်ဆောင်သည်။ ### How It Works 1. Tokenization — text ကို tokens များအဖြစ် ခွဲခြားပြီး vocabulary မှ IDs အဖြစ် ပြောင်းခြင်း 1. Pretraining — massive text corpus မှ နောက်လာမည့် token ကို ခန့်မှန်းရန် transformer ကို လေ့ကျင့်ခြင်း (self-supervised) 1. Supervised fine-tuning (SFT) — instruction-output pairs များဖြင့် human-intended behavior ကို လိုက်နာစေရန် လေ့ကျင့်ခြင်း 1. RLHF — human feedback မှ reward model ဖြင့် outputs များကို ဦးစားပေးရွေးချယ်စေရန် optimize လုပ်ခြင်း 1. Inference — generation loop ဖြင့် token တစ်ခုချင်း အဆင့်ဆင့်ထုတ်လုပ်ခြင်း ### Formulas - **Autoregressive next-token prediction**: `P(x_1..x_T) = Π_t P(x_t | x_