Simulations
Attention Heatmap
Transformer ၏ self-attention — query token က key tokens များပေါ်တွင် weights ကို မည်ကဲ့သို့ ဖြန့်ဝေသည်ကို ကြည့်ပါ။ Row တစ်ခုကို နှိပ်ပြီး attention distribution ကို ကြည့်နိုင်သည်။ attention mechanism term နှင့် တွဲလေ့လာပါ။
Attention Matrix
7 × 7
Row = query token (ကြည့်တဲ့ token) · Column = key token (အာရုံစိုက်ခံရတဲ့ token)
Sentence
Attention head
Temperature
Temperature နိမ့်လေ attention က sharp (one token ပေါ်စုလေလဲ)၊ မြင့်လေ distribution က uniform နီးလေလဲ ဖြစ်သည်။
Attention of “cat”
The
12.0%
cat
24.1%
sat
17.8%
on
12.0%
the
12.0%
mat
12.0%
.
10.3%
