ALEX / ATTENTION WINDOW
Does full attention earn its compute?
The default SSSL pattern alternates three short attention windows with one full window. Compare it with L: full attention at every layer.
Working notes
- ✓ Located WINDOW_PATTERN and the layer-window calculation in train.py.
- ✓ Confirmed the final layer always uses full attention.
- □ Run the unchanged SSSL baseline, then the L candidate.
- □ Compare validation bits per byte and tokens processed.
Next step
Change WINDOW_PATTERN only. More context may help prediction but reduce training throughput.