Abstract
We will study a theoretical model (one layer linear self attention with in-context linear task) to rigorously explain several observations (already reported in previous empirical work in the literature) about the characteristics of pretraining and post-training data that are relevant for good performance. Our analysis explain the importance of having a balanced diverse data in pretraining, and hard small set of examples for SFT, and large volume data fro RL. The experiments are on synthetic toy settings and on GPT-2 with synthetic linear regression data.