You seem to assume that autoregressive pretraining (and unfiltered behavior cloning, maybe) are the only ways to improve LLM performance.
You seem to assume that autoregressive pretraining (and unfiltered behavior cloning, maybe) are the only ways to improve LLM performance.