My read is that the user is saying that creating a world view contains trial and error. Online learning if you will. With the current setup this learning is very much not like this. While that makes a human to machine analogy I am also not too convinced by this reasoning. Just allow for a larger time lag and then the inference and learning is at the same scale.
Well to be precise the pretraining is prevented by this but the online post training does have this issue. The reinforcement learning trials show some sort of automation but I am not aware of any unsupervised methods that have been as successful as the latest batches of llms — haven’t been in this field for a while so honestly dunno.
My read is that the user is saying that creating a world view contains trial and error. Online learning if you will. With the current setup this learning is very much not like this. While that makes a human to machine analogy I am also not too convinced by this reasoning. Just allow for a larger time lag and then the inference and learning is at the same scale.
Yes but that’s related to LLMs specifically, not backpropagation. There’s plenty of ML paradigms that use backprop and have continual learning setups.
Backprop is the process of how the weights are updated.
Well to be precise the pretraining is prevented by this but the online post training does have this issue. The reinforcement learning trials show some sort of automation but I am not aware of any unsupervised methods that have been as successful as the latest batches of llms — haven’t been in this field for a while so honestly dunno.