You need to be more specific. If your point is that you literally cannot backprop while doing a forward pass, then sure. But you can certainly do inference then backprop on the outcome…
I fail to see how that is substantially different from a human doing something then reflecting and learning.
You can critique LLMs and transformers generally, but to say it’s somehow a problem with backprop is a bit of a stretch.
My read is that the user is saying that creating a world view contains trial and error. Online learning if you will. With the current setup this learning is very much not like this. While that makes a human to machine analogy I am also not too convinced by this reasoning. Just allow for a larger time lag and then the inference and learning is at the same scale.
Well to be precise the pretraining is prevented by this but the online post training does have this issue. The reinforcement learning trials show some sort of automation but I am not aware of any unsupervised methods that have been as successful as the latest batches of llms — haven’t been in this field for a while so honestly dunno.
I don’t believe there’s any requirement to learn during a particular part of the lifecycle? And if there was, the short term memory of a context window fulfils that
Backpropagation, the algorithm behind current machine learning systems, precludes an active world-model; it can’t learn during inference.
You need to be more specific. If your point is that you literally cannot backprop while doing a forward pass, then sure. But you can certainly do inference then backprop on the outcome…
I fail to see how that is substantially different from a human doing something then reflecting and learning.
You can critique LLMs and transformers generally, but to say it’s somehow a problem with backprop is a bit of a stretch.
My read is that the user is saying that creating a world view contains trial and error. Online learning if you will. With the current setup this learning is very much not like this. While that makes a human to machine analogy I am also not too convinced by this reasoning. Just allow for a larger time lag and then the inference and learning is at the same scale.
Yes but that’s related to LLMs specifically, not backpropagation. There’s plenty of ML paradigms that use backprop and have continual learning setups.
Backprop is the process of how the weights are updated.
Well to be precise the pretraining is prevented by this but the online post training does have this issue. The reinforcement learning trials show some sort of automation but I am not aware of any unsupervised methods that have been as successful as the latest batches of llms — haven’t been in this field for a while so honestly dunno.
I don’t believe there’s any requirement to learn during a particular part of the lifecycle? And if there was, the short term memory of a context window fulfils that
I thought learning was an important part of intelligence.
Huh, didn’t know there was a word for it, thanks.