Hmmm, all these LLMs are built on a particular machine learning architecture called a Transformer.
This paper from 2017 started the ball rolling - Attention Is All You Need https://arxiv.org/abs/1706.03762
and Wikipedia entry: https://en.wikipedia.org/wiki/Attention_Is_All_You_Need
Essentially (to my understanding) it takes a particular training step and shows you can do it in parallel rather than serial and ... voila! Instead of going through all of human knowledge on the internet (note my sarcasm) taking, I don't know, 12 million years, you can do it all (as long as you've got your own multi-gigawatt power supply) in an afternoon.
But here's the (it seems) proof: All these GPTs (Generative Pre-Trained Transformers) really are just dealing in probabilities, rather than actually awareness of what any of it means.
It's a rainy October day today. I'm going to go downstairs and heat up my coffee.
You and I both know what that actually means.
