A friendly tour — no scary math, lots of games.
Use → or Space to move forward · Esc for the table of contents
“How do we get a machine to understand words?”
Alan Turing was already asking a version of it in 1950: “Can machines think?”
Alan Turing, 1951
Photo: Elliott & Fry · public domain, via Wikimedia Commons
What ChatGPT is actually doing (you'll play it as a game)
How machines read: from rules to word maps
RNNs, forgetting, and the attention trick
The 2017 invention that changed everything
ChatGPT, AIs that see, and robots that act
What AI still gets wrong — and what's next
Recognizes YOUR face out of 8 billion
Photo: mikemacmarketing · CC BY 2.0 · Wikimedia Commons
Your keyboard guesses the next word… remember this one.
TikTok / Reels learned exactly what makes you scroll
Photo: Solen Feyissa · CC BY-SA 2.0 · Wikimedia Commons
Sound → words → intent → an answer, in a second
Photo: Smart Home Perfected · CC BY 2.0 · Wikimedia Commons
But one kind of AI feels different — it talks to you.
How?! Let's find out — by playing a game.
Waymo robotaxis carry paying passengers in San Francisco & Phoenix — every day, empty front seat.
Photo: Oleg Yunakov · CC BY-SA 4.0, via Wikimedia Commons
AlphaFold predicted the 3-D shape of ~200 million proteins — its creators won the 2024 Nobel Prize in Chemistry.
AlphaFold structure: SharnKB · CC BY 4.0, via Wikimedia Commons
Surgeons already operate through robotic arms; AI assistance is arriving in hospitals right now.
Photo: Marcy Sanchez, U.S. Army · public domain, via Wikimedia Commons
All different systems — but the loudest revolution is the one that learned to talk. Let's find out how.
These are real predictions from GPT-2 (OpenAI's 2019 model, running on this laptop). Click a token to add it to the story!
“Given all the words so far,
guess the next word.”
That's it. That's the secret.
Do it once — you get autocomplete.
Do it brilliantly, billions of times in a row — you get essays, code, poems, and jokes.
In math-speak: the model learns P(next token | everything so far) — one probability for every token in its vocabulary — and writing is just sampling from that distribution, over and over.
The rest of this talk: how we taught a machine to guess that well.
✓ So far: we know the one trick behind it all — guess the next word, brilliantly.
✗ Still broken: computers can't even SEE words, only numbers. Every idea in AI language starts here: how do we write meaning as numbers?
“hello”
↓
104 · 101 · 108 · 108 · 111
to a computer, text is just codes for letters
But knowing the letter codes ≠ knowing the meaning.
The codes for “cat” and “kitten” look totally unrelated.
The codes for “cat” and “car” look almost identical.
So for ~70 years, the real question has been:
how do we turn MEANING into numbers?
A fake “therapist” chatbot. Try it — click a message:
The real thing — an original ELIZA session (1966). Public domain, via Wikimedia Commons
— click a message first —
No understanding at all — just find keyword → use template.
Rules can't cover jokes, slang, typos, sarcasm…
Language has too many exceptions to list.
Humans write the rules,
machine follows them
Show the machine millions of examples —
let it figure out the patterns itself
Like learning your native language: nobody handed you a grammar book at age 2.
You just heard a LOT of sentences.
Treat a sentence like a bag of Scrabble tiles — count what's inside, ignore the order.
“Dog bites man”
becomes
dog ×1bites ×1man ×1
“Man bites dog”
becomes
dog ×1bites ×1man ×1
Same bag… completely different news story.
Word counts are useful (spam filters used this for years) — but order and meaning are lost. Next idea, please.
“You shall know a word by the company it keeps.”
Words that show up in similar sentences probably mean similar things:
“I drank a glass of milk”
“I drank a glass of juice”
→ milk and juice keep showing up in the same spots → put them close together on the map
Why is this worth doing? Before, every word was a stranger — knowing about “cat” told you nothing about “kitten”. On a map, whatever the model learns about one word rubs off on its neighbors. Learn one word, get its whole neighborhood for free.
In 2013, Word2Vec let a computer read billions of sentences and place every word on a giant map — automatically.
Under the hood: start every word at a random spot. Read the internet; each time two words appear together, nudge them a little closer (and nudge non-neighbors apart). Repeat billions of times — the map organizes itself. Each position is a vector of 300+ numbers, and “similar meaning” literally means “small angle between vectors” (cosine similarity).
each word = a list of numbers
= coordinates
cat = [0.21, −0.88, 1.03, …]
kitten = [0.25, −0.81, 0.97, …]
car = [−1.90, 0.42, −0.33, …]
cat & kitten: nearly the same numbers!
(real embeddings use hundreds of dimensions, not 2)
This is GPT-2's actual word map: positions = PCA of its real embedding vectors, similarities = real cosine values (computed on this laptop). Real data is delightfully messy — notice “robot” floating between the animals and the machines.
king − man + woman ≈ queen
The map learned that the direction from “man” to “king” means royalty. Walk the same direction from “woman”… and you land on “queen”.
We checked, for real: in GPT-2's embeddings the closest word to king − man + woman (after the inputs themselves) is queen (71%), then princess (60%). Bigger embedding models pull the same trick with Paris − France + Japan ≈ Tokyo.
Nobody programmed this. It emerged from reading billions of sentences. Meaning became geometry.
✓ So far: every word has meaningful coordinates now (embeddings — 2013's gift).
✗ Still broken: meaning lives in ORDER — “dog bites man” ≠ “man bites dog”. We need machines that read in sequence… and they'd better not forget the beginning.
Idea: read like a human — left to right, one word at a time, keeping a “memory” as you go. This is the RNN (Recurrent Neural Network).
The whole idea in one line: new memory = f(old memory, current word) — in symbols, ht = f(ht−1, xt). The same little function, applied over and over down the sentence.
Early Siri, Google Translate (2016), speech-to-text — all RNNs
Must finish word 7 before starting word 8 — no shortcuts
The memory is small — old words get squeezed out…
Who was happy? To answer, the model must remember “girl” all the way to the end…
The plan: read the whole English sentence → squeeze it into ONE memory vector → write the Chinese sentence from that.
the whole sentence…
Ilovehotpotwithmyfriendsoncolddays…crushed into this
…then rebuilt
冷天和朋友吃火锅…Like reading a whole book, closing it, and rewriting it from a single sticky note.
Long sentences? Details get lost. There had to be a better way…
New rule: while writing each output word, the model may glance back at the whole input and highlight what matters right now.
Like doing a reading-comprehension test:
you don't memorize the passage —
you look back and highlight the sentence you need.
This trick is called attention.
Remember the word — it's about to take over the world.
I love hot pot with my friends
writing “火锅” → attention shines on “hot pot”, dims the rest
Thicker line = more attention. Notice 火锅 attends to TWO English words — attention handles that effortlessly. An old RNN had to hope its sticky-note memory kept both.
This one idea (2015) made Google Translate dramatically better almost overnight. (Weights here are schematic — real measured attention arrives two chapters from now.)
✓ So far: attention lets the model look back at everything — forgetting: solved (2015).
✗ Still broken: still reading one word at a time — far too slow to ever learn from the whole internet. Time to throw away the queue entirely.
Google, 2017 — the paper
“Attention Is All You Need”
If looking back works so well… why keep the slow one-word-at-a-time reader at all?
Throw away the RNN. Keep only attention.
Reading in order, word by word
Every word looks at every other word — all at once
The Transformer — the T in GPT, the engine of ALL modern AI
Fun fact: the authors thought it was a nice translation paper. It became one of the most cited papers in history.
RNN — must wait for each word:
Transformer — everyone at once, everyone sees everyone:
Because meaning depends on context. Quick — what does “it” refer to here?
You resolved “it” instantly — attention does exactly this: the word “it” looks around the sentence and locks onto its true meaning. Swap one word at the END, and “it” points somewhere completely different!
The cat chased it — our question: what does “it” mean here?
“What am I looking for?”
“it” asks: “I'm a stand-in word… WHO do I refer to?”
“Here's what I offer.”
“cat” wears the tag: “main animal of this sentence, did the chasing”.
“Pick me, and this is what you get.”
“cat”'s package: its actual meaning — furry, animal, the chaser.
Where do these come from? Each one is just the word's coordinates run through a small learned table (a matrix). Three tables → three vectors. Those tables are what training tunes.
Question asked, name tags on, packages ready. Next slide: the matching game.
“it” walked in purple. It walks out cat-colored — the word literally absorbed the meaning it paid attention to.
1 · Match. “it”'s question is compared with every word's name tag. “cat” wins with +0.95 — everyone else goes negative. (Dot products can be negative: it just means “bad match”.)
2 · Percentages. Softmax turns the scores into shares of 100%. +0.95 → 81.6%. (That's softmax's entire job.)
3 · Collect the packages. “it” takes 82% of cat's package, plus small slices of the others, and mixes them into its new self.
That's attention — every word softly grabs meaning from the words that matter to it. And every word plays this game simultaneously, not just “it”.
These numbers are real: we ran “The cat chased it” through GPT-2 on this laptop and read out layer 5, head 4 — a head that genuinely exists and genuinely does this job.
The game you just watched is one attention head. Each head learns to hunt for one kind of connection:
Nobody assigns these jobs — heads specialize on their own during training. Researchers discover what each head does afterwards, like biologists.
Why this beat the RNN: any two words connect in one hop — no more information crawling word-by-word and fading. And it's all matrix math, so GPUs run every head, every floor, at once.
probabilities out ↑
↑ words go in at the bottom
Each floor re-reads the floor below. Low floors catch grammar; high floors catch plot, logic, tone.
The cat sat on the
Its name: a decoder-only Transformer — these days usually called the “Llama architecture”, after the open model that made it famous around 2024. Llama, Qwen, Mistral, GPT — all this same machine in different sizes.
And training? The same trip, plus a correction: the model reads real sentences, its guess is compared with the actual next word (that is the loss), and every knob gets a tiny nudge — gradient descent, repeated trillions of times.
Inference = the trip. Training = the trip + feedback.
Scale of it: every word ChatGPT has ever written to anyone = one full trip. A 500-word answer = 500 trips — each touching hundreds of billions of knobs.
The actual diagram from the 2017 paper — every box in it is something you've now watched happen.
Diagram: Daniel Voigt Godoy (dvgodoy) · CC BY 4.0, via Wikimedia Commons
Model size in “parameters” — the adjustable knobs a model tunes while learning. Each gridline is 10×!
The shock: skills nobody programmed just appeared with scale — translation, arithmetic, coding, writing poems… This is called emergence, and it's why everyone suddenly went big. (GPT-4's size is an estimate — OpenAI never confirmed it)
✓ So far: the Transformer (2017) — parallel, hungry, endlessly scalable.
✗ Still broken: it's just an engine. Now: feed it the internet, watch new skills emerge with scale, then fine-tune it into an assistant. This is where ChatGPT is born.
“I went to [MASK] and ate ramen.”
Random words are hidden and the model must guess them by reading both sides — a cloze test at internet scale. Superpower: understanding. When you search “bank by the river”, a BERT-style model is what figures out you don't mean money.
How you used it: pre-train once… then for every job — spam filter, review scoring, search ranking — collect labeled examples and fine-tune a separate copy. 100 tasks = 100 specialized models + 100 datasets.
Technical name: masked language model (an encoder).
“I went to Ichiran and ate [?]”
Reads strictly left-to-right and only ever guesses what comes next — it never peeks ahead. Sounds like a handicap, but it's exactly what lets the model write.
How you use it: just type. The task rides inside the prompt — “Translate to French: …”. GPT-3 (2020) stunned researchers by doing tasks it was never trained for, from a few examples pasted into the prompt (in-context learning). 100 tasks = 1 model + 100 prompts.
Technical name: autoregressive language model (a decoder).
So why did the GPT branch take over?
Generation turned out to be universal: any task can be phrased as “continue this text” — one model covers them all, no retraining, and the interface is just… language. A conversation is simply one more prompt, which is why the chatbot era is GPT-shaped.
(BERT-style models still quietly rank your search results — they just never became chatbots.)
It creates text — word by word, like our game in Chapter 1
Before you ever met it, it read a huge chunk of the internet
Powered by the attention engine from Chapter 4 — the decoder-only variant, which does exactly one thing: predict the next token
= a Transformer that read the internet and finishes your sentences.
Sounds simple. The scale is anything but…
Training = the guess-the-next-word game on the whole internet. Each wrong guess is measured (the loss), and every weight is nudged a tiny step in the direction that reduces it — gradient descent, repeated trillions of times.
To guess the next word on the WHOLE internet, you're forced to learn grammar, facts, logic, style, even jokes.
The guessing game was secretly the ultimate homework.
This is what “AI hardware” looks like: an NVIDIA H100 — ~$30,000 each, and top labs use tens of thousands of them.
Photo: 极客湾Geekerwan · CC BY 3.0, via Wikimedia Commons
PROMPT
“My cat” — given to raw GPT-2 (124M), which can only continue text
OUTPUT
It only knew how to CONTINUE text — not to answer you:
It saw your question and thought: “ah, a LIST of questions — I'll continue the pattern.”
A brilliant engine with zero manners. One more training step was needed.
Keep training the same model — but on a new, much smaller dataset: tens of thousands of example conversations, each a question followed by a genuinely helpful answer, written and curated by people.
The architecture doesn't change at all — it is still only predicting the next token. But after fine-tuning, the most likely continuation of a question is a helpful answer, because that is what its recent training data looks like.
Under the hood: same objective, same gradient descent — only the data changed.
This recipe — pretrain on the internet, then fine-tune on conversations — is how GPT became ChatGPT, and it's the same recipe behind Llama, Qwen and the rest.
Time each product needed to reach 100 million users:
ChatGPT: fastest-growing consumer app in history at the time. Your parents, your teachers, and your group chats all discovered it the same winter.
OpenAI — the ChatGPT family
Anthropic — strong at writing and coding (it helped build these slides)
Google — lives inside Search, Android, Docs
Meta — open weights: anyone can download & run it
Alibaba & DeepSeek — top-tier open models from China
Students today fine-tune small models on a laptop!
Different names, same DNA: Transformer + next-word prediction + fine-tuning.
But wait — these models can look at photos now. How does TEXT prediction do that?
✓ So far: LLMs have mastered reading and writing.
✗ Still broken: the world isn't made of text — it's pictures, sounds and actions. Final stretch: same brain, new senses — eyes, a paintbrush, hands.
+ eyes. Photos are chopped into tokens too.
GPT-4V · Claude · Gemini
+ a paintbrush. Your words steer a noise-remover.
DALL·E · Midjourney (the artsy cousin)
+ hands. Robot movements become tokens.
RT-2 · π0 · humanoids
+ a to-do list. LLMs that use tools: browser, code, your files.
the 2025-26 frontier
The pattern, one last time: tokenize a new sense → feed the same Transformer.
Let's meet the first three up close.
A VLM (Vision-Language Model) chops the photo into small patches (a Vision Transformer uses 16×16-pixel squares), embeds each patch exactly like a word token, and feeds them into the same Transformer alongside your text.
← drag me
Diffusion models (DALL·E, Midjourney, Stable Diffusion) train by watching millions of photos get ruined by static — and learning to reverse it.
To create art: start from pure random static, remove noise step by step, and steer with your text prompt — “a corgi astronaut, oil painting”
“Théâtre D’opéra Spatial” — made with Midjourney, it won 1st prize at a 2022 art fair before judges knew. The US Copyright Office later ruled AI images can't be copyrighted — which is why I can show it to you for free
Jason M. Allen / Midjourney, 2022 · public domain, via Wikimedia Commons
A VLA (Vision-Language-Action model) outputs movements instead of words. Actions are discretized — “rotate joint 3 by +2°” becomes a token ID — so a robot trajectory is literally a sentence the Transformer can predict.
Third time's the pattern: if you can tokenize it, you can predict it — words, pixels… now actions.
A Unitree robot dog — Google RT-2, π0, Tesla Optimus… the frontier: robots you simply talk to.
Photo: Sgt. Amber Edwards, U.S. Army, 2023 · public domain, via Wikimedia Commons
It predicts plausible-sounding words — “sounds right” and “IS right” are not the same thing!
Lying needs knowing the truth. It just… autocompletes. This is called hallucination.
Verify important facts. Great assistant, terrible oracle. Especially for homework
A person who knows ZERO Chinese sits in a room with a giant rulebook. Chinese notes slide in — they look up the symbols, follow the rules, slide perfect Chinese replies out.
Outside, it looks fluent. Inside… does anyone understand Chinese?
Philosophers still argue about this.
Team “just statistics”: it's a parrot with a gigantic memory — patterns in, patterns out.
Team “real understanding”: to predict text about the world THIS well, it must have built some internal model of the world. And hey — aren't brains pattern machines too?
There is no agreed answer. You get to make up your own mind — scientists genuinely haven't.
Models learn from OUR internet — including our stereotypes. Ask it to draw “a doctor” and “a nurse” and check who it draws…
Trained on millions of artists' work without asking. Artists are suing. Courts worldwide are deciding right now.
Calculators didn't end math careers — they changed them. The winners were people who learned to USE the new tool.
These aren't reasons to fear AI — they're reasons to understand it.
Which, congrats, you now do better than 99% of people.
Two minutes, then we compare notes.
Everything else — essays, code, conversations — is that one trick, repeated.
Rules were brittle → learn from data. Counting lost meaning → word maps. RNNs forgot → attention. Attention was slow in RNNs → Transformer. Then: scale.
GPT writes, VLMs see, VLAs act. Same engine, new senses. The story is still being written — possibly by someone in this room.
“Large Language Models explained briefly” — the most beautiful animations on this topic. Start here.
“Intro to Large Language Models” (1 hr) — by the legend himself. Then “Let's build GPT” when you're ready to code one!
“The Illustrated Transformer” — the classic visual walkthrough, one diagram at a time.
ChatGPT / Claude / Gemini are free — ask one to explain today's talk back to you — then grade it.
Know some Python? Search “nanoGPT” — you can train a tiny GPT on your laptop, on Shakespeare, tonight.
Questions?
the entire field started with “can machines think?” and nobody has answered THAT yet either.
Slides: one self-contained HTML file — arrow keys to explore, every demo is clickable. Take it home!