→ next · Esc contents · F fullscreen
00:00

Generative AI & LLMs

How does ChatGPT work?

A friendly tour — no scary math, lots of games.

Use → or Space to move forward  ·  Esc for the table of contents

Our journey today

One question, 70 years of trying to answer it:

“How do we get a machine to understand words?”
Alan Turing was already asking a version of it in 1950: “Can machines think?”

Alan Turing, 1951

Alan Turing, 1951
Photo: Elliott & Fry · public domain, via Wikimedia Commons

1 · The Secret

What ChatGPT is actually doing (you'll play it as a game)

2 · Words → Numbers

How machines read: from rules to word maps

3 · Learning to Read

RNNs, forgetting, and the attention trick

4 · The Transformer

The 2017 invention that changed everything

5 · GPT & Beyond

ChatGPT, AIs that see, and robots that act

6 · The Big Questions

What AI still gets wrong — and what's next

Chapter 1 · The Secret

You already used AI today — probably before breakfast.

face recognition

Face unlock

Recognizes YOUR face out of 8 billion

Photo: mikemacmarketing · CC BY 2.0 · Wikimedia Commons

goingtothe
Q W E R T Y U I O P
A S D F G H J K L

Autocomplete

Your keyboard guesses the next word… remember this one.

TikTok on a phone

Your feed

TikTok / Reels learned exactly what makes you scroll

Photo: Solen Feyissa · CC BY-SA 2.0 · Wikimedia Commons

smart speaker

“Hey Siri / Alexa”

Sound → words → intent → an answer, in a second

Photo: Smart Home Perfected · CC BY 2.0 · Wikimedia Commons

But one kind of AI feels different — it talks to you.
How?! Let's find out — by playing a game.

Chapter 1 · The Secret

…and it has gone far beyond homework help

Waymo robotaxi

Taxis with no driver

Waymo robotaxis carry paying passengers in San Francisco & Phoenix — every day, empty front seat.

Photo: Oleg Yunakov · CC BY-SA 4.0, via Wikimedia Commons

AlphaFold protein prediction

A Nobel Prize

AlphaFold predicted the 3-D shape of ~200 million proteins — its creators won the 2024 Nobel Prize in Chemistry.

AlphaFold structure: SharnKB · CC BY 4.0, via Wikimedia Commons

Robot-assisted surgery

In the operating room

Surgeons already operate through robotic arms; AI assistance is arriving in hospitals right now.

Photo: Marcy Sanchez, U.S. Army · public domain, via Wikimedia Commons

All different systems — but the loudest revolution is the one that learned to talk. Let's find out how.

Chapter 1 · The Secret

Let's write a story — one word at a time

These are real predictions from GPT-2 (OpenAI's 2019 model, running on this laptop). Click a token to add it to the story!

Once upon a
 Every percentage here is a real GPT-2 output (computed along the top-choice path). The full menu has ~50,000 tokens — we show the top 4.
Chapter 1 · The Secret

ChatGPT's entire job:

“Given all the words so far,
guess the next word.”

That's it. That's the secret.

Do it once — you get autocomplete.
Do it brilliantly, billions of times in a row — you get essays, code, poems, and jokes.

In math-speak: the model learns P(next token | everything so far) — one probability for every token in its vocabulary — and writing is just sampling from that distribution, over and over.

The rest of this talk: how we taught a machine to guess that well.

The map · stop 1 of 5

First stop: turn words into numbers

✓ So far: we know the one trick behind it all — guess the next word, brilliantly.
✗ Still broken: computers can't even SEE words, only numbers. Every idea in AI language starts here: how do we write meaning as numbers?

Chapter 2 · Words → Numbers

Problem: computers don't speak English.
They speak numbers.

“hello”

↓

104 · 101 · 108 · 108 · 111

to a computer, text is just codes for letters

But knowing the letter codes ≠ knowing the meaning.

The codes for “cat” and “kitten” look totally unrelated.
The codes for “cat” and “car” look almost identical.

So for ~70 years, the real question has been:
how do we turn MEANING into numbers?

Chapter 2 · Words → Numbers · 1950s–1980s

First idea: humans write rules. Meet ELIZA (1966)

A fake “therapist” chatbot. Try it — click a message:

Original ELIZA session, 1966

The real thing — an original ELIZA session (1966). Public domain, via Wikimedia Commons

The rule it used:

— click a message first —

No understanding at all — just find keyword → use template.

Rules can't cover jokes, slang, typos, sarcasm…
Language has too many exceptions to list.

Chapter 2 · Words → Numbers

The big idea that changed everything:

Old way

Humans write the rules,
machine follows them

→

New way: Machine Learning

Show the machine millions of examples —
let it figure out the patterns itself

Like learning your native language: nobody handed you a grammar book at age 2.
You just heard a LOT of sentences.

Chapter 2 · Words → Numbers

Try #1: just count the words

Treat a sentence like a bag of Scrabble tiles — count what's inside, ignore the order.

“Dog bites man”

becomes

dog ×1bites ×1man ×1

“Man bites dog”

becomes

dog ×1bites ×1man ×1

Same bag… completely different news story.

Word counts are useful (spam filters used this for years) — but order and meaning are lost. Next idea, please.

Chapter 2 · Words → Numbers · 2013: Word2Vec

Try #2: give every word a location on a map

“You shall know a word by the company it keeps.”

Words that show up in similar sentences probably mean similar things:

“I drank a glass of milk”

“I drank a glass of juice”

→ milk and juice keep showing up in the same spots → put them close together on the map

Why is this worth doing? Before, every word was a stranger — knowing about “cat” told you nothing about “kitten”. On a map, whatever the model learns about one word rubs off on its neighbors. Learn one word, get its whole neighborhood for free.

In 2013, Word2Vec let a computer read billions of sentences and place every word on a giant map — automatically.

Under the hood: start every word at a random spot. Read the internet; each time two words appear together, nudge them a little closer (and nudge non-neighbors apart). Repeat billions of times — the map organizes itself. Each position is a vector of 300+ numbers, and “similar meaning” literally means “small angle between vectors” (cosine similarity).

each word = a list of numbers
= coordinates

cat  = [0.21, −0.88, 1.03, …]
kitten = [0.25, −0.81, 0.97, …]
car  = [−1.90, 0.42, −0.33, …]

cat & kitten: nearly the same numbers!
(real embeddings use hundreds of dimensions, not 2)

Chapter 2 · Words → Numbers

The map of words (hover = neighborhood · click any TWO words = similarity score)

This is GPT-2's actual word map: positions = PCA of its real embedding vectors, similarities = real cosine values (computed on this laptop). Real data is delightfully messy — notice “robot” floating between the animals and the machines.

Chapter 2 · Words → Numbers

Wait — you can do math with words?!

king − man + woman ≈ queen

The map learned that the direction from “man” to “king” means royalty. Walk the same direction from “woman”… and you land on “queen”.

We checked, for real: in GPT-2's embeddings the closest word to king − man + woman (after the inputs themselves) is queen (71%), then princess (60%). Bigger embedding models pull the same trick with Paris − France + Japan ≈ Tokyo.

Nobody programmed this. It emerged from reading billions of sentences. Meaning became geometry.

The map · stop 2 of 5

Next stop: read a whole sentence

✓ So far: every word has meaningful coordinates now (embeddings — 2013's gift).
✗ Still broken: meaning lives in ORDER — “dog bites man” ≠ “man bites dog”. We need machines that read in sequence… and they'd better not forget the beginning.

Chapter 3 · Learning to Read · ~1990s–2015

Words have coordinates. Now: how to read a sentence?

Idea: read like a human — left to right, one word at a time, keeping a “memory” as you go. This is the RNN (Recurrent Neural Network).

The whole idea in one line: new memory = f(old memory, current word) — in symbols, ht = f(ht−1, xt). The same little function, applied over and over down the sentence.

The→ cat→ sat→ down memory updates after every word

It worked!

Early Siri, Google Translate (2016), speech-to-text — all RNNs

But it's slow

Must finish word 7 before starting word 8 — no shortcuts

And it forgets

The memory is small — old words get squeezed out…

Chapter 3 · Learning to Read

Watch an RNN forget

Who was happy? To answer, the model must remember “girl” all the way to the end…

memory of “girl”:
100%

 LSTMs (1997) patched this with learnable “keep / forget” gates — a big help, but reading stayed one-word-at-a-time, and very long texts still faded.
Chapter 3 · Learning to Read · 2014: Seq2Seq

Translation had the same problem — but worse

The plan: read the whole English sentence → squeeze it into ONE memory vector → write the Chinese sentence from that.

the whole sentence…

Ilovehotpotwithmyfriendsoncolddays
→

…crushed into this

···
→

…then rebuilt

冷天和朋友吃火锅…

Like reading a whole book, closing it, and rewriting it from a single sticky note.
Long sentences? Details get lost. There had to be a better way…

Chapter 3 · Learning to Read · 2015: Attention

The fix: stop memorizing — look back

New rule: while writing each output word, the model may glance back at the whole input and highlight what matters right now.

Like doing a reading-comprehension test:
you don't memorize the passage —
you look back and highlight the sentence you need.

This trick is called attention.
Remember the word — it's about to take over the world.

Translating word #3…

I love hot pot with my friends

writing “火锅” → attention shines on “hot pot”, dims the rest

Chapter 3 · Learning to Read

Attention in action (click a Chinese word)

Thicker line = more attention. Notice 火锅 attends to TWO English words — attention handles that effortlessly. An old RNN had to hope its sticky-note memory kept both.

This one idea (2015) made Google Translate dramatically better almost overnight. (Weights here are schematic — real measured attention arrives two chapters from now.)

The map · stop 3 of 5

Stop 3: a new architecture

✓ So far: attention lets the model look back at everything — forgetting: solved (2015).
✗ Still broken: still reading one word at a time — far too slow to ever learn from the whole internet. Time to throw away the queue entirely.

Chapter 4 · The Transformer · 2017

Then 8 researchers asked a crazy question…

Google, 2017 — the paper

“Attention Is All You Need”

If looking back works so well… why keep the slow one-word-at-a-time reader at all?
Throw away the RNN. Keep only attention.

Deleted

Reading in order, word by word

Kept

Every word looks at every other word — all at once

Result

The Transformer — the T in GPT, the engine of ALL modern AI

Fun fact: the authors thought it was a nice translation paper. It became one of the most cited papers in history.

Chapter 4 · The Transformer

The race: reading 8 words

RNN — must wait for each word:

Transformer — everyone at once, everyone sees everyone:

 Parallel = perfect for GPUs (graphics chips do thousands of things at once). Now we could train on the ENTIRE internet.
Chapter 4 · The Transformer

“Every word looks at every word” — why does that matter?

Because meaning depends on context. Quick — what does “it” refer to here?

You resolved “it” instantly — attention does exactly this: the word “it” looks around the sentence and locks onto its true meaning. Swap one word at the END, and “it” points somewhere completely different!

Chapter 4 · The mechanism · part 1 of 4

Up close: every word carries three little vectors

The cat chased it — our question: what does “it” mean here?

query — the question it asks

“What am I looking for?”

“it” asks: “I'm a stand-in word… WHO do I refer to?”

key — the name tag it wears

“Here's what I offer.”

“cat” wears the tag: “main animal of this sentence, did the chasing”.

value — the package it hands over

“Pick me, and this is what you get.”

“cat”'s package: its actual meaning — furry, animal, the chaser.

Where do these come from? Each one is just the word's coordinates run through a small learned table (a matrix). Three tables → three vectors. Those tables are what training tunes.

Question asked, name tags on, packages ready. Next slide: the matching game.

Chapter 4 · The mechanism · part 2 of 4

The matching game: “it” goes shopping (real GPT-2 numbers)

The
q·k: −1.36
8.1%
cat
q·k: +0.95
81.6%
chased
q·k: −2.03
4.2%
it
q·k: −1.63
6.2%
cat 82%
→

“it” walked in purple. It walks out cat-colored — the word literally absorbed the meaning it paid attention to.

1 · Match. “it”'s question is compared with every word's name tag. “cat” wins with +0.95 — everyone else goes negative. (Dot products can be negative: it just means “bad match”.)

2 · Percentages. Softmax turns the scores into shares of 100%. +0.95 → 81.6%. (That's softmax's entire job.)

3 · Collect the packages. “it” takes 82% of cat's package, plus small slices of the others, and mixes them into its new self.

That's attention — every word softly grabs meaning from the words that matter to it. And every word plays this game simultaneously, not just “it”.

These numbers are real: we ran “The cat chased it” through GPT-2 on this laptop and read out layer 5, head 4 — a head that genuinely exists and genuinely does this job.

Chapter 4 · The mechanism · part 3 of 4

One head is a specialist. So you stack lots of them.

The game you just watched is one attention head. Each head learns to hunt for one kind of connection:

who did what to whom
which adjective goes with which noun
how far apart words are
names ↔ facts about them
opening ↔ closing quotes
…plus dozens nobody predicted

Nobody assigns these jobs — heads specialize on their own during training. Researchers discover what each head does afterwards, like biologists.

Why this beat the RNN: any two words connect in one hop — no more information crawling word-by-word and fading. And it's all matrix math, so GPUs run every head, every floor, at once.

probabilities out ↑

Floor 126 (Llama 3's top floor)
⋮
Floor 3 …
Floor 2 — same, on floor 1's output
Floor 1 — all heads look, then a pattern-spotter cleans up

↑ words go in at the bottom

Each floor re-reads the floor below. Low floors catch grammar; high floors catch plot, logic, tone.

Chapter 4 · The mechanism · part 4 of 4

Now the whole machine, live

The cat sat on the

1 · tokenize

→

2 · embed

→

3 · think

→

4 · score

→

5 · pick

 It writes ONE token per trip, then starts over — autoregressive generation. Every ID and percentage on this slide is real GPT-2 output.
Chapter 4 · The mechanism · wrap

What you just watched, officially

Its name: a decoder-only Transformer — these days usually called the “Llama architecture”, after the open model that made it famous around 2024. Llama, Qwen, Mistral, GPT — all this same machine in different sizes.

And training? The same trip, plus a correction: the model reads real sentences, its guess is compared with the actual next word (that is the loss), and every knob gets a tiny nudge — gradient descent, repeated trillions of times.
Inference = the trip. Training = the trip + feedback.

Scale of it: every word ChatGPT has ever written to anyone = one full trip. A 500-word answer = 500 trips — each touching hundreds of billions of knobs.

Transformer architecture diagram

The actual diagram from the 2017 paper — every box in it is something you've now watched happen.

Diagram: Daniel Voigt Godoy (dvgodoy) · CC BY 4.0, via Wikimedia Commons

Chapter 4 · The Transformer · 2018–2024

So they made it bigger. Then bigger. Then BIGGER.

Model size in “parameters” — the adjustable knobs a model tunes while learning. Each gridline is 10×!

The shock: skills nobody programmed just appeared with scale — translation, arithmetic, coding, writing poems… This is called emergence, and it's why everyone suddenly went big. (GPT-4's size is an estimate — OpenAI never confirmed it)

The map · stop 4 of 5

Stop 4: scale it into an LLM

✓ So far: the Transformer (2017) — parallel, hungry, endlessly scalable.
✗ Still broken: it's just an engine. Now: feed it the internet, watch new skills emerge with scale, then fine-tune it into an assistant. This is where ChatGPT is born.

Chapter 5 · GPT & Beyond · 2018

2018: two ways to pre-train — BERT and GPT

BERT (Google) — fill in the blank

“I went to [MASK] and ate ramen.”

Random words are hidden and the model must guess them by reading both sides — a cloze test at internet scale. Superpower: understanding. When you search “bank by the river”, a BERT-style model is what figures out you don't mean money.

How you used it: pre-train once… then for every job — spam filter, review scoring, search ranking — collect labeled examples and fine-tune a separate copy. 100 tasks = 100 specialized models + 100 datasets.

Technical name: masked language model (an encoder).

GPT (OpenAI) — guess the next word

“I went to Ichiran and ate [?]”

Reads strictly left-to-right and only ever guesses what comes next — it never peeks ahead. Sounds like a handicap, but it's exactly what lets the model write.

How you use it: just type. The task rides inside the prompt — “Translate to French: …”. GPT-3 (2020) stunned researchers by doing tasks it was never trained for, from a few examples pasted into the prompt (in-context learning). 100 tasks = 1 model + 100 prompts.

Technical name: autoregressive language model (a decoder).

So why did the GPT branch take over?

Generation turned out to be universal: any task can be phrased as “continue this text” — one model covers them all, no retraining, and the interface is just… language. A conversation is simply one more prompt, which is why the chatbot era is GPT-shaped.
(BERT-style models still quietly rank your search results — they just never became chatbots.)

Chapter 5 · GPT & Beyond

OpenAI's recipe: G · P · T

Generative

It creates text — word by word, like our game in Chapter 1

Pre-trained

Before you ever met it, it read a huge chunk of the internet

Transformer

Powered by the attention engine from Chapter 4 — the decoder-only variant, which does exactly one thing: predict the next token

= a Transformer that read the internet and finishes your sentences.
Sounds simple. The scale is anything but…

Chapter 5 · GPT & Beyond · Pre-training

Step 1: read (almost) everything

Training = the guess-the-next-word game on the whole internet. Each wrong guess is measured (the loss), and every weight is nudged a tiny step in the direction that reduces it — gradient descent, repeated trillions of times.

~10 trillion
words read — books, Wikipedia, code, websites
70,000+ yrs
how long YOU would need, reading 24/7 nonstop
10,000s
of GPUs running for months
$100 M+
electricity + chips for one big training run

To guess the next word on the WHOLE internet, you're forced to learn grammar, facts, logic, style, even jokes.
The guessing game was secretly the ultimate homework.

NVIDIA H100 GPU

This is what “AI hardware” looks like: an NVIDIA H100 — ~$30,000 each, and top labs use tens of thousands of them.
Photo: 极客湾Geekerwan · CC BY 3.0, via Wikimedia Commons

Chapter 5 · GPT & Beyond

Watch it write — and meet the “creativity dial”

PROMPT

“My cat” — given to raw GPT-2 (124M), which can only continue text

Temperature T = 0.7 · balanced

OUTPUT

 All three are real GPT-2 outputs (same random seed — only T changed). Mechanically: every token's score is divided by T before softmax; small T sharpens, large T flattens. Same model, different dice.
Chapter 5 · GPT & Beyond · 2020: GPT-3

By 2020, GPT-3 was amazing. Also… kind of useless.

It only knew how to CONTINUE text — not to answer you:

YouWhat should I eat for dinner tonight?
GPT-3 (raw)What should I wear tomorrow? What is the meaning of life? Why do we ask ourselves so many questions? In this essay, I will explore the nature of questions themselves…

It saw your question and thought: “ah, a LIST of questions — I'll continue the pattern.”

A brilliant engine with zero manners. One more training step was needed.

Chapter 5 · GPT & Beyond · fine-tuning

Step 2: teach the engine to be an assistant

The fix: fine-tuning

Keep training the same model — but on a new, much smaller dataset: tens of thousands of example conversations, each a question followed by a genuinely helpful answer, written and curated by people.

Why it works

The architecture doesn't change at all — it is still only predicting the next token. But after fine-tuning, the most likely continuation of a question is a helpful answer, because that is what its recent training data looks like.

Under the hood: same objective, same gradient descent — only the data changed.

YouWhat should I eat for dinner tonight?
After fine-tuningHow about a hot pot night? If you want something quick and cheap: fried rice with whatever's in the fridge. Tell me what ingredients you have and I'll suggest a recipe.

This recipe — pretrain on the internet, then fine-tune on conversations — is how GPT became ChatGPT, and it's the same recipe behind Llama, Qwen and the rest.

Chapter 5 · GPT & Beyond · Nov 30, 2022

They put it online for free. Two months later: 100 million users.

Time each product needed to reach 100 million users:

ChatGPT: fastest-growing consumer app in history at the time. Your parents, your teachers, and your group chats all discovered it the same winter.

Chapter 5 · GPT & Beyond · 2023–today

Today's model landscape

GPT-4 / GPT-5

OpenAI — the ChatGPT family

Claude

Anthropic — strong at writing and coding (it helped build these slides)

Gemini

Google — lives inside Search, Android, Docs

Llama

Meta — open weights: anyone can download & run it

Qwen · DeepSeek

Alibaba & DeepSeek — top-tier open models from China

…and yours?

Students today fine-tune small models on a laptop!

Different names, same DNA: Transformer + next-word prediction + fine-tuning.
But wait — these models can look at photos now. How does TEXT prediction do that?

The map · stop 5 of 5

Last stop: beyond text

✓ So far: LLMs have mastered reading and writing.
✗ Still broken: the world isn't made of text — it's pictures, sounds and actions. Final stretch: same brain, new senses — eyes, a paintbrush, hands.

Chapter 5 · GPT & Beyond · the variants

The LLM family tree — same brain, new superpowers

LLM — reads & writes text

VLM

+ eyes. Photos are chopped into tokens too.

GPT-4V · Claude · Gemini

Image generators

+ a paintbrush. Your words steer a noise-remover.

DALL·E · Midjourney (the artsy cousin)

VLA

+ hands. Robot movements become tokens.

RT-2 · π0 · humanoids

Agents

+ a to-do list. LLMs that use tools: browser, code, your files.

the 2025-26 frontier

The pattern, one last time: tokenize a new sense → feed the same Transformer.
Let's meet the first three up close.

Chapter 5 · GPT & Beyond · Vision-Language Models

Plot twist: images can be words too

A VLM (Vision-Language Model) chops the photo into small patches (a Vision Transformer uses 16×16-pixel squares), embeds each patch exactly like a word token, and feeds them into the same Transformer alongside your text.

 Once pictures are tokens, one brain handles both: describe photos, solve hand-written math, read menus in foreign languages, explain memes.
Chapter 5 · GPT & Beyond · Image generation

And AI art? It paints by un-noising

pure noise picture!

← drag me

Diffusion models (DALL·E, Midjourney, Stable Diffusion) train by watching millions of photos get ruined by static — and learning to reverse it.

To create art: start from pure random static, remove noise step by step, and steer with your text prompt — “a corgi astronaut, oil painting”

Théâtre D'opéra Spatial

“Théâtre D’opéra Spatial” — made with Midjourney, it won 1st prize at a 2022 art fair before judges knew. The US Copyright Office later ruled AI images can't be copyrighted — which is why I can show it to you for free
Jason M. Allen / Midjourney, 2022 · public domain, via Wikimedia Commons

Chapter 5 · GPT & Beyond · Vision-Language-Action

Final upgrade: eyes + brain + hands

A VLA (Vision-Language-Action model) outputs movements instead of words. Actions are discretized — “rotate joint 3 by +2°” becomes a token ID — so a robot trajectory is literally a sentence the Transformer can predict.

Third time's the pattern: if you can tokenize it, you can predict it — words, pixels… now actions.

Unitree robot dog

A Unitree robot dog — Google RT-2, π0, Tesla Optimus… the frontier: robots you simply talk to.

Photo: Sgt. Amber Edwards, U.S. Army, 2023 · public domain, via Wikimedia Commons

Chapter 6 · Big Questions

AI makes things up — confidently

YouGive me 2 scientific papers about the sleeping habits of dragons.
AI

Why it happens

It predicts plausible-sounding words — “sounds right” and “IS right” are not the same thing!

It's not lying

Lying needs knowing the truth. It just… autocompletes. This is called hallucination.

Your job

Verify important facts. Great assistant, terrible oracle. Especially for homework

Chapter 6 · Big Questions

The deep one: does it actually understand?

The “Chinese Room” (1980)

A person who knows ZERO Chinese sits in a room with a giant rulebook. Chinese notes slide in — they look up the symbols, follow the rules, slide perfect Chinese replies out.

Outside, it looks fluent. Inside… does anyone understand Chinese?

Philosophers still argue about this.

Team “just statistics”: it's a parrot with a gigantic memory — patterns in, patterns out.

Team “real understanding”: to predict text about the world THIS well, it must have built some internal model of the world. And hey — aren't brains pattern machines too?

There is no agreed answer. You get to make up your own mind — scientists genuinely haven't.

Chapter 6 · Big Questions

Three honest problems

Bias in, bias out

Models learn from OUR internet — including our stereotypes. Ask it to draw “a doctor” and “a nurse” and check who it draws…

Whose art is it?

Trained on millions of artists' work without asking. Artists are suing. Courts worldwide are deciding right now.

Work is changing

Calculators didn't end math careers — they changed them. The winners were people who learned to USE the new tool.

These aren't reasons to fear AI — they're reasons to understand it.
Which, congrats, you now do better than 99% of people.

Chapter 6 · Big Questions

Your turn — pick one, argue with your neighbor

1 — If an AI writes your essay and you edit it, whose work is it?

2 — Would you trust an AI doctor that's right 99% of the time — more than a human doctor who's right 95%?

3 — The Chinese Room: fluent from outside, empty inside. Is that also true… of the AI? Could someone claim it about you?

Two minutes, then we compare notes.

Wrap-up

The whole 70-year story, one line

 Notice the pattern: every breakthrough fixed the previous one's biggest weakness.
Wrap-up

If you remember only 3 things…

1

ChatGPT = next-word guessing, done astonishingly well

Everything else — essays, code, conversations — is that one trick, repeated.

2

Every breakthrough fixed the last one's weakness

Rules were brittle → learn from data. Counting lost meaning → word maps. RNNs forgot → attention. Attention was slow in RNNs → Transformer. Then: scale.

3

Words, images, actions — it's all tokens now

GPT writes, VLMs see, VLAs act. Same engine, new senses. The story is still being written — possibly by someone in this room.

Wrap-up

Want more? Start here (all free)

3Blue1Brown

“Large Language Models explained briefly” — the most beautiful animations on this topic. Start here.

Andrej Karpathy

“Intro to Large Language Models” (1 hr) — by the legend himself. Then “Let's build GPT” when you're ready to code one!

Jay Alammar

“The Illustrated Transformer” — the classic visual walkthrough, one diagram at a time.

Play right now

ChatGPT / Claude / Gemini are free — ask one to explain today's talk back to you — then grade it.

Build one

Know some Python? Search “nanoGPT” — you can train a tiny GPT on your laptop, on Shakespeare, tonight.

Thank you!

Questions?
the entire field started with “can machines think?” and nobody has answered THAT yet either.

Slides: one self-contained HTML file — arrow keys to explore, every demo is clickable. Take it home!