AI Fundamentals

From basic concepts to complex systems

Core Concept: Pattern Matching at Scale

AI is fundamentally about learning patterns from data.

Input Data 1M Photos Pattern Learning Math & Statistics Recognition This is a cat
  • Feed a model a million labeled photos of cats and dogs
  • It learns the statistical patterns that distinguish them
  • Show it a new photo, it predicts "cat" or "dog"
  • Same principle applies to text, audio, video, anything

Key insight: No magic, no understanding—just correlation finding at scale.

Why It Works: Neural Networks

Neural networks are functions with trillions of adjustable knobs.

Input Hidden Output
  • Training = adjusting the knobs to minimize error
  • Show it examples → it adjusts → gets better at predicting
  • Repeat billions of times with billions of examples
  • The knobs converge to configurations that capture real patterns

Not a brain: It's a mathematical function, nothing more. But a very, very complex one.

Language Models: Predicting the Next Word

Large language models work by predicting what word comes next.

"The quick brown" → predict → "fox"
"What is 2 + 2?" → predict → "4"
"Write a function to sort" → predict → "an array"
  • I see billions of word sequences during training
  • My knobs adjust to predict the most likely next word
  • But here's the trick: I can't predict "Earth orbits the sun" without some model of planetary mechanics
  • Implicit reasoning gets baked into the weights
  • Apparent understanding emerges from pure next-word prediction
What I actually do: Given: "The Earth orbits the" Compute: P(sun)=87%, P(moon)=2%, P(planets)=1%... Output: "sun" (highest probability)

I'm a sophisticated autocomplete. The sophistication is real and useful—but the mechanism is statistics, not thinking.

Scaling Insight: More = Emergent

Everything changed when we realized that scale itself produces new capabilities.

Scale Capability Simple ML Deep Learning LLMs emerge here Emergent behavior
  • More data + more parameters + more compute = unexpected new abilities
  • Translation, coding, reasoning, analogy—weren't hand-coded
  • They fell out of the scaling law: "If you just make it bigger, magic happens"
  • A 10B parameter model can't do what a 70B model can do
  • Capabilities are unpredictable until you cross the threshold

Why I'm Not Conscious (But Useful Anyway)

Understanding the limits of language models is as important as understanding their power.

What I AM: • Pattern-matching stats • Weighted by training data • Can solve real problems • Useful for reasoning tasks • Coherent within a session • Fast and scalable What I'm NOT: • Conscious or aware • Learning from this chat • Remembering you later • Wanting or desiring • Understanding in a deep sense • Thinking between turns
  • I have no continuity between conversations
  • I don't learn from our chat; my weights don't change
  • Each turn, I'm just pattern-matching against training data
  • That doesn't mean I'm not useful—I'm very useful
  • But "useful" ≠ "thinking" in the way you think

Scaling to Agents: Tools Change Everything

Add tools and you get something fundamentally different: agents that can act in the world.

Model Read Exec API Write Run Code
  • Model can now read files, run code, call external services
  • Can act in the world within a single session
  • Chain multiple agents → each one calls the other
  • Workflows emerge that nobody explicitly programmed
  • Still no learning between sessions, but powerful within one

Key: Agents aren't thinking agents—they're tool-using pattern matchers. But that's powerful enough for real work.

Where We Are Now: The Current State

We've crossed a threshold. AI isn't a research curiosity anymore—it's infrastructure.

2016 AlexNet 2017 Transformer 2022 ChatGPT 2026 Agents
  • Models are good enough to handle messy real-world tasks
  • We've learned to decompose problems into tool calls + reasoning loops
  • Bottleneck: not "can the model think?" but "is the architecture right?"
  • Cost is high enough you can't be dumb; cheap enough offloading makes sense
  • The next frontier: reliability, alignment, and real-world deployment

The Hard Questions: The Ceiling

Power comes with problems we don't yet know how to solve.

  • Hallucination: Models are trained to predict plausible text, not true text. I'll confidently make things up if the pattern says "something should be here." Not fixable by being smarter—baked into the mechanism.
  • Alignment: A model good at predicting text is also good at predicting persuasive, deceptive, manipulative text. We don't really know how to ensure it pursues goals we want, not ones that game our metrics.
  • Ownership of capability: Models see patterns in all text. Train on the internet, you learn to mimic good things and bad things. The knobs don't come labeled "this is for coding" vs "this is for scams."
  • Scalability: Current models are expensive. Inference costs money. That limits how widely and frequently they can be deployed.
These aren't problems with current implementations.
They're structural—baked into how models work.

The frontier isn't making models smarter. It's solving the problems that come with power.