Skip to main content

Command Palette

Search for a command to run...

How LLMs Actually Work — An Easy Guide

Updated
•5 min read•View as Markdown
How LLMs Actually Work — An Easy Guide

“Large Language Models aren’t magic. They’re just really good at predicting the next word.”

Imagine you’re talking to an extremely smart friend who has read millions of books, newspapers, WhatsApp messages, and movie subtitles in every language. That’s what an LLM is — a model trained to understand and respond like a human.


🇮🇳 1. What is an LLM?

LLM = Large Language Model
It’s a type of Artificial Intelligence that can:

  • Understand human language 🧏‍♂️

  • Generate meaningful text 📝

  • Hold conversations 🤝

Think of it like a digital brain trained on a massive library of the internet.

Popular examples:

  • ChatGPT

  • Claude

  • Gemini

These models power:

  • 🤖 Chatbots

  • 👨‍🏫 AI tutors

  • 👩‍💻 Code assistants

  • 🧾 Business analysis tools

They’re fast ⚡, multilingual 🌏, and always available 🕒.


🪙 2. How the LLM Pipeline Works

When you type something into an AI chatbot, here’s what actually happens behind the curtain:

preview

🪄 Analogy:
Imagine sending a courier through the Indian postal system:

  1. You write letters (your prompt).

  2. The post office breaks it down into packets (tokenization).

  3. They put addresses on each (embeddings).

  4. A network of sorting centers processes it (transformer).

  5. The right packet reaches the right person (prediction).

  6. The receiver rebuilds the message (text output).


🧩 3. Step-by-Step Inside the LLM

Step 1: Tokenization – 📦 Breaking the Message

LLMs can’t read words directly — they break everything down into tokens (words or subwords).

Example:

"The most famous cheese in Paris."
→ ["The", " most", " famous", " cheese", " in", " Paris", "."]

🪄 Indian Analogy:
Like breaking a large Laddu into small pieces to share with friends — easier to handle, same sweetness.

preview


Step 2: Embeddings – 🧭 Giving Tokens a Meaning

Each token is converted into a vector — a list of numbers.
This is how the model understands meaning mathematically.

"I love AI"
Tokens → [I] [love] [AI]
Vectors → [0.21, 0.55, 0.71]  (example numbers)

🪄 Indian Analogy:
Think of each token like a parcel in a courier system. Embedding is the address and PIN code — without it, the system wouldn’t know where to send it.

preview


Step 3: Transformer (Attention) – 🕸️ The Brain of the Model

This is where the real intelligence kicks in.
The Transformer architecture uses attention to find relationships between words.

Example:

  • “Cheese” and “Paris” are related → Transformer gives them more weight.

🪄 Indian Analogy:
Like a wedding in India — the bride, groom, relatives, and friends are all connected. Everyone is paying more attention to certain key people (the bride & groom) but still knows who else is in the room.

preview


Step 4: Prediction – 🎯 Choosing the Next Word

Now the model predicts the most likely next token based on:

  • Context

  • Learned probabilities

Input → Tokenization → Embeddings → Transformer → Output Embedding → Tokens → Text

It does this one token at a time, very fast.

🪄 Indian Analogy:
Like how we finish each other’s sentences in a conversation —
If someone says “Chai garam garam…” you instantly think of “aayi hai” ☕.

preview


🧠 4. How LLMs Balance Accuracy & Creativity

LLMs don’t always pick the top token. That would make their answers sound robotic.
They use sampling strategies to make outputs more natural and human-like.


🌡️ Temperature – Controlling Randomness

  • Low temperature (e.g., 0.2) → Safe, factual, predictable

  • High temperature (e.g., 0.9) → More creative, unpredictable

🪄 Indian Analogy:
Like ordering tea:

  • Low temp = Fixed “cutting chai” from your regular tapri ☕

  • High temp = Experimental “kulhad chai with elaichi and masala” 😄

⚠️ Too high → model may hallucinate or give silly answers.

preview


🏆 Top-K Sampling – Pick the Top Few

  • Take only the K most probable tokens (e.g., top 40).

  • Randomly pick one from them.

  • Discard the rest.

🪄 Analogy:
Like choosing your favourite samosa shop among 5 trusted ones, not from 100 random stalls.

preview


🌊 Top-P (Nucleus) Sampling – Pick Smartly

  • Take the smallest set of tokens whose cumulative probability ≥ P (e.g., 0.9).

  • Pick one from that set.

  • More adaptive than Top-K.

🪄 Analogy:
Like inviting only your closest 20 relatives to a wedding, not all 200 people you know.

preview


🪶 Min-P Sampling – Cut the Junk

  • Discard tokens below a minimum probability threshold.

  • This avoids picking unlikely nonsense tokens.

🪄 Analogy:
Like removing burnt pakoras from the oil before serving 😄

preview


🏋️ 5. Training vs. Inference

  • Training:
    The model learns from massive datasets.
    It’s like a student reading all NCERT books, newspapers, novels, and WhatsApp forwards for years.

  • Inference:
    The model uses what it has learned to predict answers in real time.
    Like giving an answer in a viva exam based on everything studied earlier.

Training Data + Transformer → Trained Model
Prompt + Sampling → Final Output

preview

🪄 Quick Revision (TLDR)

preview

StepWhat HappensAnalogy (India)
TokenizationBreak text into piecesBreaking a laddu into small bits 🍬
EmbeddingConvert tokens to vectorsGiving each parcel an address 📮
TransformerFind relationships between tokensPaying attention at a wedding 👰🤵
PredictionChoose the next wordFinishing a friend’s sentence 🗣️
TemperatureControl randomnessTapri chai vs. experimental kulhad chai ☕
Top-K / Top-P / Min-PSampling strategiesChoosing from trusted samosa shops 🍽️

🌟 Final Thoughts

LLMs are not thinking like humans.
They’re just very powerful pattern predictors trained on a mountain of data.

But when you understand their building blocks — tokenization, embeddings, transformers, and sampling — they stop being “mysterious” and start being understandable.

👉 Every time you chat with ChatGPT, this exact dance of tokens, vectors, and probabilities is happening behind the scenes.