3 months ago|
Generative Ai

Gen AI Explained

Gen AI in a nutshell

Blog Image

Generative AI has moved from laboratory curiosity to board-level infrastructure in a handful of years. Most executives now use it daily. Far fewer can explain what is actually happening when a model writes a strategy memo, drafts code, or produces an image from a sentence.

That gap matters. If you cannot name the underlying technologies, you cannot judge risk, cost, quality, or where the next capability leap will come from. This article sets out the core machinery in plain language, without pretending the mathematics is trivial, and without requiring a PhD to follow it.

What generative AI actually is

Generative AI is a family of models that learn the statistical structure of data and then sample from that structure to produce new artefacts: text, code, images, audio, video, or combinations of those.

It is not a search engine. It does not retrieve a stored document and hand it back. It estimates what is likely to come next or what a coherent sample should look like, given everything it has seen during training and everything you have just provided as context.

Two consequences follow immediately.

First, fluency is not the same as truth. A model can produce a polished paragraph that is factually wrong because the training objective rewarded plausibility, not verification.

Second, “intelligence” in these systems is tightly bound to architecture and data. Change the architecture, the objective, or the post-training process, and behaviour changes with it.

The stack, from the ground up

1. Tokenisation: turning the world into units a model can count

Models do not read words as humans do. Text is first broken into tokens sub-word chunks such as “gener”, “ative”, and “ AI”. Images are often compressed into latent patches or discrete visual tokens. Audio is sliced into frames or codec tokens.

Tokenisation is unglamorous and decisive. It determines vocabulary size, how well the model handles rare terms and other languages, how expensive a prompt is, and how much meaning survives the first conversion into numbers.

Those tokens are then mapped into embeddings: dense vectors in a high-dimensional space, where “bank” near “river” sits in a different neighbourhood from “bank” near “interest rate”. Everything afterwards is arithmetic on those vectors.

2. The transformer: the architecture that made scale possible

Almost every leading language model, and a growing share of image and video systems is built on the transformer, introduced in the 2017 paper Attention Is All You Need.

Earlier sequence models (RNNs and LSTMs) processed tokens one after another. That made long-range context expensive and training hard to parallelise. Transformers discarded recurrence. They let every token look at every other token in the sequence at once.

The key mechanism is self-attention.

For each token the model creates three learned projections:

  • a query “what am I looking for?”
  • a key “what do I contain?”
  • a value “what information should I pass on if I am relevant?”

Relevance is measured by comparing queries with keys (scaled dot-product attention). The result is a set of weights. Those weights are used to mix the values. In effect, the representation of the word “it” can absorb information from “the animal” several sentences earlier, because attention discovered that the two belong together.

A few design details make this usable at industrial scale:

  • Multi-head attention runs several of these comparisons in parallel, so different heads can specialise (syntax, coreference, topic, code structure).
  • Positional information is added because raw attention is permutation-sensitive; modern systems typically use rotary position embeddings (RoPE) rather than the original sinusoidal encodings.
  • Stacked layers refine representations step by step. A large model is not one clever trick. It is dozens or hundreds of attention and feed-forward blocks, each adjusting the embeddings a little further.

Most production language models today are decoder-only: they generate left to right, each new token conditioned on all previous tokens. Encoder–decoder designs still appear in translation and some multimodal pipelines, but the dominant generative pattern for text is next-token prediction.

That single objective predict the next token well, at enormous scale turned out to be a surprisingly general learning signal.

3. Large language models: scale, pre-training and emergence

A large language model is a transformer trained on vast corpora of text and code until its parameters encode statistical regularities of language, knowledge and form.

Training usually proceeds in stages.

Pre-training is self-supervised. The model sees unfinished sequences and must predict what follows. No human labels are required at this stage. The result is a foundation model: fluent, broad, and not yet a product. Ask a raw base model a question and it may continue the question rather than answer it.

Supervised fine-tuning (SFT) then teaches the model the shape of useful behaviour: question then answer, instruction then completion, tool call then result. Relatively small, high-quality datasets can shift a base model from “text completer” to “assistant”.

Preference alignment goes further. In classic RLHF (reinforcement learning from human feedback), people rank candidate answers; a reward model learns those rankings; the language model is then optimised to produce answers the reward model scores highly, usually with a penalty that stops it drifting too far from the SFT policy. Variants now include direct preference optimisation and reinforcement learning with verifiable rewards (useful where an answer can be checked, as in mathematics or code).

This pipeline explains a great deal of product behaviour. Helpfulness, refusal style, verbosity and “personality” are largely post-training choices, not mysterious properties of the transformer itself.

Two further engineering patterns now dominate the frontier:

  • Mixture of Experts (MoE) routes each token through a subset of specialised feed-forward networks. Total parameter counts can be enormous while the compute per token stays manageable.
  • Multimodal models share or align transformers across text, images, audio and video, so a single system can read a chart, write a summary and propose the next slide.

4. Diffusion models: how images (and increasingly video) are made

Text generation is mostly autoregressive: one token after another. High-quality image generation took a different path.

Diffusion models learn to reverse a gradual noising process. During training, clean images are corrupted with Gaussian noise, step by step, until they become static. The network learns to denoise — to estimate and remove that noise — conditioned on a text embedding of the prompt. At generation time the model starts from noise and walks backwards until a coherent image appears.

Why this displaced earlier methods:

  • GANs (generative adversarial networks) pit a generator against a discriminator. They can be fast at inference, but training is unstable and they are prone to mode collapse — producing a narrow range of samples.
  • VAEs (variational autoencoders) learn a compressed latent space and decode from it. Training is more stable, but naïve VAEs often yield blurrier images.

Modern systems usually combine ideas. Latent diffusion first compresses the image with an autoencoder, then runs the expensive denoising process in that smaller latent space. That is why high-resolution synthesis became practical. Increasingly, the denoising backbone itself is a transformer (Diffusion Transformers, or DiTs) rather than a convolutional U-Net.

Video models extend the same logic across space and time — which is why they are so computationally hungry.

5. The supporting machinery organisations actually feel

The headline architectures sit on a stack that determines cost and reliability:

  • Context windows how much of a document, codebase or conversation the model can attend to at once.
  • Retrieval-augmented generation (RAG) fetching approved documents at query time so the model is grounded in your corpus rather than only in training data.
  • Tool use and agents letting the model call search, databases, code interpreters or workflow systems instead of improvising an answer.
  • Guardrails and evaluation toxicity filters, policy classifiers, factuality checks, red-teaming and human review.

From an adoption standpoint, this last layer is where most programmes succeed or stall. The model is a component. The system around it is the product.

What this means for organisations

Understanding the stack changes how you buy, govern and build.

Quality failures have different causes. Hallucinations are a predictable result of next-token sampling without retrieval or tools. Blurry or inconsistent images often trace to latent compression or insufficient conditioning. Sycophantic or overly cautious answers are frequently alignment artefacts. Treat the symptom as a systems problem, not a personality problem.

Cost follows architecture. Autoregressive decoding is sequential, so long outputs are expensive. Diffusion is iterative, so image and video quality costs steps and compute. MoE models shift the trade-off between total capacity and active compute. Procurement conversations should include tokens, context length, latency and energy not only “which brand”.

Data and post-training are strategy. Two organisations can use the same base model and obtain very different behaviour once fine-tuning, RAG corpora, tools and evaluation harnesses diverge. Competitive advantage increasingly sits in proprietary data, workflow integration and evaluation, not in access to a public chatbot.

Governance has to match the mechanism. If the model samples from a probability distribution, you need human accountability for high-stakes outputs, provenance for training and retrieval data, and clear rules for synthetic content. Ethics is not an appendix. It is an operating constraint on systems that can generate convincing falsehoods at scale.

A short map of the field as it stands

FamilyCore ideaTypical outputsStrengthLimitationTransformers / LLMsAttention + next-token predictionText, code, reasoning, some multimodalGeneral-purpose, scalableHallucination; sequential decoding costDiffusionIterative denoisingImages, video, audioHigh fidelity, controllableSlow sampling; compute-heavyGANsGenerator vs discriminatorImages, style transferFast inferenceUnstable training; mode collapseVAEsEncode decode a latent distribution Compression, latents, some generationStable, structured latentsOften less sharp samples


The commercially important pattern in 2026 is hybrid: transformers for language and control, latent diffusion or autoregressive visual tokens for media, retrieval and tools for grounding, and post-training for behaviour.

Conclusion

Generative AI looks like magic only if you stop at the interface. Underneath it is a small set of ideas executed at extreme scale: tokens and embeddings, attention, next-token prediction, denoising, and a training pipeline that turns raw statistical models into systems people can work with.

Leaders do not need to derive the attention equation. They do need enough fluency to ask better questions: What is the base model? How is it grounded? What was the alignment objective? Where does retrieval stop and generation start? Who evaluates the outputs that matter?

Those questions separate theatre from adoption.

If this briefing is useful, I write and teach on AI strategy, agentic workflows and responsible implementation including practical programmes inside The AI Training Academy.



AI Strategy and Adoption Specialist


IMD Lausanne and Zurich

What part of the stack is least well understood in your organisation — the models themselves, or the operating system you wrap around them?

Share on:

0 comments

No comments yet

Your Views Please!

Your email address will not be published. Required fields are marked *
Please Login to Comment

You need to be logged in to post a comment on this blog post.

Login Sign Up