Generative AI & LLM Foundations | Build Toward a Small GPT | SRAI Book 4 Lesson 1
SRAI — Statistics, Reasoning and AI
0:00 / 0:00
Generative AI & LLM Foundations | Build Toward a Small GPT | SRAI Book 4 Lesson 1
38 просмотров · 13 дней назад
SRAI — Statistics, Reasoning and AI
16 подписчиков
38 просмотров · 13 дней назад
How do large language models generate text, and why does fluent output not automatically constitute reliable evidence?
This opening lesson of SRAI Book 4 introduces the technical and governance foundations of generative AI and large language models. It explains how text becomes tokens, how tokens become vector representations, how autoregressive models predict successive tokens, and how attention combines information across a retained context.
The lesson also examines the practical limitations of language models, including confabulation, unsupported citations, altered numbers, missing qualifications, context-window constraints, bias, privacy risks and prompt injection.
Learners will explore:
• Generative and discriminative learning
• Tokenization, identifiers and embeddings
• Context windows and information boundaries
• Autoregressive next-token prediction
• Scaled dot-product attention
• Prompt assembly and decoding policies
• Temperature, top-k and top-p sampling
• Retrieval-augmented generation
• Claim-level evidence and verification
• Tool calling and agentic control boundaries
• Groundedness, correctness and decision fitness
• Privacy, security and multilingual evaluation
• The SRAI generative-AI evidence chain
The accompanying practical notebook contains 41 cells, including 21 executable code cells. It has been validated in both Visual Studio Code and Google Colab and does not require a private SRAI software package.
BOOK 4 LEARNING DESTINATION
Across the complete Book 4 pathway, ambitious learners will progress from these foundations toward building, training, evaluating and explaining a small GPT-style language model.
This is an educational model rather than a frontier-scale system. Learners will nevertheless work with the essential mechanisms: tokenization, embeddings, causal attention, decoder blocks, training loss, text generation and systematic evaluation.
SRAI PRINCIPLE NO. 96
A language model generates statistically plausible continuations. Factual authority must come from evidence, validation and accountable human use.
Production Unit: PU-B04-C01
Book: Book 4 — Generative AI and Large Language Models
Lesson: Lesson 1 — Generative AI and LLM Foundations
Instructor: Mbaye Kebe
Ecosystem: Statistical Research and AI Institute — SRAI
Supporting learning materials include the controlled lesson, executable notebook, exercises, solutions and marking guide, executive brief and presentation.
#GenerativeAI #LargeLanguageModels #LLM #GPT #Transformers #MachineLearning #ArtificialIntelligence #ResponsibleAI #AIGovernance #DataScience #SRAI