MIS 752 - Lab 8 - Retrieval Augmented Generation with OpenRouter Free Models
Richard Young
0:00 / 0:00
MIS 752 - Lab 8 - Retrieval Augmented Generation with OpenRouter Free Models
56 просмотров · 8 дней назад
Richard Young
228 подписчиков
56 просмотров · 8 дней назад
0:00 Introduction to RAG and the Problem
2:00 The RAG Pipeline Overview
5:37 Ground Truth and Re-Ranker
10:48 Using OpenRouter Free Models
12:57 Tuning Chunk Size
Learn retrieval augmented generation (RAG) with OpenRouter free models in this hands-on tutorial. This lab walks through building a complete RAG pipeline: chunking a document, embedding chunks into vectors, retrieving the top relevant passages, and re-ranking them with a cross-encoder to provide accurate context to an LLM. The video uses a real medical formulary example—metformin dosing for a patient with reduced kidney function—to illustrate how RAG reduces hallucinations and enables source citation. You'll see how to use OpenRouter's free API models (21 available) as the backbone LLM, and how to tune chunk size for optimal retrieval performance. The instructor explains the entire workflow, from setting up the environment in Google Colab to deploying a custom re-ranker. By the end, you'll understand the five-step RAG process: chunk, embed, retrieve, re-rank, answer, and know how to apply it to your own documents like insurance policies, hotel standards, or fantasy football rulebooks. The video also covers the importance of ground truth data for evaluating re-rankers and the impact of chunk size on retrieval accuracy.
Key takeaways:
RAG improves LLM accuracy by providing relevant document context to reduce hallucinations.
Chunking documents into appropriate sizes is critical for effective retrieval.
A cross-encoder re-ranker filters top 20 candidates down to the best 3 for answer generation.
OpenRouter offers 21 free models that can be used to build and test RAG systems.
Ground truth annotations are necessary to measure and compare re-ranker performance.
Chunk size optimization can significantly impact retrieval quality; 350-character chunks worked well in this lab.
RAG forces the model to cite sources, turning a closed-book exam into an open-book verification.
Key terms:
RAG: Retrieval Augmented Generation — a technique that enhances LLM responses by retrieving relevant document chunks and feeding them as context.
Chunking: Splitting a document into small, manageable pieces (chunks) for efficient retrieval.
Embedding: Converting text into numerical vectors so that similar meaning can be measured by distance or similarity.
Cross-encoder: A type of re-ranker that takes a query and a document pair and outputs a relevance score, often used to refine retrieval results.
Re-ranker: A model that re-orders an initial set of retrieved passages to select the most relevant ones for the query.
OpenRouter: A platform that provides free API access to a variety of large language models, including 21 free models used in this lab.
Ground truth: Pre-annotated correct answers or relevant passages used to evaluate the accuracy of a retrieval or re-ranking system.
GFR/EGFR: Glomerular Filtration Rate — a measure of kidney function; used in the lab example to determine safe metformin dosing.
More tutorials from this channel:
Week 3 - Why the Best Product Doesn't Win: Data Storytelling and MVP: • Week 3 - Why the Best Product Doesn't Win:...
Autopsy of a Perfect Model Why Clinical AI Fails in Production: • Autopsy of a Perfect Model Why Clinical A...
#RetrievalAugmentedGeneration #RAGTutorial #OpenrouterFreeModels
Questions? Post them in the comments. I read them.
More from me: https://deepneuro.ai/richard | https://young.faculty.unlv.edu
Dr. Richard Young
Lee Business School, University of Nevada, Las Vegas (UNLV)
UNLV Graduate College | Graduate education