Перейти к содержимому

Lab 10 - Why Regex Missed 60% of Sensitive Data in This OpenRouter Test | 60% FAILED

Richard Young

0:00 / 0:00

Lab 10 - Why Regex Missed 60% of Sensitive Data in This OpenRouter Test | 60% FAILED

3 просмотра · 12 часов назад
Richard Young
228 подписчиков
3 просмотра · 12 часов назад
0:00 Introduction & Lab Overview 3:24 Setting Up OpenRouter and Dependencies 5:24 Layer One: Regex Filter 7:04 Layer Two: LLM Judge 9:45 Layer Three: Answer Scoring 13:54 Audit Results and Tradeoffs 15:47 Exporting as Python and Tuning In this OpenRouter tutorial, you will learn how to build AI guardrails that block sensitive data like SSNs, phone numbers, and emails in a pharmacy application. We implement three validation layers: a fast regex filter, an LLM judge using a model like Qwen 2.5 via OpenRouter, and a final scoring layer that evaluates answer quality. Through hands-on testing with synthetic data, you'll see why regex catches only 40% of sensitive items, why an LLM can have perfect recall but high over-refusal, and how a third layer balances safety and usability. The demonstration uses a pharmacy counter example where we measure refusal rates, wrongly blocked requests, and items that leak through. You'll learn to tune the refusal threshold to match your risk tolerance—whether you're protecting healthcare data or any other high-stakes application. We also walk through exporting the notebook as a Python file for easier iteration in ChatGPT or other LLMs. By the end, you'll understand the tradeoffs between speed, accuracy, and safety when designing guardrails, and how to choose where to err when the stakes are high. Key takeaways: Build three validation layers: regex, LLM judge, and answer scoring. Regex is fast and fully auditable but inaccurate; it missed 6 of 10 required refusals. An LLM judge caught all 10 refusals but wrongly blocked 8 legitimate requests (over-refusal). Layer three scores answers to catch cleverly worded harmful inputs while allowing safe ones. Tune the refusal threshold to balance data protection against user access. Export your notebook as a Python file to iterate faster in ChatGPT or other LLMs. Key terms: Guardrails: Predefined rules or validation layers that block or flag unsafe AI inputs and outputs, such as sensitive data or harmful instructions. Regex: Regular expression – a pattern-matching technique used to quickly detect specific strings like phone numbers or SSNs. LLM Judge: A language model (e.g., Qwen 2.5) that evaluates whether an input or output should be refused based on context. Layer 1 / Layer 2 / Layer 3: Three validation stages: Layer 1 uses regex for fast blocking; Layer 2 uses an LLM for contextual judgment; Layer 3 scores the answer quality. Refusal Rate: Percentage of requests the guardrail system blocks. High refusal can prevent data leaks but may also block legitimate use. Over-refusal: When a guardrail incorrectly blocks a safe request, causing user frustration or harm (e.g., a patient not getting needed medication information). Data Leakage: Exposure of sensitive information (SSN, phone, email, MRN) through AI responses that should have been blocked. OpenRouter: A service that provides access to multiple LLMs (like Qwen, Llama) via a single API, used in this video to call the judging model. More tutorials from this channel: De-identify Patient Data with Python, LLMs & Hugging Face | Lab 2:    • De-identify Patient Data with Python, LLMs...   #OpenrouterTutorial #OpenrouterFreeModels #Openrouter Questions? Post them in the comments. I read them. More from me: https://deepneuro.ai/richard | https://young.faculty.unlv.edu Dr. Richard Young Lee Business School, University of Nevada, Las Vegas (UNLV) UNLV Graduate College | Graduate education ##