Перейти к содержимому

Apache Pinot Powers Millisecond Scale Analytics | TechTalk Podcast Ep. 12

Codestreamlab

0:00 / 0:00

Apache Pinot Powers Millisecond Scale Analytics | TechTalk Podcast Ep. 12

15 просмотров · 11 дней назад
Codestreamlab
4 подписчика
15 просмотров · 11 дней назад
Welcome! In this video, we dive deep into **Apache Pinot**, the open-source, column-oriented, distributed OLAP data store written in Java designed for ultra-low latency, real-time analytics at massive scale[1][2]. Originally developed at LinkedIn to overcome the limits of traditional relational and key-value stores for products like Who's Viewed Your Profile and internal A/B testing analytics[3][4], Apache Pinot powers high-concurrency, sub-second queries across massive streaming and batch datasets[2]. --- 📌 Timestamps / Chapter Outline: *00:00* \- Introduction to Apache Pinot & Real-Time OLAP[2] *02:15* \- Core Architecture: Controllers, Brokers, Servers, and Minions[7] *05:30* \- Cluster Management with Apache Helix & ZooKeeper[7][8] *08:10* \- Ingestion Pipelines: Streaming (Kafka, Pulsar, Kinesis) vs. Batch (S3, GCS, Hadoop)[9] *11:45* \- How Consuming Segments Work for Real-Time Streaming Data[10][13] *14:20* \- Pluggable Indexing: Star-Tree, Inverted, Bitmap, and Range Indexes[14] *17:50* \- Visualizing Real-Time Metrics with Apache Superset[15][16] *21:10* \- LinkedIn Engineering Case Study & Scaling Lessons[3][4] --- 🚀 Key Topics & Features Covered: *Distributed Architecture:* Managed by Apache Helix for cluster orchestration and ZooKeeper for maintaining cluster state and metadata[7]. *Broker & Query Routing:* Brokers scatter SQL queries across offline and real-time servers, then aggregate results to serve low-latency responses[8][18]. *Real-Time Stream Ingestion:* Ingests streaming data directly into memory via low-level consumers (e.g., Kafka) using consuming segments that periodically flush to deep storage[10]. *Pluggable Indexing:* Optimizes query performance using advanced indexing technologies such as Sorted, Bitmap, Inverted, and Star-Tree indexing[14]. *Visualization:* Integrates smoothly with tools like Apache Superset to create dashboards with auto-refreshing real-time streaming data[15][16]. --- 🔖Hashtags: #ApachePinot #RealTimeAnalytics #BigData #DataEngineering #ApacheKafka #ApacheSuperset #OLAP #DistributedSystems #LinkedInEngineering