Embedding every paper costs less than one day of keeping search online
Kemmu Draws Tech
0:00 / 0:00
Embedding every paper costs less than one day of keeping search online
4 просмотра · 7 дней назад
Kemmu Draws Tech
9 подписчиков
4 просмотра · 7 дней назад
Hugging Face's write-up on Papers with Code search splits the stack three ways: a Job runs the batch embedding pass, an Inference Endpoint encodes live queries, and a Bucket holds the vectors. This episode prices each leg off Hugging Face's own published rate card and works out how the bill really divides between building the index, storing it, and waiting for someone to type.
Chapters
0:00 The guess everyone makes
0:40 Three products, one GPU
2:06 What the reading costs
3:24 What the waiting costs
4:43 What the storing costs
6:40 The bill
7:27 The only lever
Sources
How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code — https://huggingface.co/blog/pwc-search
Pricing and Billing · Hugging Face — https://huggingface.co/docs/hub/en/jo...
Pricing · Hugging Face — https://huggingface.co/docs/inference...
Storage limits · Hugging Face — https://huggingface.co/docs/hub/en/st...
Storage Buckets · Hugging Face — https://huggingface.co/docs/hub/en/st...
Drawn with Inkstack.